Skip to main content
100%

Gist 55fc864ec9cea1107845d3ed7a3dab8b

✓ Published0🌍 Public
AArizonaTay
Last edited Dec 27, 2021
Created on Dec 27, 2021

This example demonstrates a text-cleaning pipeline for caption data, showing how raw captions are converted into tokenized, filtered lists for downstream analysis. The visualization itself is not rendered; instead, the focus is on the data preprocessing logic. The code uses `pandas` with `apply` to run natural language processing via `spaCy` (`nlp`), extracting tokens from the `Caption` column. It then applies a custom `tokens` function that removes stop words and non-alphanumeric characters using `re.match`, while preserving tokens containing "@" symbols. The result is a new `clean_caption` column, highlighting a common step in text mining workflows.

AI-generated description

Similar vizzes