Skip to main content
100%

Gist e14e45693dae94f60b953c09a641f6a4

✓ Published0🌍 Public
AArizonaTay
Last edited Dec 27, 2021
Created on Dec 27, 2021

This example shows a data-processing pipeline that identifies and extracts potential sponsor names and tags from caption text. It classifies each row as “Yes” or “No” for sponsorship based on whether any sponsor or tag tokens are present. The code uses Python’s `re` module for emoji removal and text cleaning, and relies on spaCy’s named entity recognition (via `list_words.ents`) to detect organizations and products. It applies custom string-manipulation functions like `remove_emoji`, `pairList`, and `getTag` to filter and deduplicate extracted entities, then combines sponsor and tag columns before checking for non-empty results.

AI-generated description

Similar vizzes