Gist fbc8e33e01b5be01bc6c1ab2d1d08cfc
✓ Published0🌍 Public
AArizonaTay
Last edited Dec 27, 2021
Created on Dec 27, 2021
The visualization displays the 30 most frequently occurring words from a set of cleaned Instagram captions, plotted as a bar chart that ranks terms by raw count. It shows a sharp drop-off in frequency after the top few words, highlighting the long-tail distribution common in text data. The chart is generated using Python’s matplotlib, with the caption data loaded into a pandas DataFrame and tokenized via the Natural Language Toolkit (nltk). Frequencies are computed with `nltk.FreqDist`, and the top 30 entries are selected for plotting, with a filter excluding the characters ‘1’, ‘2’, and ‘3’ from the hashtag tokens. The figure is sized at 14 by 6 inches for readability.
AI-generated description