Gist 7518328
✓ Published0🌍 Public
TTheLady
Last edited Jan 13, 2012
Created on Nov 17, 2013
This example computes and displays the most relevant terms (including single words, bigrams, and trigrams) from a pair of documents by calculating term frequency–inverse document frequency (TF-IDF). It first tokenizes Portuguese text using NLTK’s `RegexpTokenizer` and removes stopwords, then builds a vocabulary of unigrams, bigrams, and trigrams using `nltk.bigrams` and `nltk.trigrams`. For each document, the code outputs each term’s raw frequency, TF, IDF, and TF-IDF score, and finally prints a ranked list of the highest scoring terms across all documents. The script relies solely on Python’s `math` module and NLTK, with no external visualization library.
AI-generated description