import seenthis into gensim
This example demonstrates how to transform a dump of the social network Seenthis into a format usable by the gensim topic-modeling library, specifically adapting the Wikipedia-corpus script to generate bag-of-words and TF-IDF vector representations. The code shows a pipeline that reads raw posts, strips markup using regex-based cleaning (including removal of comments, templates, and links), tokenizes the text, and then builds a dictionary and sparse matrix outputs in Matrix Market format. It relies on gensim’s `Dictionary`, `TextCorpus`, and `MmCorpus` classes, plus utility functions like `filter_wiki` for preprocessing, mirroring the original `wikicorpus.py` structure. The output files—a word-ID mapping and two corpus files—allow downstream topic modeling with gensim, though the visualization itself is not shown; the code’s focus is on data preparation rather than graphical display.
AI-generated description