Spark Notes
✓ Published0🌍 Public
CCliffordAnderson
Last edited Feb 14, 2020
Created on Feb 14, 2020
This example shows a text-analytics pipeline that annotates a CSV-loaded corpus of academic articles. It renames raw columns to meaningful fields such as `article`, `journal`, `date`, and `text`, then uses Spark NLP’s `PretrainedPipeline("explain_document_ml")` to transform the data, adding tokens and other linguistic annotations. The code demonstrates schema inspection via `printSchema` and previews the resulting token column with `select("token").show()`, highlighting how Spark DataFrames integrate with a pretrained NLP model for scalable document processing.
AI-generated description