Gist d003ece0b6ca171be8e095e0bc90af7e
✓ Published0🌍 Public
GGeorgeMcIntire
Last edited Oct 9, 2017
Created on Oct 9, 2017
This notebook example demonstrates how to programmatically scrape, parse, and structure a corpus of New York Times articles into a pandas DataFrame. The code loads a pre-saved pickle file of 50 article links, iterates through each URL using the `newspaper` library's `Article` class, and extracts the article’s text, title, publication date, keywords, and summary. It then stores the results in a list of dictionaries and converts that list into a structured DataFrame. The script includes error handling for unparseable articles, inserting `NaN` values, and prints the total execution time, which took roughly 38.5 seconds for this run. The displayed table shows the resulting columns: date, keywords, link, summary, text, and title.
AI-generated description