Gist d003ece0b6ca171be8e095e0bc90af7e
This notebook example demonstrates how to programmatically scrape, parse, and structure a corpus of New York Times articles into a pandas DataFrame. The code loads a pre-saved pickle file of 50 article links, iterates through each URL using the `newspaper` library's `Article` class, and extracts the article’s text, title, publication date, keywords, and summary. It then stores the results in a list of dictionaries and converts that list into a structured DataFrame. The script includes error handling for unparseable articles, inserting `NaN` values, and prints the total execution time, which took roughly 38.5 seconds for this run. The displayed table shows the resulting columns: date, keywords, link, summary, text, and title.