Gist 2559ad6fd17831c9259a8af6f38ec027
✓ Published0🌍 Public
NN0taN3rd
Last edited Jun 14, 2018
Created on Jun 14, 2018
This example filters a Common Crawl index file to identify archived web pages. It reads a CDXJ file using the `read_cdxj.py` script, which parses each line with the `CDXObject` class from `pywb.warcserver.index.cdxobject`. For each record, it checks whether the MIME type contains "html" and the HTTP status code equals 200, then prints the corresponding URL. The visualization shows which URLs meet these criteria, highlighting successful HTML captures from the crawl data.
AI-generated description