2017-07-25: Swimming In A Ocean Of JavaScript. Using Nodejs, Electron And Chrome To Preserve The Web
✓ Published0🌍 Public
NN0taN3rd
Last edited Jul 23, 2017
Created on Jul 22, 2017
This example demonstrates a Node.js toolkit for web archiving, showing how to parse, stream, and generate WARC and CDXJ files to preserve web content. It uses the `cdxj` library to read CDXJ indexes asynchronously, the `node-warc` parser with event listeners for streaming WARC records, and a configuration object for the `squid` crawler that controls Chrome via its WebSocket API. The code highlights handling HTTP/2 limitations in archives and supports features like page-same-domain crawling and auto-scrolling.
AI-generated description