XQuery for Partitioning the Blue Mountain Project
✓ Published0🌍 Public
CCliffordAnderson
Last edited Apr 29, 2017
Created on Apr 28, 2017
This example demonstrates a pipeline for cleaning, partitioning, and exporting articles from the Blue Mountain Project’s TEI transcriptions into language-specific databases. It shows how XQuery 3.1 scripts transform raw periodical XML by removing facsimile and hyphenation artifacts, splitting documents into individual articles, and grouping them by `mainLang` attributes into new databases. The workflow then exports French-language texts as plain .txt files, and R code using the `tm` and `SnowballC` libraries performs text mining, including term-document matrices and stemming.
AI-generated description