Import of corpus data from XML
This utility imports XML documents into Sparv by parsing the XML tree, extracting textual content, handling elements and attributes, separating body text from header metadata, optionally removing namespaces, and creating annotations for both corpus text and structured metadata. The import process can be configured to rename or skip specific elements and attributes, and to include or exclude header metadata.
