Import of corpus data from PDF
This utility enables Sparv to process PDF files by extracting their text using pypdfium2.
This utility enables Sparv to process PDF files by extracting their text using pypdfium2.
This importer is used with Sparv. Check out Sparv's quick start guide to get started!
To use this importer, run Sparv with the argument 'pdf_import:parse':
sparv run pdf_import:parse
For more info on how to use Sparv, check out the Sparv documentation.