This exports content as XML, which matters when another system has to consume it. XML carries structure alongside the text, so downstream software can pick out the parts it needs rather than parsing prose.
It reflects the layout the PDF exposes: pages, blocks and text runs. It is not a semantic schema for your domain.
You will need a transformation step, usually XSLT or a script.
Run OCR PDF first, or the XML will describe images rather than text.
These tools handle one document at a time. When the job is thousands of files, Greenbooks handles document digitization as a managed service, including bulk scanning, OCR and metadata capture, with the output loaded into DocuVenta DMS. Talk to us about volume work.