PDFx

← Back to Open-Source PDF Software
PythonDepreciatedStaleTested

CLI tool and Python package that extracts metadata (creation date, creator, page count, producer) and detects references (other PDFs, URLs, DOIs, arXiv IDs) within a PDF, with parallel downloading of referenced PDFs, plain-text extraction, and broken-hyperlink detection; outputs plain text or JSON, and accepts local files or URLs.

Sample Results

Real output from running this tool against a sample PDF, as part of PDFog's open-source PDF tooling benchmark — pick one to see the commands and full output.