Skip to content

NameetP/pdfmux

v1.8.2 Maintenance

This release keeps dependencies and maintenance posture current for teams operating this tool.

✓ No known CVEs patched
Read the diff → Tool health → What is this tool? →

✓ No known CVEs patched in this version

Topics

ai-agent docling document-parsing llm mcp ocr
+8 more
opendataloader pdf pdf-extraction pdf-to-json pdf-to-markdown python self-healing structured-extraction

Summary

AI summary

Removed unverified benchmark claim; documentation now reflects a successful 433-document batch with zero silent failures.

Full changelog

Docs-only, no code changes. Removes an unverified "#2 on opendataloader-bench" claim from the README/PyPI description and replaces it with the real, git-dated 433-document customer batch: the naive early pipeline silently dropped 16 documents, 11 with no log line at all; rebuilt with the per-page audit, 433/433 processed with zero silent failures.

Also aligns the patent-pending method description with LICENSING.md, and seeds the expanded proposition — Certify Anything: audit any extractor's output for silently-dropped pages.

Full notes: CHANGELOG.md

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

Share this release

Track NameetP/pdfmux

Get notified when new releases ship.

Sign up free

About NameetP/pdfmux

PDF extraction router with built-in MCP server. Classifies each page (digital, scanned, tables) and routes to the best backend (PyMuPDF, Docling, OCR, or optional LLM fallback)

All releases →

Related context

Earlier breaking changes

  • v1.7.0 `--strict` flag is ON by default; `--min-confidence` defaults to 0.75.

Beta — feedback welcome: [email protected]