NameetP/pdfmux
Developer ProductivitySelf‑healing PDF extractor that audits each page, re‑extracts failures and supports OCR, table, LLM fallback backends for clean RAG input
Features
- Per‑page confidence scoring and automatic self‑audit of extraction quality
- Routed extraction through multiple rule‑based backends (PyMuPDF, OpenDataLoader, RapidOCR, Docling, Surya, Marker) plus BYOK LLM fallback
- Zero‑config CLI and Python API with chunking, schema‑guided structured output, cost estimation and CI‑friendly strict mode
Recent releases
View all 19 releases →Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
About
Stars
75
Forks
12
Languages
Python
JavaScript
Shell
Downloads/week
7
↑550%
NPM Maintainers
1
Single npm maintainer
Contributors
3
Install & Platforms
Install via
pip
Similar tools
Alternative to
LlamaParse