Skip to content

NameetP/pdfmux

Developer Productivity

Self‑healing PDF extractor that audits each page, re‑extracts failures and supports OCR, table, LLM fallback backends for clean RAG input

Python Latest v1.8.5 · 1h ago Security brief →

Features

  • Per‑page confidence scoring and automatic self‑audit of extraction quality
  • Routed extraction through multiple rule‑based backends (PyMuPDF, OpenDataLoader, RapidOCR, Docling, Surya, Marker) plus BYOK LLM fallback
  • Zero‑config CLI and Python API with chunking, schema‑guided structured output, cost estimation and CI‑friendly strict mode

Recent releases

View all 19 releases →
No immediate action
v1.8.1 New feature

Verification audit tool

No immediate action
v1.8.4 Bug fix

Version string fix

No immediate action
v1.8.3 Bug fix

Dead link fix + timeout termination

Config change
v1.7.0 Breaking risk
Breaking upgrade

--strict on by default

No immediate action
v1.6.4 Breaking risk

audit command + upsell line

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

About

Stars
75
Forks
12
Languages
Python JavaScript Shell
Downloads/week
7 ↑550%
NPM Maintainers
1 Single npm maintainer
Contributors
3

Install & Platforms

Install via
pip

Alternative to

LlamaParse

Beta — feedback welcome: [email protected]