GoldenMatch
Data Pipelines & ETLZero‑config entity resolution that scales – dedupe and match messy records from CSVs up to 100 M rows without training data or tuning
Features
- Exact, fuzzy (Jaro‑Winkler), probabilistic (Fellegi‑Sunter) and LLM‑based matching
- Handles unstructured input – extract records from PDFs/images then dedupe
- Runs in Python, edge‑safe TypeScript, SQL, Rust for Postgres/DuckDB, optional WebAssembly acceleration
- Integrates as a resolution stage into GraphRAG pipelines (neo4j‑graphrag, LlamaIndex, Graphiti)
- Scales to 100 M records in ~9 minutes on a Ray cluster with low memory footprint
Recent releases
View all 150 releases →Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
About
Stars
123
Forks
13
Languages
Python
TypeScript
Rust
Downloads/week
39
↑5600%
NPM Maintainers
1
Single npm maintainer
Contributors
4
TypeScript
Types included ✓
Install & Platforms
Install via
pip
npm
Community & Support
Similar tools
Alternative to
Splink