Skip to content

GoldenMatch

Data Pipelines & ETL

Zero‑config entity resolution that scales – dedupe and match messy records from CSVs up to 100 M rows without training data or tuning

Python Latest v3.8.0 · 5d ago Security brief →

Features

  • Exact, fuzzy (Jaro‑Winkler), probabilistic (Fellegi‑Sunter) and LLM‑based matching
  • Handles unstructured input – extract records from PDFs/images then dedupe
  • Runs in Python, edge‑safe TypeScript, SQL, Rust for Postgres/DuckDB, optional WebAssembly acceleration
  • Integrates as a resolution stage into GraphRAG pipelines (neo4j‑graphrag, LlamaIndex, Graphiti)
  • Scales to 100 M records in ~9 minutes on a Ray cluster with low memory footprint

Recent releases

View all 150 releases →
No immediate action
v3.8.0 New feature

FS columnar path + kernelized scorers

No immediate action
v3.7.0 Mixed

FS recall scaling + email/phone admission

Review required
v3.6.0 New feature

Default FS routing for fuzzy data

No immediate action
golden-suite-v0.3.1 Breaking risk

Version bumps

No immediate action
goldenflow-native-v0.27.0 Feature

arrow-native kernel + numeric-input ops

Weekly OSS security release digest.

The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.

No spam, unsubscribe anytime.

About

Stars
123
Forks
13
Languages
Python TypeScript Rust
Downloads/week
39 ↑5600%
NPM Maintainers
1 Single npm maintainer
Contributors
4
TypeScript
Types included ✓

Install & Platforms

Install via
pip npm

Community & Support

Alternative to

Splink

Beta — feedback welcome: [email protected]