
Business intelligence platform
Enterprise data platform • Analytics engineering

000
snyamson
Registry service • Data quality
Multiple partners registering the same households at intake, inflating reach figures and risking duplicate assistance.
Names were transliterated inconsistently, dates of birth were often approximate, and no shared identifier existed across partners. Exact matching found almost nothing; loose matching produced false positives that would have wrongly excluded real households.
A probabilistic matching service scores candidate pairs across name, date of birth, location and household composition, then routes anything in the uncertain band to a human reviewer rather than deciding automatically.
Exclusion is never automatic. The service flags; a caseworker decides.
Reach figures became defensible, and the review queue gave partners a shared, auditable record of why a household was or was not treated as a duplicate.