Semantic Name Normalizer: Matching Cascade

AboutOctober 2025

03· Data Engineering, AI EngineeringInternship project, PT Telkom Indonesia · 3 of 4, recommended for production

The version recommended for production: three cheap matching stages run first and the AI model only sees what none of them could resolve, cutting token use to a sixth of the naive approach while letting every result be traced back to the stage that produced it.

ETLLLM
How it worksv3
Both the raw input names and the knowledge base get normalized first (capitalization fixes on words like "Kabupaten"/"Kelurahan", prefix/suffix cleanup, and for the knowledge base, merging its columns into one full name per place). Each name then runs through up to three matching stages in order: a canonical-list match, then a substring match, then fuzzy matching (Ratcliff-Obershelp, cutoff around 0.7), stopping at whichever stage first finds a hit. Only names that fail all three go to the AI model, along with a batch of the nearest candidates and a set of static examples to choose from.
Results side by sidev3
The same ten input names run through all three versions so far. This one finally gets "kota BandaAceh city" fully right ("Kota Banda Aceh," properly spaced), something neither earlier version managed. It also changes two results neither earlier version touched: "kota banyuwangi city" and "kota LUMAJANG city" both come back as "Kabupaten" instead of "Kota." And on "kabupaten YOGYAKARTA," it reverts to "Kota Yogyakarta", the same answer the "No Knowledge" version gave, rather than keeping the "Kabupaten Yogyakarta" the previous version got right.
Write-up

This is the version I recommended for production, and the reason is cost structure rather than accuracy. Every messy name goes through up to three matching stages, a canonical-list lookup, then a substring match, then fuzzy matching, and each stage only runs if the one before it came up empty. The model is the last resort, not the first thing the data hits.

Both the input names and the register are normalized the same way first: capitalization fixed on words like "Kabupaten" and "Kelurahan", stray prefixes and suffixes stripped, the register's separate columns merged into one full name per place. Only names that survive every stage without a hit reach the model, and they arrive with a batch of the nearest candidates attached, so it is choosing from a shortlist instead of inventing an answer.

That ordering pays twice. Token use drops to roughly 150,000 a run against 960,000 for the naive version, and the same batch finishes in about a minute where the semantic-search version takes 28. It also makes the output auditable: every row records which stage resolved it and why, which is what lets a data team debug a bad batch instead of shrugging at it. At $0.52 a run it costs about three cents more than the cheapest variant, and those three cents are what buy the traceability.

The fuzzy-matching threshold is a real balancing act. Too strict and it rejects names that were correct; too loose and false positives slip through. The examples handed to the model fallback are sampled dynamically rather than fixed, so that sampling carries its own bias into which candidates the model ever sees.

Things to underline
  • Uses AI only as a last resort, three cheaper matching stages run first, so the model gets called far less often
  • Averages only about 150,000 tokens per run, a fraction of what the AI-only versions burn
  • Every match can be traced back to which stage found it and why, not a black box
  • Barely needs adjusting when a new dataset gets added
  • The fuzzy-match threshold is a genuine trade-off: too strict rejects correct results, too loose lets false positives through
  • Dynamically-sampled examples fed to the AI fallback step introduce their own sampling bias
  • Recognizing a valid but unusual name still depends on how the fuzzy threshold happens to be tuned
  • Runs the full batch in about 1 minute 7 seconds
  • Costs about $0.52 per run ($0.40 compute, $0.1153 AI inference, $0.0074 storage), around Rp 8,658, to check 10,693 rows against an 80,534-entry knowledge base
Built with
PythonDatabricks & Hive MetastoreHugging Face