Semantic Name Normalizer: Grounded in a Register

AboutOctober 2025

02· Data Engineering, AI EngineeringInternship project, PT Telkom Indonesia · 2 of 4

The second of four: the same AI cleaner, but checked against an official register of every Indonesian province, regency and district, which also generates the messy examples used to teach it what real messiness looks like.

ETLLLMRAG
How it worksv2
A real knowledge base of place names is downloaded from Hugging Face and normalized, then used to generate realistic noisy-to-clean example pairs (lowercase, extra suffixes, missing spaces, mixed case) that get bundled into the prompt alongside a batch of 10 unclean rows. If the model returns fewer results than it was sent, the missing ones are flagged as "[MISSING]"; if it returns too many, the extras are trimmed, either way, the count going into Hive Metastore always matches the batch that went in.
Results side by sidev2
The same ten input names run through two versions: the earlier "No Knowledge" version's output in the middle column, this version's in the right column. The difference shows up clearest on "kabupaten YOGYAKARTA", No Knowledge reclassifies it as "Kota Yogyakarta," while this version correctly keeps it "Kabupaten Yogyakarta." Neither gets "kota BandaAceh city" fully right, though: one leaves it as "Kota BandaAceh," the other as "Kota Bandaaceh".
Write-up

The second version keeps the AI model but takes the guesswork away. Before anything is sent, an official register of every province, regency and district name in Indonesia, 80,534 entries, is pulled from Hugging Face and cleaned up: capitalization fixed, key columns merged into one full name per place.

That register does double duty as teaching material. For each clean name, a few deliberately messy variants are generated, lowercased, suffixed with "city" or "regency", stripped of spaces, title-cased only, so the model learns the shapes real messiness takes instead of inferring them from two hand-written examples.

It also handles a failure that is easy to miss. Sometimes the model returns fewer results than it was sent, and a batch that comes back with 9 answers for 10 rows silently shifts every row after it. Short batches are padded with "[MISSING]" and long ones trimmed, so what lands in Hive Metastore always lines up with what went in. That one check is the difference between a usable pipeline stage and a quietly corrupted table.

What it still cannot do is tell administrative types apart from context. "Jember City" and "Kabupaten Jember" are genuinely different places and it does not reliably know that. Noisy input can also make it over-cautious, rejecting a name that was valid, and a formatting pattern it has never seen still needs a rule written by hand. It runs at about $0.49 and 550,000 to 750,000 tokens per pass.

Things to underline
  • Builds its own training examples from a real 80,534-entry knowledge base, instead of relying on hard-coded examples
  • Catches a mismatched batch, flags missing rows as "[MISSING]" and trims any extras, so the output count always matches the input
  • Barely needs adjusting when a new dataset gets added
  • Tracks its own coverage ratio and normalization stats, so monitoring isn't just raw logs
  • Still a black box: a prediction can't be traced back to why the model chose it
  • Can't tell administrative types apart from context alone, "Jember City" and "Kabupaten Jember" read as the same thing to it
  • Noisy input can make it too cautious, sometimes failing to recognize a name that's actually valid
  • Any genuinely new formatting pattern still needs a manual rule written by hand
  • Burns roughly 550,000 to 750,000 tokens per run
  • Costs about $0.49 per run ($0.40 compute, $0.084 AI inference, $0.0026 storage), around Rp 8,060 to check 10,693 rows against an 80,534-entry knowledge base.
Built with
PythonDatabricks & Hive MetastoreHugging Face