Semantic Name Normalizer: Grounded in a Register
The second of four: the same AI cleaner, but checked against an official register of every Indonesian province, regency and district, which also generates the messy examples used to teach it what real messiness looks like.
The second version keeps the AI model but takes the guesswork away. Before anything is sent, an official register of every province, regency and district name in Indonesia, 80,534 entries, is pulled from Hugging Face and cleaned up: capitalization fixed, key columns merged into one full name per place.
That register does double duty as teaching material. For each clean name, a few deliberately messy variants are generated, lowercased, suffixed with "city" or "regency", stripped of spaces, title-cased only, so the model learns the shapes real messiness takes instead of inferring them from two hand-written examples.
It also handles a failure that is easy to miss. Sometimes the model returns fewer results than it was sent, and a batch that comes back with 9 answers for 10 rows silently shifts every row after it. Short batches are padded with "[MISSING]" and long ones trimmed, so what lands in Hive Metastore always lines up with what went in. That one check is the difference between a usable pipeline stage and a quietly corrupted table.
What it still cannot do is tell administrative types apart from context. "Jember City" and "Kabupaten Jember" are genuinely different places and it does not reliably know that. Noisy input can also make it over-cautious, rejecting a name that was valid, and a formatting pattern it has never seen still needs a rule written by hand. It runs at about $0.49 and 550,000 to 750,000 tokens per pass.
- Builds its own training examples from a real 80,534-entry knowledge base, instead of relying on hard-coded examples
- Catches a mismatched batch, flags missing rows as "[MISSING]" and trims any extras, so the output count always matches the input
- Barely needs adjusting when a new dataset gets added
- Tracks its own coverage ratio and normalization stats, so monitoring isn't just raw logs
- Still a black box: a prediction can't be traced back to why the model chose it
- Can't tell administrative types apart from context alone, "Jember City" and "Kabupaten Jember" read as the same thing to it
- Noisy input can make it too cautious, sometimes failing to recognize a name that's actually valid
- Any genuinely new formatting pattern still needs a manual rule written by hand
- Burns roughly 550,000 to 750,000 tokens per run
- Costs about $0.49 per run ($0.40 compute, $0.084 AI inference, $0.0026 storage), around Rp 8,060 to check 10,693 rows against an 80,534-entry knowledge base.