Semantic Name Normalizer: Unaided LLM

AboutOctober 2025

01· Data Engineering, AI EngineeringInternship project, PT Telkom Indonesia · 1 of 4

The control in a four-way study of cleaning messy Indonesian place names: an AI model rewrites each one with no reference list to check its answer against, which is how it ends up confidently wrong rather than unsure.

LLM
How it worksv1
Rows from the unclean_indonesian_places table in Hive Metastore processed in batches of 10, go into an AI request step (build a prompt from the batch plus a hard-coded example, parse the JSON response), producing cleaned names that become a new column in a Spark DataFrame, which merges with the original table and saves to Hive Metastore as a new table.
Sample resultsv1
Ten sample place names before and after cleaning. Most came out right, like "kota TEGAL city" becoming "Kota Tegal", but "kabupaten Yogyakarta" came back as "Kota Yogyakarta", the mix-up mentioned above.
Write-up

This is the control in a four-part experiment, not a version anyone should deploy. It cleans messy location names, things like "kota TEGAL city" or "kabupaten YOGYAKARTA", into a consistent form like "Kota Tegal", by sending each one to an AI model (Meta Llama 3.1-8B Instruct, through Hugging Face) in batches of ten and asking it to fix the spelling and formatting.

The model is given no list of correct place names. It gets two examples of what a clean name looks like, and everything after that is an educated guess from patterns it already carries. That sets the floor the other three versions are measured against.

The guessing shows. In a small test batch it got 9 of 10 right, and it turned "kabupaten Yogyakarta" into "Kota Yogyakarta", swapping one kind of administrative area for another, roughly how confusing a county with a city would read in English. Nothing in the design could have caught that, and nothing in the output marks it as uncertain. Being confidently wrong rather than unsure is the failure mode that makes unaided model cleaning dangerous inside a pipeline, because address normalization sits underneath logistics routing, government service delivery and most customer databases in Indonesia.

It is slow and expensive at scale too, since every name needs its own request: around 10,700 names took just over 5 minutes and about 960,000 tokens. It suits small, tidy lists where names barely vary, a set of provinces or postal codes, and not free-form text.

Things to underline
  • Cleaned about 10,700 place names in just over 5 minutes, using only two example corrections as a guide
  • Mixed up "kabupaten Yogyakarta" with "Kota Yogyakarta", an unwitting mistake caused by having no reference list to check itself against
  • Needs a separate AI request for every single name, so time and cost both climb directly with the size of the list
  • Only as fast as the model itself, so every batch has to wait on a full AI response, so there's no way to speed it up short of a faster model
  • Often confuses places with similar names, since it has no way to tell two similarly-spelled regions apart
  • Works entirely on assumption, with nothing to check its answers against, so it can confidently guess wrong and never signal it's unsure
  • Doesn't learn between batches, so the exact same input can come back with a different answer depending on the run
  • Burns about 960,000 tokens per run on average
  • Costs roughly $0.57 per run ($0.40 compute, $0.137 AI inference, $0.035 storage), around Rp 9,723 to clean all 10,693 rows
Built with
PythonDatabricks & Hive MetastoreHugging Face