Semantic Name Normalizer: Vector Search & Rerank
The fourth of four: each name is embedded and matched by meaning through FAISS, reranked for precision, and escalated to an AI model only when nothing scores confidently enough. The best explainability of the four, and five times slower than the crudest.
The fourth version replaces both guesswork and string similarity with real semantic search. Every name in the register is turned into a vector embedding and indexed with FAISS, so a messy input can be matched by what it means rather than by how closely its characters happen to line up.
To clean a name it is embedded the same way, FAISS returns the closest candidates by cosine similarity, and a cross-encoder reranks them for precision. If the top result clears a 0.75 confidence score that is the answer; if it does not, the name goes to the model as a fallback, the same safety net the cascade used, triggered by a similarity score instead of a string-distance cutoff.
The grounding buys accuracy and reproducibility. The same FAISS index is reused every run, so the same input reliably gives the same output, the confidence threshold is a genuine tunable lever, it runs almost entirely on CPU, and every decision can be explained by which candidates ranked highest and by how much.
It is also the slowest and, per run, the most expensive: 28 minutes for the batch the naive version cleaned in 5, at about $0.60, rebuilding its embeddings from scratch each time. Sophistication turning out not to be the right choice is the finding that closed the study.
One failure survived all four versions. "Kabupaten Yogyakarta" still resolves to "Kota Yogyakarta" here, exactly as it did with no knowledge base at all, which locates the problem in the reference data rather than in the matching technique. That is the kind of conclusion that saves a team a quarter spent on better retrieval when the fix is a data correction.
- Matches names by real semantic meaning (embeddings + FAISS), not just spelling or string similarity
- Same input reliably produces the same output every run, since the FAISS index doesn't change between runs
- Every match can be explained by which candidates ranked highest and by how much
- Ready for a much bigger dataset without needing structural changes
- By far the slowest version yet, 28 minutes for the same batch the simplest version cleaned in about 5
- Rebuilds embeddings from scratch on every run, so cost and compute load climb further at scale
- The confidence threshold is a genuine trade-off: too strict rejects correct matches, too loose biases the ranking toward the wrong one
- Still gets "kabupaten Yogyakarta" wrong, the same mistake the simplest version made, suggesting the knowledge base itself is ambiguous there, not just the matching method
- Burns roughly 75,000 tokens per run, the lowest of any version, since the AI model is only a rare fallback
- Costs about $0.60 per run ($0.40 compute, $0.01237 AI inference, $0.189 storage), around Rp 9,962, to check 10,693 rows against an 80,534-entry knowledge base
- Still needs response caching and smarter context windowing to be practical at bigger scale, and the current double retrieval pass could likely be trimmed to one