Indonesian Face Recognition Model (IR-50 + ArcFace)
A face-recognition model trained from scratch on Indonesian faces, built because the usual pretrained options learned what a face looks like from datasets that barely contain any, and then tested honestly enough to find out it still lost to one of them.
Nearly every off-the-shelf face-recognition model is trained on datasets built in the West. CASIA-WebFace, LFW and VGGFace2 between them are overwhelmingly Caucasian, with East Asian faces a distant second and Southeast Asian faces barely represented at all. A model learns what dimensions of a face are worth paying attention to from the faces it is shown, so pointing one of those models at Indonesian riders means asking it to discriminate along axes that were tuned on a population these people aren't in. This is a well-documented fairness problem, and it is also a plain accuracy problem for anyone deploying biometrics in Indonesia.
So I trained one. The backbone is IR-50, a 50-layer residual network that compresses a face into 512 numbers, trained with ArcFace loss on 26,670 photographs of 68 Indonesian students. ArcFace is what makes those 512 numbers comparable: instead of only asking the network to classify each face correctly, it demands a fixed angular gap between every identity and every other one, so same-person and different-person pairs end up in genuinely separated regions rather than merely on the correct side of a boundary.
The training itself has two decisions worth naming. First, that angular margin is ramped from zero to full over the first ten epochs rather than applied from the start. A fresh network has no useful features yet, and demanding wide separation immediately destabilises it; you can see the cost of the ramp as a deliberate bump in the loss curve around epoch 6 to 10, after which loss collapses cleanly. Second, every training image is re-run through a face detector before training and only kept if it passes a confidence and size floor, because an earlier run had non-face crops in the training set quietly corrupting the embedding. Those cleaned crops are cached, so the filtering costs one pass, not one per epoch.
On held-out images it works: 0.9967 ROC-AUC, a 1.59% equal error rate, 99.5% rank-1 accuracy, measured across 81,913 same-person pairs against 5.2 million different-person ones. Performance is even across identities too, with the single worst-performing person still at 92.9% rank-1, so no one identity is collapsing and dragging an average along behind it. The checkpoint that shipped is epoch 30 of a possible 100, chosen on validation score rather than training loss, which kept the 40 subsequent epochs of overfitting out of the final model.
Then it met real CCTV, and the equal error rate went from 1.59% to 28.7%. Same model, same preprocessing, verified as correctly implemented by re-running the production pipeline against the clean test set and getting the clean numbers back. The difference is entirely the data: the test set is closed-set, drawn from the same pool as training, while the gate sees people the model has never encountered, through helmets, motion blur and hard sunlight, at whatever resolution survives being cropped out of a wide CCTV frame. Genuine pairs average 0.666 similarity and impostor pairs 0.284, so the model is still separating them on average, but the impostor tail reaches 0.706 and overlaps the genuine distribution. That overlap is the whole error rate.
The most uncomfortable finding is the comparison I ran anyway. Swapping this purpose-trained model back out for the general-purpose VGGFace2 model it was meant to improve on gave better live results, not worse: 15.4% equal error rate against 28.7%. The likely reason is dull and specific. VGGFace2 is trained on millions of faces across enormous variation in pose, lighting and image quality, whereas this model saw 68 identities under mild colour jitter, which is nothing like a gate camera at noon. It is a data-scale and augmentation problem, not an architecture one. The right fix is more identities and augmentation that actually imitates CCTV degradation, upscaling artefacts and contrast correction included, rather than a different backbone.
That result is in the thesis with a chart of its own, and I would rather have it there than not. Reporting only the 1.59% would have been technically true and practically misleading, and a face-recognition number that overstates itself is the kind of number that gets someone wrongly detained at a gate. The gap is also what shaped the system this model sits inside: because the face could not be trusted alone, Smart Gate was built so that the plate has to match before a face score is even consulted, and 40% of its successful matches ended up going through with no usable face at all. The model being worse than hoped is precisely why the system around it still works.
- Trained a 512-dimension face embedding on 26,670 images of 68 Indonesian identities, reaching 0.9967 ROC-AUC and a 1.59% equal error rate over 5.3 million verification pairs
- Ramped the ArcFace angular margin from 0 to 0.5 across the first 10 epochs to keep early training stable, then selected the checkpoint on validation score (epoch 30 of 100) rather than training loss, leaving 40 epochs of overfitting out of the shipped model
- Re-filtered every training image through a face detector first after finding non-face crops corrupting an earlier run, and cached the result so the cleanup cost one pass rather than one per epoch
- Held performance even across identities, with the weakest of 68 still at 92.9% rank-1, so the headline average isn't hiding a collapsed class
- Degraded to a 28.7% equal error rate on live CCTV against riders it had never seen, a closed-set to open-set gap of more than 18x
- Lost to the off-the-shelf VGGFace2 model it was meant to replace on that same live footage (28.7% vs 15.4% equal error rate), traced to training augmentation that never imitated CCTV conditions
- Published the unfavourable comparison in the thesis rather than reporting the 1.59% alone, and used it to justify demoting the face to supporting evidence in the system it feeds