Dual-Head Plate Detector (YOLOv8n)
A modified YOLOv8n that learns to spot Indonesian license plates without losing the 80 everyday object classes it already knew, by growing a second detection head instead of retraining the original one.
The straightforward way to make a detector recognize license plates is to fine-tune a pretrained model on plate photographs. It also breaks the model. Retraining the detection head on one new class overwrites what it knew about the other eighty, so you get a plate detector that can no longer find the motorcycle the plate is attached to, or the person riding it. For a gate system that needs all three at once, that trade is unacceptable, and running two entirely separate models to avoid it doubles the memory and the latency.
This takes the other route: leave the original detection head alone, still predicting COCO's 80 classes, and graft a second head beside it that learns only license_plate. Both heads read the same features from the shared network body, so the new class learns from scratch without disturbing anything the old head knows. A custom layer then merges the two heads' predictions into one output tensor, which matters more than it sounds: it means every downstream consumer, from non-maximum suppression to the inference API, keeps working as though nothing had changed. Getting this in required patching five files inside the Ultralytics framework, including the model config, the trainer and the module registry.
The first 22 layers are frozen during training, so the general visual features a COCO-pretrained network already has are kept and only the later layers adapt to plates. Augmentation leans on what a gate camera actually does to a plate: hue, saturation and brightness shifts for the time of day, rotation, shear, translation and scaling for the angle a motorcycle approaches at, and mosaic augmentation, which tiles several training images into one and is specifically useful for teaching small-object detection.
The dataset work behind it is the unglamorous half. Two plate datasets were merged with duplicates caught by hashing image contents rather than comparing filenames, then re-indexed so their annotation IDs didn't collide, then merged again with MS-COCO 2017 and converted from COCO's absolute pixel corners to YOLO's centre-and-size fractions. 73,752 deduplicated images came out the other end. Every one of those steps is a place where a merged detection dataset breaks silently rather than loudly.
The result: 0.811 mAP@50 on license_plate. That is higher than the same model scores on person (0.695) or motorcycle (0.486), classes it has been pretrained on since COCO, despite the plate class having only 470 validation images against person's 2,693. Plates are simply easier geometry, being rigid rectangles with a narrow aspect ratio and high text contrast, while a motorcycle changes shape completely depending on the angle you see it from. The headline 81-class average of 0.387 looks much worse and is essentially an artifact: dozens of COCO classes have a handful of examples or fewer in this validation split, and a class with one object scores 0 or 1 on noise alone, dragging the mean around.
Two numbers that look like flaws are choices. Recall (0.795) sits above precision (0.704) on purpose, because the errors are not symmetric. A false positive, a box that isn't really a plate, gets filtered downstream by aspect-ratio checks, by OCR failing to find text, and by a format regex. A false negative, a plate never detected at all, is gone permanently and no later stage can recover it. So the detection threshold is set deliberately low and the pipeline is built to clean up after it. Likewise the gap between mAP@50 (0.811) and mAP@50-95 (0.469) says boxes are reliably found but loosely drawn, which is fine here because the next stage expands and straightens the crop anyway. The confusion matrix confirms the shape of the failures: mistakes cluster in the background row rather than off-diagonal, meaning the model misses plates rather than mistaking them for something else.
Keeping COCO's 80 classes alive has direct operational value here rather than being a nice extra: the person and motorcycle boxes are what let the system pair a plate to a rider, estimate where a plate should be when it is not found, and reject a detection that turns up somewhere a plate could never be. One model, 81 classes, one forward pass, on a nano-sized backbone light enough for gate hardware. The dual-head pattern itself answers a common problem, "I need this pretrained detector plus one domain-specific class I care about", which is most applied computer vision: retail shelf monitoring, industrial inspection and agricultural sorting all hit the same wall.
The limits are honest ones. The 470-image validation split for plates is small, and all of it is Indonesian motorcycle plates, so car plates and other countries' formats are untested. Localization is loose enough that the rectification stage downstream isn't optional, it's load-bearing. And the frozen-backbone approach means the model inherits whatever COCO's features are bad at, which includes very small, very dark objects, exactly the condition a plate hits at dusk.
- Reached 0.811 mAP@50 on a newly added license-plate class, beating the same model's pretrained person (0.695) and motorcycle (0.486) classes despite 5x less validation data
- Added an 81st class without degrading the original 80, by growing a second detection head on the shared backbone and merging both outputs through a custom layer, so downstream code needed no changes
- Patched five files inside Ultralytics (model config, trainer, task registry and module definitions) to make the dual-head architecture trainable at all
- Deduplicated and merged two plate datasets plus MS-COCO 2017 into 73,752 images, catching duplicates by hashing image contents rather than filenames
- Tuned deliberately for recall over precision (0.795 vs 0.704) because downstream geometry, OCR and format checks can discard a false plate but can never recover a missed one
- Localizes loosely: the gap between 0.811 mAP@50 and 0.469 mAP@50-95 means boxes are found reliably but drawn imprecisely, making the downstream crop-expansion step mandatory rather than optional
- Validated on only 470 plate images, all Indonesian motorcycle plates, so car plates and other national formats are untested