Dual-Head Plate Detector (YOLOv8n)

About2026

11· Machine Learning, AI EngineeringUndergraduate thesis · the plate half of Smart Gate

A modified YOLOv8n that learns to spot Indonesian license plates without losing the 80 everyday object classes it already knew, by growing a second detection head instead of retraining the original one.

computer-visionobject-detection
Building the training setFig. 3.2
Two separate plate datasets are merged first, with duplicate photographs caught by hashing the image contents rather than trusting filenames, then re-indexed so their annotation IDs don't collide. The result is merged again with MS-COCO 2017, with license_plate appended as class 80. Output: 73,752 deduplicated images.
Normalizing the annotationsFig. 3.3
COCO stores a box as absolute pixel corners; YOLO wants a centre point and size expressed as fractions of the image. Each image gets a stable unique ID, its annotations are re-indexed and converted, and the result is one plain text line per object. Unglamorous, and the single most common place a merged detection dataset silently breaks.
Two heads on one backboneFig. 3.6
The original detection head keeps predicting COCO's 80 classes and is left alone. A second head is added beside it, fed the same features, and learns license_plate by itself. A custom ConcatHead layer merges both sets of predictions into one output, so everything downstream (non-maximum suppression, the inference API) carries on working unchanged. Implementing this meant patching five files inside Ultralytics.
What the merged dataset looks likeFig. 4.4
Box positions cluster at the image centre, the COCO habit of photographing the subject in the middle, while box widths and heights pile up hard against zero. Most objects, plates very much included, occupy a tiny fraction of the frame. This is the distribution that makes small-object detection the actual problem here.
Training curvesFig. 4.5
All three loss components fall smoothly on both training and validation data, with no divergence between them, and every detection metric plateaus before epoch 20. Early stopping cut training at epoch 33 of a possible 100. Nothing dramatic, which is the point: the run is clean enough that the results can be read at face value.
Precision-recall across all 81 classesFig. 4.6
Each grey line is one class; the blue line is the 81-class average at 0.387 mAP@50. That average looks poor until you notice the spread, and that it is dragged down by dozens of COCO classes with almost no examples in this validation split (the toaster class has exactly one). The three classes this system actually uses all sit well above it.
The errors are misses, not mix-upsFig. 4.7
A strong diagonal, including for the newly added license_plate class, and the errors concentrated in the bottom background row rather than scattered off-diagonal. That means the model's mistakes are almost entirely 'didn't see it' rather than 'thought it was something else'. A missed plate is unrecoverable downstream; a mislabelled one would be worse.
Working and failing, side by sideFig. 4.8
Left, everything the pipeline needs: rider, motorcycle and plate all boxed. Right, the same camera on a different rider, where the plate is detected but the motorcycle is not, and the plate sits deep in shadow behind shopping bags. The fallback logic in Smart Gate exists for exactly the right-hand case.
Write-up

The straightforward way to make a detector recognize license plates is to fine-tune a pretrained model on plate photographs. It also breaks the model. Retraining the detection head on one new class overwrites what it knew about the other eighty, so you get a plate detector that can no longer find the motorcycle the plate is attached to, or the person riding it. For a gate system that needs all three at once, that trade is unacceptable, and running two entirely separate models to avoid it doubles the memory and the latency.

This takes the other route: leave the original detection head alone, still predicting COCO's 80 classes, and graft a second head beside it that learns only license_plate. Both heads read the same features from the shared network body, so the new class learns from scratch without disturbing anything the old head knows. A custom layer then merges the two heads' predictions into one output tensor, which matters more than it sounds: it means every downstream consumer, from non-maximum suppression to the inference API, keeps working as though nothing had changed. Getting this in required patching five files inside the Ultralytics framework, including the model config, the trainer and the module registry.

The first 22 layers are frozen during training, so the general visual features a COCO-pretrained network already has are kept and only the later layers adapt to plates. Augmentation leans on what a gate camera actually does to a plate: hue, saturation and brightness shifts for the time of day, rotation, shear, translation and scaling for the angle a motorcycle approaches at, and mosaic augmentation, which tiles several training images into one and is specifically useful for teaching small-object detection.

The dataset work behind it is the unglamorous half. Two plate datasets were merged with duplicates caught by hashing image contents rather than comparing filenames, then re-indexed so their annotation IDs didn't collide, then merged again with MS-COCO 2017 and converted from COCO's absolute pixel corners to YOLO's centre-and-size fractions. 73,752 deduplicated images came out the other end. Every one of those steps is a place where a merged detection dataset breaks silently rather than loudly.

The result: 0.811 mAP@50 on license_plate. That is higher than the same model scores on person (0.695) or motorcycle (0.486), classes it has been pretrained on since COCO, despite the plate class having only 470 validation images against person's 2,693. Plates are simply easier geometry, being rigid rectangles with a narrow aspect ratio and high text contrast, while a motorcycle changes shape completely depending on the angle you see it from. The headline 81-class average of 0.387 looks much worse and is essentially an artifact: dozens of COCO classes have a handful of examples or fewer in this validation split, and a class with one object scores 0 or 1 on noise alone, dragging the mean around.

Two numbers that look like flaws are choices. Recall (0.795) sits above precision (0.704) on purpose, because the errors are not symmetric. A false positive, a box that isn't really a plate, gets filtered downstream by aspect-ratio checks, by OCR failing to find text, and by a format regex. A false negative, a plate never detected at all, is gone permanently and no later stage can recover it. So the detection threshold is set deliberately low and the pipeline is built to clean up after it. Likewise the gap between mAP@50 (0.811) and mAP@50-95 (0.469) says boxes are reliably found but loosely drawn, which is fine here because the next stage expands and straightens the crop anyway. The confusion matrix confirms the shape of the failures: mistakes cluster in the background row rather than off-diagonal, meaning the model misses plates rather than mistaking them for something else.

Keeping COCO's 80 classes alive has direct operational value here rather than being a nice extra: the person and motorcycle boxes are what let the system pair a plate to a rider, estimate where a plate should be when it is not found, and reject a detection that turns up somewhere a plate could never be. One model, 81 classes, one forward pass, on a nano-sized backbone light enough for gate hardware. The dual-head pattern itself answers a common problem, "I need this pretrained detector plus one domain-specific class I care about", which is most applied computer vision: retail shelf monitoring, industrial inspection and agricultural sorting all hit the same wall.

The limits are honest ones. The 470-image validation split for plates is small, and all of it is Indonesian motorcycle plates, so car plates and other countries' formats are untested. Localization is loose enough that the rectification stage downstream isn't optional, it's load-bearing. And the frozen-backbone approach means the model inherits whatever COCO's features are bad at, which includes very small, very dark objects, exactly the condition a plate hits at dusk.

Things to underline
  • Reached 0.811 mAP@50 on a newly added license-plate class, beating the same model's pretrained person (0.695) and motorcycle (0.486) classes despite 5x less validation data
  • Added an 81st class without degrading the original 80, by growing a second detection head on the shared backbone and merging both outputs through a custom layer, so downstream code needed no changes
  • Patched five files inside Ultralytics (model config, trainer, task registry and module definitions) to make the dual-head architecture trainable at all
  • Deduplicated and merged two plate datasets plus MS-COCO 2017 into 73,752 images, catching duplicates by hashing image contents rather than filenames
  • Tuned deliberately for recall over precision (0.795 vs 0.704) because downstream geometry, OCR and format checks can discard a false plate but can never recover a missed one
  • Localizes loosely: the gap between 0.811 mAP@50 and 0.469 mAP@50-95 means boxes are found reliably but drawn imprecisely, making the downstream crop-expansion step mandatory rather than optional
  • Validated on only 470 plate images, all Indonesian motorcycle plates, so car plates and other national formats are untested
Built with
PyTorchUltralytics YOLOv8MS-COCO 2017Roboflow