Smart Gate, Two-Factor Vehicle Access Control
A campus-gate system that checks a motorcycle at exit against whoever actually rode it in, reading the plate and recognizing the rider's face, so a borrowed or stolen bike is refused the moment a different person tries to ride it out.
Gated housing complexes, campuses and paid parking in Indonesia mostly still verify vehicles the way they did thirty years ago: a guard glances at a plate, maybe at a ticket, and lifts the barrier. That check has a specific hole in it. It confirms the vehicle, never the person on it. Ride a borrowed or stolen motorcycle back out through the same gate and a plate-only system, human or automated, waves it straight through. Motorcycle theft in Indonesian parking areas works almost entirely through that hole.
Smart Gate closes it by checking two things at once when a bike leaves, and comparing both against whatever was recorded coming in: the plate, and the rider's face. When both agree the gate opens and the visit closes. When the plate matches but the face doesn't, the exit is refused, logged as a security event, and the motorcycle stays marked as still inside, exactly as if it never left. Nothing about the vehicle alone is ever sufficient.
The plate side is a pipeline, not a model. A detector finds the plate, a rotated outline recovers the angle it sits at, the crop is flattened in perspective, and only then is it read. The reading itself is split into an Indonesian plate's three zones (region prefix, digits, letter suffix) and each zone is corrected against its own alphabet, so a character that has to be a digit can never come back as the letter O. A primary text engine goes first, and two backup engines are called in only when that first reading fails the format check. Individually those three engines return a valid reading 21%, 27% and 43% of the time. Voting zone by zone across all of them reaches 69%, better than the best engine on its own, and swapping which engine leads changed the final match count by exactly zero.
The face side is a model I trained myself rather than a pretrained one, and its results are the most useful thing in the whole project. On its own curated test set it is close to perfect: an equal error rate of 1.59% and a ROC-AUC of 0.9967. Pointed at real gate CCTV, with riders it had never seen, helmets, motion and harsh sun, that same model degrades to a 28.7% equal error rate. The gap between those two numbers is the entire lesson. A benchmark score measured on faces drawn from the same pool as the training data says almost nothing about a model at a real gate.
So the architecture was built around distrusting the face rather than around the face being good. Every accepted pair has to clear a strict plate gate first (both sides read, all four core digits identical) before any similarity score is even considered. Only then does the face decide which tier the match gets labelled: both agreeing, plate-only, or a partial plate rescued by a strong face. A face score can never let a pair through by itself; it can only corroborate, or contradict. The closest published comparison, TRI-GATE, fuses its three signals into one number and opens the gate when that number clears a threshold. That is simpler, but a single number can't tell you afterwards which signal actually let a vehicle out. Here every decision carries an explicit evidence label, which is what makes the log worth anything in an investigation.
Evaluated on 142 real CCTV captures covering 71 entry-exit pairs, the system matched 47 of them (66.2%) with zero format or digit violations, and the plate detector reached 0.811 mAP@50. 19 of those 47 matches went through on the plate alone, with the face absent or too weak to use, which is only possible because the face was never load-bearing. The remaining 24 pairs were refused on purpose, each with a stated reason. That is the priority the whole thing is tuned to: refusing an unconvincing pair is cheap, and releasing the wrong vehicle is not.
66.2% only means something next to a comparison. TRI-GATE reports 97%, but on a dedicated gate camera with the vehicle stopped and the plate filling a good part of the frame. This ran on passive CCTV, plates 60-120 pixels wide, faces routinely behind helmets, and identities never seen during training. Most of the gap is task difficulty, and the rest is a reporting choice: TRI-GATE publishes a single combined accuracy, while this reports false-accept and false-reject rates in full, including the unflattering 28.7%.
A second input path takes short video clips instead of single frames, pooling evidence across frames before deciding, and adds a rider-continuity check aimed squarely at the theft case. A staged test (same motorcycle, different rider at exit) was correctly refused: the face scored 0.793, below the 0.85 needed to confirm but above the floor that would mark it as a different person entirely, so the pair was pushed to manual review rather than released. Three filmed hand-off scenarios ran the same way, with every mismatched rider caught, 12 out of 12, and no wrong rider ever let out. A clip that can't be resolved always fails closed, into the manual queue that already exists today.
The honest limits are all load-bearing, not footnotes. Image resolution is the real ceiling: 12 of 142 captures never produced a readable plate at all, and past roughly 66% the next gains come from moving or upgrading the camera, not from better code. Every threshold was tuned on one camera over one simulated day and would need revalidating elsewhere. Matching is greedy and order-dependent, so an early entry can claim an exit that suited a later one better. And the face model's open-set gap is unresolved; an earlier VGGFace2 backbone actually scored better on live footage (15.4% versus 28.7% equal error rate), which is documented in the thesis rather than quietly dropped.
The parts that make it deployable are the unglamorous ones. It runs on passive CCTV a complex already owns rather than imported gate hardware, which is the cost barrier keeping this class of system out of Indonesian residential and campus parking in the first place. It is built to Indonesian plate format rather than adapted from a Western one. Every decision, accepted or refused, leaves an auditable reason a guard or an investigator can read back. And when it is not sure, it hands the case to the human process already running the gate today, so it can be switched on beside existing staff instead of replacing them on day one.
- Matched 47 of 71 entry-exit pairs from real CCTV (66.2%) with zero false accepts: no accepted pair violated plate format or digit agreement
- Built the decision around a mandatory plate gate so a face score can never open the gate alone, which is why 19 of 47 matches still succeeded with the face absent or unusable
- Reached 0.811 mAP@50 on the license-plate class, higher than the pretrained person (0.695) and motorcycle (0.486) classes in the same model
- Raised usable plate readings from 43% (best single OCR engine) to 69% by voting three engines zone by zone; swapping the lead engine changed the final match count by zero
- Labelled every accepted match with the evidence behind it rather than a single fused score, so an audit can tell which signal actually released a vehicle
- Caught every staged rider-mismatch across three filmed hand-off scenarios (12 of 12), including a same-plate different-rider theft test refused at 0.793 face similarity
- Face recognition fell from 1.59% equal error rate on its own test set to 28.7% on live CCTV, an open-set generalisation gap reported rather than hidden
- Only 73.2% of captures produced a complete plate reading; 12 of 142 produced none at all, a limit set by CCTV resolution rather than by the algorithm
- Every similarity threshold was tuned on one camera over one simulated day, so another site needs revalidation before the numbers carry over
- Pairs entries to exits greedily in arrival order, so an early entry can consume an exit that would have suited a later one better