Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Geometry-Constrained Multi-Frame Character Association for License Plate Recognition on Moving Cameras.

Created on 01 Oct 2026

Authors

Ufuk Asil, İlker Yoncacı

Published in

Sensors (Basel, Switzerland). Volume 26. Issue 18. Sep 08, 2026. Epub Sep 08, 2026.

Abstract

Multi-frame fusion is standard for converting frame-by-frame license plate character detections into stable, reliable readings. Character Time-Series Matching (CTM), a leading approach, associates characters across frames using the Hungarian algorithm with a fixed Euclidean distance threshold and a translation-only motion model, reporting 96.7% accuracy on the UFPR-ALPR dataset. In this work, we demonstrate that this high performance is protocol-dependent: when ground-truth static plate crops and pre-segmented tracks are used, CTM performs strongly. However, in real-world scenarios involving moving cameras (such as drone-mounted cameras, helmet-mounted cameras, and mobile platforms) where inter-frame geometry changes dynamically, baseline multi-frame association frameworks that combine fixed spatial gates with unconstrained translation propagation fail. In these cases, temporal fusion provides no benefit and degrades plate recognition performance below the single-frame baseline. Indeed, under its own Intersection over Union (IoU) tracker, this literature method correctly reads only 15.1% of plates in traffic videos recorded with a real moving camera (86 human-verified tracks). To address this vulnerability, we propose Geo-CTM (Geometry-Constrained CTM), an association pipeline integrating height-scaled adaptive matching gates, inter-frame similarity estimation via Random Sample Consensus (RANSAC), transform-guided character coasting, and co-occurrence-constrained duplicate track elimination. Systematic motion-model ablation demonstrates that while the complete association pipeline provides the primary foundation for robustness (raising mean accuracy from 85.40% to over 91.5%), estimating a similarity transform (91.82%) delivers the most physically grounded and identifiable representation on planar plates without estimation degeneration. While our method performs comparably to CTM on ideal data when using the same detector and detections, it minimizes performance loss under geometric distortion conditions where CTM is inadequate. For instance, a statistically significant improvement is achieved under a 0 → 60° perspective change; in real traffic videos, with the tracker held fixed so that the fusion layer is the only variable, performance rises from 15.1% to 26.7% under the IoU tracker of the original system and from 16.3% to 29.1% under ByteTrack (+11.6 and +12.8 points; exact McNemar p=0.021 and p=0.013), whereas changing the tracker alone while holding the fusion layer fixed moves accuracy by only 1-2 points and is not statistically significant. Finally, our error taxonomy analysis demonstrates that on the undistorted benchmark all residual errors correspond to zero-evidence cases beyond the reach of decision-level fusion, while under dynamic perspective distortion errors are dominated by association misalignment, highlighting the specific development areas that future performance improvements must target.

PMID:
42817257
Bibliographic data and abstract were imported from PubMed on 01 Oct 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement