Fine-Tuning TrOCR for Automatic Number Plate Recognition: A YOLOv11-Based Comparison with Tesseract and EasyOCR
DOI:
https://doi.org/10.61467/2007.1558.2026.v17i4.1294Keywords:
Computer vision, Deep Learning, Transformer architecture, visión por computadora, aprendizaje profundo, arquitectura de transformadoresAbstract
TrOCR, a recently proposed object character recognition engine, is implemented along with the YOLOv11 object detector into an Automatic Number Plate Recognition (ANPR) pipeline to identify plates at an entrance of one of the Universidad Juárez Autónoma de Tabasco campus. First tests show moderate results of TrOCR, obtaining a performance of 0.4605 using the Character Error Rate (CER) metric. To improve this result, we perform a fine-tune of the base TrOCR engine retraining the model using a dataset of 651 custom license plate images. The retrained TrOCR reduced CER to 0.2576, demonstrating the effectiveness of retraining for improving recognition accuracy in challenging conditions. The approach shows that combining YOLOv11 with a fine-tuned OCR engine allows accurate recognition under varied conditions, highlighting its potential for robust, real-world deployment.
Spanish-language metadata / Metadatos en español
Título en español:
Ajuste fino de TrOCR para el reconocimiento automático de matrículas: una comparación basada en YOLOv11 con Tesseract y EasyOCR
Resumen:
Los sistemas de reconocimiento automático de matrículas (ANPR, por sus siglas en inglés) suelen constar de dos etapas: la localización de la matrícula en una imagen y el reconocimiento de sus caracteres mediante reconocimiento óptico de caracteres (OCR). Aunque el aprendizaje profundo ha mejorado significativamente la precisión de la detección —con modelos como YOLOv11, que ofrecen un rendimiento robusto y en tiempo real—, la etapa de OCR continúa siendo un problema abierto, especialmente en condiciones adversas como desenfoque por movimiento, baja iluminación u oclusión parcial. Este artículo evalúa y compara tres motores de OCR —Tesseract, EasyOCR y TrOCR— dentro de una arquitectura ANPR que utiliza YOLOv11 para la detección en un entorno real de un campus universitario. Una contribución principal es el ajuste fino de TrOCR mediante un conjunto de datos personalizado de 651 imágenes de matrículas, diseñado para representar patrones visuales y ruido específicos del dominio. El rendimiento se cuantifica mediante métricas de detección —precisión y exhaustividad— y mediante la tasa de error de caracteres (CER) para evaluar el reconocimiento. YOLOv11 mostró un alto rendimiento de detección, con una precisión de 0.938 y una exhaustividad de 0.995. En la comparación inicial de los motores de OCR, TrOCR superó a los demás, con una CER de 0.461, frente a 0.626 para EasyOCR y 1.027 para Tesseract. Después del ajuste fino, la CER de TrOCR mejoró aproximadamente un 44 %, hasta alcanzar un valor de 0.258, lo que confirma que el reentrenamiento de un modelo robusto puede mejorar la precisión del reconocimiento de caracteres en escenarios reales y complejos.
Palabras Claves:
visión por computadora; aprendizaje profundo; arquitectura de transformadores.
Smart citations:
https://scite.ai/reports/10.61467/2007.1558.2026.v17i4.1294
Dimensions.
Open Alex.
References
Al-Hasan, T. M., Bonnefille, V., & Bensaali, F. (2024). Enhanced YOLOv8-based system for automatic number plate recognition. Technologies, 12(9), Article 164. https://doi.org/10.3390/technologies12090164
Anwar, N., Khan, T., & Mollah, A. F. (2022). Text detection from scene and born images: How good is Tesseract? In A. K. S. Pundir, N. Yadav, H. Sharma, & S. Das (Eds.), Recent trends in communication and intelligent systems (pp. 115–122). Springer. https://doi.org/10.1007/978-981-19-1324-2_13
Baek, J., Kim, G., Lee, J., Park, S., Han, D., Yun, S., Oh, S. J., & Lee, H. (2019). What is wrong with scene text recognition model comparisons? Dataset and model analysis. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 4714–4722). IEEE. https://doi.org/10.1109/ICCV.2019.00481
Bakhtan, M. A. H., Abdullah, M., & Rahman, A. A. (2016). A review on license plate recognition system algorithms. In 2016 International Conference on Information and Communication Technology Management (ICICTM) (pp. 84–89). IEEE. https://doi.org/10.1109/ICICTM.2016.7890782
Dhyani, S., & Kumar, V. (2023). Real-time license plate detection and recognition system using YOLOv7x and EasyOCR. In 2023 Global Conference on Information Technologies and Communications (GCITC) (pp. 1–5). IEEE. https://doi.org/10.1109/GCITC60406.2023.10425814
Drobac, S., & Lindén, K. (2020). Optical character recognition with neural networks and post-correction with finite state methods. International Journal on Document Analysis and Recognition, 23, 279–295. https://doi.org/10.1007/s10032-020-00359-9
Google Cloud. (n.d.). Colab Enterprise documentation. Retrieved 31 July 2026, from https://docs.cloud.google.com/colab/docs
Jocher, G., & Qiu, J. (2024). Ultralytics YOLO11 (Version 11.0.0) [Computer software]. GitHub. https://github.com/ultralytics/ultralytics
K, T. D., James, J., Gopinath, D. P., & K, M. A. (2024). Advocating character error rate for multilingual ASR evaluation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2410.07400
Laroca, R., Zanlorensi, L. A., Gonçalves, G. R., Todt, E., Schwartz, W. R., & Menotti, D. (2021). An efficient and layout-independent automatic license plate recognition system based on the YOLO detector. IET Intelligent Transport Systems, 15(4), 483–503. https://doi.org/10.1049/itr2.12030
Lauar, F., & Laurent, V. (2024). Spanish TrOCR: Leveraging transfer learning for language adaptation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2407.06950
Li, M., Lv, T., Chen, J., Cui, L., Lu, Y., Florencio, D., Zhang, C., Li, Z., & Wei, F. (2023). TrOCR: Transformer-based optical character recognition with pre-trained models. Proceedings of the AAAI Conference on Artificial Intelligence, 37(11), 13094–13102. https://doi.org/10.1609/aaai.v37i11.26538
Menezes, D., Patel, S., Shaikh, A., & Parikh, N. (2025). Robust ANPR system using YOLOv8 and TrOCR with aspect-ratio-based splitting and post-processing. International Journal for Research in Applied Science & Engineering Technology, 13(9), 1911–1914. https://doi.org/10.22214/ijraset.2025.74265
Nguyen, D. D., Vo, T. S., Le, M. H., & Nguyen, M. S. (2025). An integration of segmentation technique on edge devices for license plate recognition. Vietnam Journal of Science, Technology and Engineering, 67(1), 3–13. https://doi.org/10.31276/VJSTE.2023.0099
Rather, I. H., Kumar, S., & Gandomi, A. H. (2024). Breaking the data barrier: A review of deep learning techniques for democratizing AI with small datasets. Artificial Intelligence Review, 57, Article 226. https://doi.org/10.1007/s10462-024-10859-3
Sarhan, A., Abdel-Rahem, R., Darwish, B., Abou-Attia, A., Sneed, A., Hatem, S., Badran, A., & Ramadan, M. (2024). Egyptian car plate recognition based on YOLOv8, Easy-OCR, and CNN. Journal of Electrical Systems and Information Technology, 11, Article 32. https://doi.org/10.1186/s43067-024-00156-y
Satya, B., Manongga, D., Hendry, & Aminuddin, A. (2025). Optimized YOLOv8 for automatic license plate recognition on resource-constrained devices. Engineering, Technology & Applied Science Research, 15(2), 21976–21981. https://doi.org/10.48084/etasr.9983
Shu, M. (2025). Utilizing transfer learning for deep learning based image classification. Journal of Intelligence Technology and Innovation, 3(1), 58–73. https://itip-submit.com/index.php/JITI/article/view/108
Sonnara, F., Chihaoui, H., & Filali, F. (2025). Efficient real-time license plate recognition using deep learning on edge devices. Journal of Real-Time Image Processing, 22, Article 159. https://doi.org/10.1007/s11554-025-01738-3
Sporici, D., Cușnir, E., & Boiangiu, C.-A. (2020). Improving the accuracy of Tesseract 4.0 OCR engine using convolution-based preprocessing. Symmetry, 12(5), Article 715. https://doi.org/10.3390/sym12050715
Spruck, A., Hawesch, M., Maier, A., Rieß, C., Seiler, J., & Kaup, A. (2021). 3D rendering framework for data augmentation in optical character recognition (Invited paper). In 2021 International Symposium on Signals, Circuits and Systems (ISSCS) (pp. 1–4). IEEE. https://doi.org/10.1109/ISSCS52333.2021.9497438
Ströbel, P. B., Clematide, S., Volk, M., & Hodel, T. (2022). Transformer-based HTR for historical documents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2203.11008
Tao, L., Hong, S., Lin, Y., Chen, Y., He, P., & Tie, Z. (2024). A real-time license plate detection and recognition model in unconstrained scenarios. Sensors, 24(9), Article 2791. https://doi.org/10.3390/s24092791
Tavares, R. A. (2024). Comparison of image preprocessing techniques for vehicle license plate recognition using OCR: Performance and accuracy evaluation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2410.13622
Vizcarra, C., Alhamed, S., Algosaibi, A., Alnaeem, M., Aldalbahi, A., Aljaafari, N., Sawalmeh, A., Nazzal, M., Khreishah, A., Alhumam, A., & Anan, M. (2024). Deep learning adversarial attacks and defenses on license plate recognition system. Cluster Computing, 27, 11627–11644. https://doi.org/10.1007/s10586-024-04513-4
Wang, C.-Y., Bochkovskiy, A., & Liao, H.-Y. M. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2207.026967:
Zherzdev, S., & Gruzdev, A. (2018). LPRNet: License plate recognition via deep neural networks [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1806.10447
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Combinatorial Optimization Problems and Informatics

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.