RAÍZ: A Hybrid Architecture Based on Linguistic Knowledge and Deep Learning for the Machine Translation of Low-Resource Indigenous Languages

Authors

DOI:

https://doi.org/10.61467/2007.1558.2026.v17i5.1451

Keywords:

Indigenous Language Translation, Low-Resource NLP, Hybrid Neural Architecture, Multilingual Transformers, Linguistic Preprocessing, Community-Based Validation, Culturally Grounded AI, traducción de lenguas indígenas, PLN de bajos recursos, arquitectura neuronal híbrida, inteligencia artificial con fundamento cultural, Transformers multilingües

Abstract

Indigenous languages continue to experience significant technological marginalisation due to limited digital corpora, orthographic inconsistency, dialectal variation, morphological complexity, and insufficient representation in contemporary natural language processing systems. This study introduces RAÍZ, a hybrid machine translation architecture designed for low-resource Indigenous languages. The proposed framework combines Transformer-based multilingual modelling with linguistically informed components, including phonetic normalisation, morphological segmentation, semantic contextualisation, lexical retrieval mechanisms, and community-centred validation procedures involving native speakers. The architecture was evaluated using Spanish–Nahuatl and Spanish–Mazahua translation tasks through a mixed evaluation strategy integrating BLEU, chrF, and COMET metrics alongside qualitative assessment focused on clarity, naturalness, semantic fidelity, and cultural appropriateness. Experimental results indicate that RAÍZ achieved 28.6 BLEU, 56.8 chrF, and 74% COMET in the Spanish–Nahuatl task, while obtaining 27.4 BLEU, 54.9 chrF, and 71% COMET for Spanish–Mazahua translation. The findings suggest that hybrid and linguistically grounded approaches can improve translation robustness in low-resource settings while supporting culturally sensitive language technologies. The study further highlights the importance of combining neural architectures with community participation and explicit linguistic knowledge in the development of more inclusive machine translation systems for Indigenous languages.

 

Spanish-language metadata / Metadatos en español
Título en español:
RAÍZ: Una arquitectura híbrida basada en conocimiento lingüístico y aprendizaje profundo para la traducción automática de lenguas indígenas con pocos recursos

Resumen:
Las lenguas indígenas continúan experimentando una importante marginación tecnológica debido a la disponibilidad limitada de corpus digitales, la inconsistencia ortográfica, la variación dialectal, la complejidad morfológica y su insuficiente representación en los sistemas contemporáneos de procesamiento del lenguaje natural.

Este estudio presenta RAÍZ, una arquitectura híbrida de traducción automática diseñada para lenguas indígenas con pocos recursos. El marco propuesto combina modelado multilingüe basado en Transformers con componentes fundamentados en conocimiento lingüístico, entre ellos la normalización fonética, la segmentación morfológica, la contextualización semántica, mecanismos de recuperación léxica y procedimientos de validación centrados en la comunidad con participación de hablantes nativos.

La arquitectura se evaluó mediante tareas de traducción español–náhuatl y español–mazahua, utilizando una estrategia de evaluación mixta que integró las métricas BLEU, chrF y COMET con una evaluación cualitativa centrada en la claridad, naturalidad, fidelidad semántica y adecuación cultural.

Los resultados experimentales indican que RAÍZ alcanzó 28,6 BLEU, 56,8 chrF y 74 % COMET en la tarea español–náhuatl, mientras que para la traducción español–mazahua obtuvo 27,4 BLEU, 54,9 chrF y 71 % COMET.

Los hallazgos sugieren que los enfoques híbridos con fundamento lingüístico pueden mejorar la robustez de la traducción en contextos de bajos recursos y, al mismo tiempo, favorecer el desarrollo de tecnologías lingüísticas culturalmente sensibles. El estudio también destaca la importancia de combinar arquitecturas neuronales con la participación comunitaria y el conocimiento lingüístico explícito para desarrollar sistemas de traducción automática más inclusivos destinados a las lenguas indígenas.

Palabras Claves:
traducción de lenguas indígenas, PLN de bajos recursos, arquitectura neuronal híbrida, Transformers multilingües, preprocesamiento lingüístico, validación basada en la comunidad, inteligencia artificial con fundamento cultural


Smart citations:

SciteAI
Dimensions.
Open Alex.

References

Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with Python. O'Reilly Media.

De Gibert, O., Pugh, R., Marashian, A., Vazquez, R., Ebrahimi, A., Denisov, P., Rice, E., Gow-Smith, E., Prieto, J., Robles, M., Manrique, R., Moreno, O., Lino, A., Coto-Solano, R., Alvarez, A., Agüero-Torales, M., Ortega, J. E., Chiruzzo, L., Oncevay, A., . . . Mager, M. (2025). Findings of the AmericasNLP 2025 shared tasks on machine translation, creation of educational material, and translation metrics for Indigenous languages of the Americas. In Proceedings of the Fifth Workshop on NLP for Indigenous Languages of the Americas (AmericasNLP) (pp. 134–152). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.americasnlp-1.16Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1702.08608

Freitag, M., Rei, R., Mathur, N., Lo, C.-K., Stewart, C., Avramidis, E., Kocmi, T., Foster, G., Lavie, A., & Martins, A. F. T. (2022). Results of WMT22 metrics shared task: Stop using BLEU—Neural metrics are better and more robust. In Proceedings of the Seventh Conference on Machine Translation (WMT) (pp. 46–68). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.wmt-1.2

Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735

Jurafsky, D., & Martin, J. H. (2023). Speech and language processing (3rd ed. draft, January 7, 2023) [Draft manuscript]. https://web.stanford.edu/~jurafsky/slp3/old_jan23/ed3book_jan72023.pdf

Mager, M., Oncevay, A., Ebrahimi, A., Ortega, J., Rios, A., Fan, A., Gutierrez-Vasques, X., Chiruzzo, L., Giménez-Lugo, G., Ramos, R., Meza Ruiz, I. V., Coto-Solano, R., Palmer, A., Mager-Hois, E., Chaudhary, V., Neubig, G., Vu, N. T., & Kann, K. (2021). Findings of the AmericasNLP 2021 shared task on open machine translation for Indigenous languages of the Americas. In Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas (pp. 202–217). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.americasnlp-1.23

Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.

NLLB Team, Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Mejia Gonzalez, G., Hansanti, P., . . . Wang, J. (2022). No language left behind: Scaling human-centered machine translation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2207.04672

Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135

Popović, M. (2015). chrF: Character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation (pp. 392–395). Association for Computational Linguistics. https://doi.org/10.18653/v1/W15-3049

Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1–67.

Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 2685–2702). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.213

Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215. https://doi.org/10.1038/s42256-019-0048-x

Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1715–1725). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1162

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.

Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., & Raffel, C. (2022). ByT5: Towards a token-free future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics, 10, 291–306. https://doi.org/10.1162/tacl_a_00461

Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., & Raffel, C. (2021). mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 483–498). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.41

Downloads

Published

2026-09-06

How to Cite

Salas-Cabrera, G., Gallegos-Macías, A. A., & Trejo-Macotela, F. R. (2026). RAÍZ: A Hybrid Architecture Based on Linguistic Knowledge and Deep Learning for the Machine Translation of Low-Resource Indigenous Languages. International Journal of Combinatorial Optimization Problems and Informatics, 17(5), 168–182. https://doi.org/10.61467/2007.1558.2026.v17i5.1451

Issue

Section

Articles

Most read articles by the same author(s)

<< < 1 2