Dynamic Load Balancing and Fault Tolerance for Machine Learning Deployments Using a Process Control Table
DOI:
https://doi.org/10.61467/2007.1558.2026.v17i4.1296Keywords:
MLOPS, Load balancing, Model Monitoring, Request optimization, Process Management, balanceo de carga, monitorización de modelos, optimización de solicitudesAbstract
Deploying machine learning models in production environments presents significant challenges when systems must process high request volumes while maintaining stable real-time performance. Existing solutions often address request distribution, performance monitoring, and fault management separately, which may contribute to bottlenecks, increased latency, and service saturation. This study presents the design of a Process Control Table that acts as a central coordinator for orchestrating the real-time execution of machine learning models. The proposed mechanism dynamically routes incoming tasks to the least occupied worker available at the time of assignment. Prototype testing indicates that the Process Control Table reduces latency and maintains more stable response times under high demand than the traditional schemes evaluated in the study, which tended to become saturated. The mechanism also promotes equitable workload distribution and incorporates automatic failover capabilities to improve deployment reliability. These findings indicate that centralised process coordination can support dynamic load balancing, request optimisation, and fault tolerance in machine learning deployments.
Spanish-language metadata / Metadatos en español
Título en español:
Balanceo dinámico de carga y tolerancia a fallos en despliegues de aprendizaje automático mediante una tabla de control de procesos
Resumen:
El despliegue de modelos de aprendizaje automático en entornos de producción presenta desafíos significativos cuando los sistemas deben procesar un gran volumen de solicitudes y, al mismo tiempo, mantener un rendimiento estable en tiempo real. Las soluciones existentes suelen abordar por separado la distribución de solicitudes, la monitorización del rendimiento y la gestión de fallos, lo que puede contribuir a la aparición de cuellos de botella, al aumento de la latencia y a la saturación del servicio. Este estudio presenta el diseño de una tabla de control de procesos que actúa como coordinador central para orquestar la ejecución en tiempo real de modelos de aprendizaje automático. El mecanismo propuesto dirige dinámicamente las tareas entrantes al trabajador disponible con menor carga en el momento de la asignación. Las pruebas realizadas con un prototipo indican que la tabla de control de procesos reduce la latencia y mantiene tiempos de respuesta más estables bajo una alta demanda que los esquemas tradicionales evaluados en el estudio, los cuales tendieron a saturarse. El mecanismo también favorece una distribución equitativa de la carga de trabajo e incorpora capacidades de conmutación automática por error para mejorar la fiabilidad de los despliegues. Estos hallazgos indican que la coordinación centralizada de procesos puede respaldar el balanceo dinámico de carga, la optimización de solicitudes y la tolerancia a fallos en los despliegues de aprendizaje automático.
Palabras Claves:
MLOps; balanceo de carga; monitorización de modelos; optimización de solicitudes; gestión de procesos.
Smart citations:
https://scite.ai/reports/10.61467/2007.1558.2026.v17i4.1296
Dimensions.
Open Alex.
References
Aliev, K., & Antonelli, D. (2021). Proposal of a monitoring system for collaborative robots to predict outages and to assess reliability factors exploiting machine learning. Applied Sciences, 11(4), Article 1621. https://doi.org/10.3390/app11041621
Dal Pozzolo, A., Caelen, O., Johnson, R. A., & Bontempi, G. (2015). Calibrating probability with undersampling for unbalanced classification. In 2015 IEEE Symposium Series on Computational Intelligence (SSCI) (pp. 159–166). IEEE. https://doi.org/10.1109/SSCI.2015.33
de la Rúa Martínez, J. (2020). Scalable architecture for automating machine learning model monitoring [Master’s thesis, KTH Royal Institute of Technology]. DiVA. https://urn.kb.se/resolve?urn=urn:nbn:se:kth:diva-280345
Elgamal, Z. S., El Fangary, L., & Fahmy, H. (2025). The impact of using MLOps and DevOps on container-based applications: A survey. Informatics Bulletin, 7(1), 51–63.
Ibadov, N., Akgün, F. M., Üncü, İ. S., Davraz, M., & Koru, M. (2025). Real-time service life estimation of vacuum insulated panels via embedded sensing and machine learning models. Buildings, 15(16), Article 2879. https://doi.org/10.3390/buildings15162879
Jin, Y., & Yang, Z. (2025). Scalability optimization in cloud-based AI inference services: Strategies for real-time load balancing and automated scaling [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2504.15296
Kamila, N. K., Frnda, J., Pani, S. K., Das, R., Islam, S. M. N., Bharti, P. K., & Muduli, K. (2022). Machine learning model design for high performance cloud computing & load balancing resiliency: An innovative approach. Journal of King Saud University–Computer and Information Sciences, 34(10), 9991–10009. https://doi.org/10.1016/j.jksuci.2022.10.001
Low, W. K., Ramasamy, R. K., & Rajendran, V. (2024). Adaptive load balancing strategies in service composition for improved system performance. Journal of Infrastructure, Policy and Development, 8(13), Article 8967. https://doi.org/10.24294/jipd8967
Maddireddy, K., Methukula, S. K., Sridhar, C., & Vaidhyanathan, K. (2025). LoCoML: A framework for real-world ML inference pipelines [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2501.14165
Mahdizadeh, M., Montazerolghaem, A., & Jamshidi, K. (2025). Task scheduling and load balancing in SDN-based cloud computing: A review of relevant research. Journal of Engineering Research, 13(4), 3132–3146. https://doi.org/10.1016/j.jer.2024.11.002
McClure, S., Cohen, E., Shpiner, A., Silberstein, M., Ratnasamy, S., Shenker, S., & Keslassy, I. (2025). Load balancing for AI training workloads [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.21372
Moens, P., Bracke, V., Soete, C., Vanden Hautte, S., Nieves Avendano, D., Ooijevaar, T., Devos, S., Volckaert, B., & Van Hoecke, S. (2020). Scalable fleet monitoring and visualization for smart machine maintenance and industrial IoT applications. Sensors, 20(15), Article 4308. https://doi.org/10.3390/s20154308
Mora, S. (2024). Intrusion detection using machine learning with real-time dashboard [Master’s thesis, National College of Ireland]. NORMA@NCI. https://norma.ncirl.ie/8752/
Puranik, S., & Mahmood, Y. (2025). Real-time machine performance visualisation & monitoring in smart manufacturing: A dual-tool approach [Master’s thesis, Mälardalen University]. DiVA. https://urn.kb.se/resolve?urn=urn:nbn:se:mdh:diva-71784
Sanjalawe, Y., Fraihat, S., Al-E’mari, S., Abualhaj, M., Makhadmeh, S., & Alzubi, E. (2025). Smart load balancing in cloud computing: Integrating feature selection with advanced deep learning models. PLOS ONE, 20(9), Article e0329765. https://doi.org/10.1371/journal.pone.0329765
SAP. (n.d.). What is the industrial Internet of Things (IIoT)? Retrieved July 30, 2026, from https://www.sap.com/resources/what-is-iiot
Shafiq, D. A., Jhanjhi, N. Z., & Abdullah, A. (2022). Load balancing techniques in cloud computing environment: A review. Journal of King Saud University–Computer and Information Sciences, 34(7), 3910–3933. https://doi.org/10.1016/j.jksuci.2021.02.007
Sliwko, L. (2024). Cluster workload allocation: A predictive approach leveraging machine learning efficiency. IEEE Access, 12, 194091–194107. https://doi.org/10.1109/ACCESS.2024.3520422
Vashistha, D., Mehta, D., Kumhar, M., Bhatia, J., & Alkhayyat, A. (2025). Deep learning based load balancing in cloud computing: A survey. Procedia Computer Science, 259, 1963–1972. https://doi.org/10.1016/j.procs.2025.04.152
Walker, E. (2023). Challenges and solutions in deploying machine learning models at scale. American Journal of Machine Learning, 4(5), 13–22.
Yalamati, S. (2025). AI-enhanced fault tolerance in microservices: Predictive failure models for resilient software systems. International Journal of Applied Engineering & Technology, 7(2), 26–35.
Yao, J., Zhang, L., & Huang, J. (2025). Evaluation of large language model-driven AutoML in data and model management from human-centered perspective [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.05962
Yao, Z., Desmouceaux, Y., Townsley, M., & Heide Clausen, T. (2021). Towards intelligent load balancing in data centers [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2110.15788
Yu, Y., Abadi, M., Barham, P., Brevdo, E., Burrows, M., Davis, A., Dean, J., Ghemawat, S., Harley, T., Hawkins, P., Isard, M., Kudlur, M., Monga, R., Murray, D., & Zheng, X. (2018). Dynamic control flow in large-scale machine learning. In Proceedings of the Thirteenth EuroSys Conference (pp. 1–15). Association for Computing Machinery. https://doi.org/10.1145/3190508.3190551
Zhang, Z., Wu, Z., Rincon, D., & Christofides, P. D. (2019). Real-time optimization and control of nonlinear processes using machine learning. Mathematics, 7(10), Article 890. https://doi.org/10.3390/math7100890
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Combinatorial Optimization Problems and Informatics

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.