Comparative Benchmark of Eleven Regression Models for Software Effort Estimation on a COCOMO-Like Dataset

Authors

DOI:

https://doi.org/10.61467/2007.1558.2026.v17i5.1441

Keywords:

Software effort estimation, Software cost estimation, Gaussian process regression, Model tree, Stacking, benchmark, MMRE, PRED(25), COCOMO-like dataset, Regression ranking, estimación del esfuerzo de software, estimación del coste de software, regresión mediante procesos gaussianos

Abstract

This study develops a comparative benchmark for software effort estimation using the benchmark suite implemented in Python and the result package generated by that suite. Eleven regressors were compared under a leakage-safe protocol on a COCOMO-like dataset of 62 projects and 20 numeric predictors, with 49 projects reserved for model development and 13 for final hold-out testing. The evaluated methods were a three-hidden-layer deep neural network, CatBoost, XGBoost, stacked ensemble regression, random forest, support vector regression with radial basis kernel, Gaussian process regression, LightGBM, an optimized M5P model tree, gradient boosting, and AdaBoost. Model quality was judged through repeated 5×2 cross-validation on the training partition and through an independent hold-out test set, using MAE, RMSE, R², MMRE, MdMRE, and PRED(25). The repeated-CV ranking identified Gaussian process regression as the most accurate and stable model, with mean rank = 1.0000, RMSE = 0.2152, R² = 0.8998, and PRED(25) = 0.9289. The same model also dominated the hold-out evaluation, achieving MAE = 0.1444, RMSE = 0.1835, R² = 0.9084, MMRE = 0.0850, MdMRE = 0.0634, and PRED(25) = 0.9231. The overall order of precision from repeated cross-validation was: Gaussian process regression, optimized model tree (M5P), stacked ensemble regressor, gradient boosting regressor, XGBoost regressor, random forest regressor, CatBoost regressor, AdaBoost regressor, support vector regression, deep neural network, and LightGBM regressor. Permutation importance of the best model showed KDSI, ACAP, PCAP, RELY, and AAF as the most influential predictors. The evidence indicates that, for the present small-sample COCOMO-like setting, kernel-based and piecewise linear/tree-based methods outperform the deeper neural alternative.

 

Spanish-language metadata / Metadatos en español
Título en español:
Benchmark comparativo de once modelos de regresión para la estimación del esfuerzo de software en un conjunto de datos similar a COCOMO

Resumen:
Este estudio desarrolla un benchmark comparativo para la estimación del esfuerzo de software utilizando la batería de evaluación implementada en Python y el paquete de resultados generado por dicha batería. Se compararon once regresores mediante un protocolo diseñado para evitar fugas de información (data leakage) sobre un conjunto de datos similar a COCOMO compuesto por 62 proyectos y 20 predictores numéricos, de los cuales 49 proyectos se reservaron para el desarrollo de los modelos y 13 para la prueba final con un conjunto hold-out.

Los métodos evaluados fueron una red neuronal profunda con tres capas ocultas, CatBoost, XGBoost, regresión mediante un ensamble apilado (stacked ensemble), bosque aleatorio, regresión de vectores de soporte con kernel de base radial, regresión mediante procesos gaussianos, LightGBM, un árbol de modelos M5P optimizado, gradient boosting y AdaBoost.

La calidad de los modelos se evaluó mediante validación cruzada repetida 5×2 sobre la partición de entrenamiento y mediante un conjunto independiente de prueba hold-out, utilizando MAE, RMSE, R², MMRE, MdMRE y PRED(25).

La clasificación obtenida mediante validación cruzada repetida identificó la regresión mediante procesos gaussianos como el modelo más preciso y estable, con un rango medio = 1,0000, RMSE = 0,2152, R² = 0,8998 y PRED(25) = 0,9289. El mismo modelo también dominó la evaluación hold-out, alcanzando MAE = 0,1444, RMSE = 0,1835, R² = 0,9084, MMRE = 0,0850, MdMRE = 0,0634 y PRED(25) = 0,9231.

El orden global de precisión obtenido mediante validación cruzada repetida fue el siguiente: regresión mediante procesos gaussianos, árbol de modelos optimizado (M5P), regresor de ensamble apilado, regresor de gradient boosting, regresor XGBoost, regresor de bosque aleatorio, regresor CatBoost, regresor AdaBoost, regresión de vectores de soporte, red neuronal profunda y regresor LightGBM.

El análisis de importancia por permutación del mejor modelo mostró que KDSI, ACAP, PCAP, RELY y AAF fueron los predictores más influyentes. La evidencia indica que, para el presente escenario de tamaño muestral reducido y similar a COCOMO, los métodos basados en kernels y los enfoques lineales por tramos o basados en árboles superan a la alternativa neuronal más profunda.

Palabras Claves:
estimación del esfuerzo de software, estimación del coste de software, regresión mediante procesos gaussianos, árbol de modelos, stacking, benchmark, MMRE, PRED(25), conjunto de datos similar a COCOMO, clasificación de regresores


Smart citations:

SciteAI. 
Dimensions.
Open Alex.

References

Arora, S., & Mishra, N. (2017). Software cost estimation using single layer artificial neural network. International Journal of Advanced Engineering Research and Science, 4(9), 22–26. doi:10.22161/ijaers.4.9.6

Boehm, B. W. (1981). Software engineering economics. Prentice-Hall.

Boehm, B. W., Abts, C., Brown, A. W., Chulani, S., Clark, B. K., Horowitz, E., Madachy, R., Reifer, D. J., & Steece, B. (2000). Software cost estimation with COCOMO II. Prentice Hall.

Boehm, B. W., Abts, C., Brown, A. W., Chulani, S., Clark, B. K., Horowitz, E., Madachy, R., Reifer, D. J., & Steece, B. (2000). Software cost estimation with COCOMO II. Prentice Hall.

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. doi:10.1023/A:1010933404324

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). Association for Computing Machinery. doi:10.1145/2939672.2939785

Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20, 273–297. doi:10.1007/BF00994018

Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232. doi:10.1214/aos/1013203451

Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.

Hoerl, A. E., & Kennard, R. W. (1970). Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1), 55–67. doi:10.1080/00401706.1970.10488634

Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In F. Bach & D. Blei (Eds.), Proceedings of the 32nd International Conference on Machine Learning (Vol. 37, pp. 448–456). PMLR.

Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30, 3146–3154.

Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. International Conference on Learning Representations.

Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. International Conference on Learning Representations.

Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., & Gulin, A. (2018). CatBoost: Unbiased boosting with categorical features. Advances in Neural Information Processing Systems, 31, 6638–6648.

Quinlan, J. R. (1992). Learning with continuous classes. In A. Adams & L. Sterling (Eds.), AI ’92: Proceedings of the 5th Australian Joint Conference on Artificial Intelligence (pp. 343–348). World Scientific.

Rasmussen, C. E., & Williams, C. K. I. (2006). Gaussian processes for machine learning. MIT Press. doi:10.7551/mitpress/3206.001.0001

Smola, A. J., & Schölkopf, B. (2004). A tutorial on support vector regression. Statistics and Computing, 14(3), 199–222. doi:10.1023/B:STCO.0000035301.49549.88

Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15, 1929–1958.

Wang, Y., & Witten, I. H. (1996). Induction of model trees for predicting continuous classes (Working Paper 96/23). University of Waikato, Department of Computer Science.

Wang, Y., & Witten, I. H. (1997). Induction of model trees for predicting continuous classes. In M. van Someren & G. Widmer (Eds.), Poster papers of the 9th European Conference on Machine Learning (ECML’97) (pp. 128–137). Laboratory of Intelligent Systems.

Wolpert, D. H. (1992). Stacked generalization. Neural Networks, 5(2), 241–259. doi:10.1016/S0893-6080(05)80023-1

Downloads

Published

2026-09-06

How to Cite

Aguilar-Ortiz, J., Zamudio-García, V. M., Gómez-Ramos, M. Y., & Domínguez-Mayorga, C. R. (2026). Comparative Benchmark of Eleven Regression Models for Software Effort Estimation on a COCOMO-Like Dataset. International Journal of Combinatorial Optimization Problems and Informatics, 17(5), 80–105. https://doi.org/10.61467/2007.1558.2026.v17i5.1441

Issue

Section

Articles

Most read articles by the same author(s)