Evaluating the SEMMA Methodology with AutoML Across Phishing Detection, Pollen Classification, and Voice Identification Tasks
DOI:
https://doi.org/10.61467/2007.1558.2026.v17i4.1290Keywords:
SEMMA methodology, automated machine learning, phishing detection, pollen classification, voice identification, Support Vector Machine, metodología SEMMA, aprendizaje automático automatizado, clasificación de polenAbstract
The exponential growth of data has led to an increasing demand for effective methodologies in the field of machine learning. SEMMA (Sample, Explore, Modify, Model, and Assess) stands out among these methodologies due to its clear and efficient structure for pattern extraction. The contribution of this study lies in the application of the SEMMA methodology to a variety of case studies involving historical data, images, and audio. Specific datasets were used to address phishing detection, pollen image analysis, and voice classification in audio recordings. The results obtained highlight the effectiveness of SEMMA in solving data analysis problems across various domains and disciplines. Furthermore, potential future research directions in this field are identified.
Spanish-language metadata / Metadatos en español
Título en español:
Evaluación de la metodología SEMMA con AutoML en tareas de detección de phishing, clasificación de polen e identificación de voz
Resumen:
El crecimiento exponencial de los datos ha generado una demanda cada vez mayor de metodologías eficaces en el campo del aprendizaje automático. SEMMA (Muestrear, Explorar, Modificar, Modelar y Evaluar) destaca entre estas metodologías debido a su estructura clara y eficiente para la extracción de patrones. La contribución de este estudio radica en la aplicación de la metodología SEMMA a diversos estudios de caso que involucran datos históricos, imágenes y audio. Se utilizaron conjuntos de datos específicos para abordar la detección de phishing, el análisis de imágenes de polen y la clasificación de voz en grabaciones de audio. Los resultados obtenidos destacan la eficacia de SEMMA para resolver problemas de análisis de datos en diversos ámbitos y disciplinas. Además, se identifican posibles líneas futuras de investigación en este campo.
Palabras Claves: metodología SEMMA; aprendizaje automático automatizado; selección de modelos; detección de phishing; clasificación de polen; identificación de voz; bosque aleatorio; máquina de vectores de soporte.
Smart citations:
https://scite.ai/reports/10.61467/2007.1558.2026.v17i4.1290
Dimensions.
Open Alex.
References
Ahmad, Z., Yaacob, S., Ibrahim, R., & Wan Fakhruddin, W. F. (2022). The review for visual analytics methodology. In 2022 International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA) (pp. 1–10). IEEE. https://doi.org/10.1109/HORA55278.2022.9800100
Azevedo, A., & Santos, M. F. (2008). KDD, SEMMA and CRISP-DM: A parallel overview. In H. Weghorn & A. P. Abraham (Eds.), Proceedings of the IADIS European Conference on Data Mining 2008 (pp. 182–185). IADIS Press.
Benesty, J., Chen, J., Huang, Y., & Cohen, I. (2009). Pearson correlation coefficient. In Noise reduction in speech processing (pp. 1–4). Springer. https://doi.org/10.1007/978-3-642-00296-0_5
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
Cios, K. J., Pedrycz, W., & Swiniarski, R. W. (1998). Machine learning. In Data mining methods for knowledge discovery (pp. 229–308). Springer.
Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297. https://doi.org/10.1007/BF00994018
Costa, C. J., & Aparicio, J. T. (2020). POST-DS: A methodology to boost data science. In 2020 15th Iberian Conference on Information Systems and Technologies (CISTI) (pp. 1–6). IEEE. https://doi.org/10.23919/CISTI49556.2020.9140932
De Ville, B. (2013). Decision trees. Wiley Interdisciplinary Reviews: Computational Statistics, 5(6), 448–455. https://doi.org/10.1002/wics.1278
De-La-Hoz-Correa, E., Mendoza-Palechor, F. E., De-La-Hoz-Manotas, A., Morales-Ortega, R. C., & Sánchez Hernández, B. A. (2019). Obesity level estimation software based on decision trees. Journal of Computer Science, 15(1), 67–77. https://doi.org/10.3844/jcssp.2019.67.77
Dong, X., Yu, Z., Cao, W., Shi, Y., & Ma, Q. (2020). A survey on ensemble learning. Frontiers of Computer Science, 14(2), 241–258. https://doi.org/10.1007/s11704-019-8208-z
Duville, M. M., Alonso-Valerdi, L. M., & Ibarra-Zarate, D. I. (2021). The Mexican Emotional Speech Database (MESD): Elaboration and assessment based on machine learning. In 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) (pp. 1644–1647). IEEE. https://doi.org/10.1109/EMBC46164.2021.9629934
Fayyad, U., Piatetsky-Shapiro, G., & Smyth, P. (1996). From data mining to knowledge discovery in databases. AI Magazine, 17(3), 37–54. https://doi.org/10.1609/aimag.v17i3.1230
Fern, A., & Givan, R. (2003). Online ensemble learning: An empirical study. Machine Learning, 53(1–2), 71–109. https://doi.org/10.1023/A:1025619426553
Gama, J. (2010). Knowledge discovery from data streams. Chapman & Hall/CRC.
Hansun, S., Suryadibrata, A., Nurhasanah, R., & Fitra, J. (2022). Tweets sentiment on PPKM policy as a COVID-19 response in Indonesia. Indian Journal of Computer Science and Engineering, 13(1), 51–58.
Holzinger, A. (2013). Human-computer interaction and knowledge discovery (HCI-KDD): What is the benefit of bringing those two fields to work together? In A. Cuzzocrea, C. Kittl, D. E. Simos, E. Weippl, & L. Xu (Eds.), Availability, reliability, and security in information systems and HCI (Lecture Notes in Computer Science, Vol. 8127, pp. 319–328). Springer. https://doi.org/10.1007/978-3-642-40511-2_22
Jacob, D., & Henriques, R. (2023). Educational data mining to predict bachelors students’ success. Emerging Science Journal, 7, 159–171. https://doi.org/10.28991/ESJ-2023-SIED2-013
Jadrić, M., Garača, Ž., & Čukušić, M. (2010). Student dropout analysis with application of data mining methods. Management: Journal of Contemporary Management Issues, 15(1), 31–46.
Li, G., Tan, J., & Chaudhry, S. S. (2019). Industry 4.0 and big data innovations. Enterprise Information Systems, 13(2), 145–147. https://doi.org/10.1080/17517575.2018.1554190
Lin, F. Y., & McClean, S. (2001). A data mining approach to the prediction of corporate failure. Knowledge-Based Systems, 14(3–4), 189–195. https://doi.org/10.1016/S0950-7051(01)00096-X
López-Torres, S., López-Torres, H., Rocha-Rocha, J., Butt, S. A., Tariq, M. I., Collazos-Morales, C., & Piñeres-Espitia, G. (2020). IoT monitoring of water consumption for irrigation systems using SEMMA methodology. In U. S. Tiwary & S. Chaudhury (Eds.), Intelligent human computer interaction (pp. 222–234). Springer. https://doi.org/10.1007/978-3-030-44689-5_20
Mohd Selamat, S. A., Prakoonwit, S., Sahandi, R., Khan, W., & Ramachandran, M. (2018). Big data analytics—A review of data-mining models for small and medium enterprises in the transportation sector. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 8(3), Article e1238. https://doi.org/10.1002/widm.1238
Palacios, H. J. G., Toledo, R. A. J., Pantoja, G. A. H., & Navarro, Á. A. M. (2017). A comparative between CRISP-DM and SEMMA through the construction of a MODIS repository for studies of land use and cover change. Advances in Science, Technology and Engineering Systems Journal, 2(3), 598–604. https://doi.org/10.25046/aj020376
Panov, P., & Džeroski, S. (2007). Combining bagging and random subspaces to create better ensembles. In M. R. Berthold, J. Shawe-Taylor, & N. Lavrač (Eds.), Advances in intelligent data analysis VII (Lecture Notes in Computer Science, Vol. 4723, pp. 118–129). Springer. https://doi.org/10.1007/978-3-540-74825-0_11
Rogalewicz, M., & Sika, R. (2016). Methodologies of knowledge discovery from data and data mining methods in mechanical engineering. Management and Production Engineering Review, 7(4), 97–108. https://doi.org/10.1515/mper-2016-0040
Sarker, O., Jayatilaka, A., Haggag, S., Liu, C., & Babar, M. A. (2024). A multi-vocal literature review on challenges and critical success factors of phishing education, training and awareness. Journal of Systems and Software, 208, Article 111899. https://doi.org/10.1016/j.jss.2023.111899
SAS Institute. (2003). Data mining using SAS Enterprise Miner: A case study approach.
Sazonau, V. (2012). Implementation and evaluation of a random forest machine learning algorithm [Unpublished report]. University of Manchester.
Shalev-Shwartz, S., & Ben-David, S. (2014). Understanding machine learning: From theory to algorithms. Cambridge University Press. https://doi.org/10.1017/CBO9781107298019
Tello-Mijares, S., & Flores, F. (2016). A novel method for the separation of overlapping pollen species for automated detection and classification. Computational and Mathematical Methods in Medicine, 2016, Article 5689346. https://doi.org/10.1155/2016/5689346
Thornton, C., Hutter, F., Hoos, H. H., & Leyton-Brown, K. (2013). Auto-WEKA: Combined selection and hyperparameter optimization of classification algorithms. In Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 847–855). Association for Computing Machinery. https://doi.org/10.1145/2487575.2487629
Vistro, D. M., Rasheed, F., & David, L. G. (2019). The cricket winner prediction with application of machine learning and data analytics. International Journal of Scientific & Technology Research, 8(9), 985–990.
Vrbančič, G., Fister, I., Jr., & Podgorelec, V. (2020). Datasets for phishing websites detection. Data in Brief, 33, Article 106438. https://doi.org/10.1016/j.dib.2020.106438
Zavaleta-Sánchez, E., Domínguez-Sánchez, G., Loeza-Mejía, C.-I., & Sánchez-DelaCruz, E. (2024). Comparative study of KDD and CRISP-DM methodologies for phishing identification. In X.-S. Yang, S. Sherratt, N. Dey, & A. Joshi (Eds.), Proceedings of Ninth International Congress on Information and Communication Technology (Lecture Notes in Networks and Systems, Vol. 1013, pp. 317–330). Springer. https://doi.org/10.1007/978-981-97-3559-4_25
Zurada, J., & Karwowski, W. (2011). Knowledge discovery through experiential learning from business and other contemporary data sources: A review and reappraisal. Information Systems Management, 28(3), 258–274. https://doi.org/10.1080/10580530.2010.493846
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Combinatorial Optimization Problems and Informatics

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.