Does size matter? The influence of reduct size on the performance of supervised classifiers
DOI:
https://doi.org/10.61467/2007.1558.2026.v17i5.1273Keywords:
Dimensionality reduction, reduct, classifiers, reduct lengthAbstract
This study investigates the role of reduct-based attribute subsets (derived from rough set theory) with respect to supervised classification performance. Considering twenty-one datasets, two representations were assessed: minimum-length reducts containing the smallest attribute sets to retain consistency, and maximum-length reducts comprising the largest irredundant subsets. A Friedman–Nemenyi test was used to assess eleven classifiers with five quality measures. For the majority of classifiers, both types of reducts systematically degrade the predictive performance when compared with using all attributes. JRIP, Logistic Regression, and MLP are consistent and have the same performance except for PRC area, where maximum-length reducts reduce performance. For SMO, minimum-length reducts perform better than maximum-length reducts. These observations show that the efficacy of reduct-based feature selection largely depends on the characteristics of classifiers; it can provide significant dimensionality reduction without loss of accuracy for some models.
Spanish-language metadata / Metadatos en español
Título en español:
¿Importa el tamaño? La influencia del tamaño del reducto en el rendimiento de los clasificadores supervisados
Resumen:
Este estudio investiga el papel de los subconjuntos de atributos basados en reductos, derivados de la teoría de conjuntos aproximados (rough set theory), en relación con el rendimiento de la clasificación supervisada. Considerando veintiún conjuntos de datos, se evaluaron dos representaciones: reductos de longitud mínima, que contienen los conjuntos de atributos más pequeños necesarios para conservar la consistencia, y reductos de longitud máxima, que comprenden los subconjuntos irredundantes de mayor tamaño.
Se utilizó una prueba de Friedman–Nemenyi para evaluar once clasificadores mediante cinco medidas de calidad. Para la mayoría de los clasificadores, ambos tipos de reductos reducen sistemáticamente el rendimiento predictivo en comparación con el uso de todos los atributos. JRIP, Regresión Logística y MLP presentan un comportamiento consistente y muestran el mismo rendimiento, excepto en el área PRC, donde los reductos de longitud máxima reducen el rendimiento. En el caso de SMO, los reductos de longitud mínima presentan un mejor rendimiento que los reductos de longitud máxima.
Estas observaciones muestran que la eficacia de la selección de características basada en reductos depende en gran medida de las características de los clasificadores; para algunos modelos, puede proporcionar una reducción significativa de la dimensionalidad sin pérdida de exactitud.
Palabras Claves:
reducción de dimensionalidad, reducto, clasificadores, longitud del reducto
Smart citations:
References
Chen, Y., Zhu, Q., & Xu, H. (2015). Finding rough set reducts with fish swarm algorithm. Knowledge-Based Systems, 81, 22–29. https://doi.org/10.1016/j.knosys.2015.02.002
Choromański, M., Grześ, T., & Hońko, P. (2020). Breadth search strategies for finding minimal reducts: Towards hardware implementation. Neural Computing and Applications, 32, 14801–14816. https://doi.org/10.1007/s00521-020-04833-7
Frank, E., Hall, M. A., & Witten, I. H. (2016). The WEKA workbench: Online appendix for “Data mining: Practical machine learning tools and techniques” (4th ed.). Morgan Kaufmann.
González-Díaz, Y., Martínez-Trinidad, J. F., Carrasco-Ochoa, J. A., & Lazo-Cortés, M. S. (2024). An algorithm for computing all rough set constructs for dimensionality reduction. Mathematics, 12(1), Article 90. https://doi.org/10.3390/math12010090
Guyon, I., & Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of Machine Learning Research, 3, 1157–1182.
Hu, K., Diao, L., Lu, Y., & Shi, C. (2000). A heuristic optimal reduct algorithm. In K. S. Leung, L.-W. Chan, & H. Meng (Eds.), Intelligent data engineering and automated learning—IDEAL 2000 (Lecture Notes in Computer Science, Vol. 1983, pp. 139–144). Springer. https://doi.org/10.1007/3-540-44491-2_21
Jensen, R., & Shen, Q. (2003). Finding rough set reducts with ant colony optimization. In Proceedings of the 2003 UK Workshop on Computational Intelligence (pp. 15–22).
Kelly, M., Longjohn, R., & Nottingham, K. (n.d.). The UCI Machine Learning Repository. Retrieved September 1, 2026, from https://archive.ics.uci.edu
Lazo-Cortés, M. S., Martínez-Trinidad, J. F., Carrasco-Ochoa, J. A., & Sanchez Diaz, G. (2016). A new algorithm for computing reducts based on the binary discernibility matrix. Intelligent Data Analysis, 20(2), 317–337. https://doi.org/10.3233/IDA-160807
Niu, J., Chen, D., Li, J., & Wang, H. (2022). Fuzzy rule-based classification method for incremental rule learning. IEEE Transactions on Fuzzy Systems, 30(9), 3748–3761. https://doi.org/10.1109/TFUZZ.2021.3128061
Pawlak, Z. (1982). Rough sets. International Journal of Computer & Information Sciences, 11(5), 341–356. https://doi.org/10.1007/BF01001956
Pawlak, Z. (1991). Rough sets: Theoretical aspects of reasoning about data. Kluwer Academic Publishers. https://doi.org/10.1007/978-94-011-3534-4
Rodríguez-Diez, V., Martínez-Trinidad, J. F., Carrasco-Ochoa, J. A., Lazo-Cortés, M. S., & Olvera-López, J. A. (2020). MinReduct: A new algorithm for computing the shortest reducts. Pattern Recognition Letters, 138, 177–184. https://doi.org/10.1016/j.patrec.2020.07.004
Rodríguez-Diez, V., Martínez-Trinidad, J. F., Carrasco-Ochoa, J. A., Lazo-Cortés, M. S., & Olvera-López, J. A. (2021). A comparative study of two algorithms for computing the shortest reducts: MiLIT and MinReduct. In E. Roman-Rangel, Á. F. Kuri-Morales, J. F. Martínez-Trinidad, J. A. Carrasco-Ochoa, & J. A. Olvera-López (Eds.), Pattern recognition: 13th Mexican Conference, MCPR 2021 (Lecture Notes in Computer Science, Vol. 12725, pp. 57–67). Springer. https://doi.org/10.1007/978-3-030-77004-4_6
Stańczyk, U. (2023). How transformations of representation for input data can affect the properties of induced decision reducts and rules. Procedia Computer Science, 225, 3603–3612. https://doi.org/10.1016/j.procs.2023.10.355
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Combinatorial Optimization Problems and Informatics

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.