Cayley–Hamilton-Guided Krylov Regularization for Machine Learning: Theory, Computational Architecture, and Empirical Evidence

Authors

DOI:

https://doi.org/10.61467/2007.1558.2026.v17i5.1420

Keywords:

Cayley–Hamilton theorem, matrix polynomials, Krylov subspaces, regularized least-squares classification, conjugate gradient, machine learning operators, Teorema de Cayley–Hamilton, subespacios de Krylov, gradiente conjugado

Abstract

This study develops and evaluates a Cayley–Hamilton-based framework for regularised machine learning. Its objective is to formulate a practical Krylov-subspace architecture for classification that replaces repeated dense inversion with stable operator iteration. The formulation draws on the theorem’s finite polynomial closure for square matrices. The analysis connects the characteristic-polynomial identity with regularised least-squares learning, derives full and truncated conjugate-gradient solvers, and evaluates the proposed method on four datasets. Full conjugate gradients reproduce the direct solution to machine precision, whereas truncated iteration preserves competitive predictive accuracy while reducing training cost and showing greater tolerance to perturbations in ill-conditioned data. These findings support the interpretation of the Cayley–Hamilton theorem not merely as an algebraic identity, but as a design principle for finite-dimensional learning operators, efficient computation and interpretable matrix-polynomial models.

 

Spanish-language metadata / Metadatos en español
Título en español:
Regularización de Krylov guiada por Cayley–Hamilton para aprendizaje automático:
Teoría, arquitectura computacional y evidencia empírica

Resumen:
Este estudio desarrolla y evalúa un marco basado en Cayley–Hamilton para el aprendizaje automático regularizado. Su objetivo es formular una arquitectura práctica de subespacios de Krylov para clasificación que sustituya la inversión densa repetida por una iteración estable de operadores. La formulación se fundamenta en el cierre polinómico finito del teorema para matrices cuadradas. El análisis vincula la identidad del polinomio característico con el aprendizaje regularizado mediante mínimos cuadrados, deriva solucionadores de gradiente conjugado completos y truncados, y evalúa el método propuesto en cuatro conjuntos de datos. Los gradientes conjugados completos reproducen la solución directa con precisión de máquina, mientras que la iteración truncada mantiene una precisión predictiva competitiva al tiempo que reduce el coste de entrenamiento y muestra una mayor tolerancia a perturbaciones en datos mal condicionados. Estos resultados respaldan la interpretación del teorema de Cayley–Hamilton no meramente como una identidad algebraica, sino como un principio de diseño para operadores de aprendizaje de dimensión finita, computación eficiente y modelos interpretables basados en polinomios matriciales.

Palabras Claves:
Teorema de Cayley–Hamilton, polinomios matriciales, subespacios de Krylov, clasificación regularizada por mínimos cuadrados, gradiente conjugado, operadores de aprendizaje automático.


Smart citations:

https://scite.ai/reports/10.61467/2007.1558.2026.v17i5.1420
Dimensions.
Open Alex.

References

Björck, Å. (1996). Numerical methods for least squares problems. Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9781611971484

Engl, H. W., Hanke, M., & Neubauer, A. (1996). Regularization of inverse problems. Kluwer Academic Publishers. https://doi.org/10.1007/978-94-009-1740-8

Gazzola, S., & Sabaté Landman, M. (2020). Krylov methods for inverse problems: Surveying classical, and introducing new, algorithmic approaches. GAMM-Mitteilungen, 43(4), e202000017. https://doi.org/10.1002/gamm.202000017

Gazzola, S., Novati, P., & Russo, M. R. (2015). On Krylov projection methods and Tikhonov regularization. Electronic Transactions on Numerical Analysis, 44, 83–123. ETNA full text

Gohberg, I., Lancaster, P., & Rodman, L. (2009). Matrix polynomials. Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9780898719024

Golub, G. H., Hansen, P. C., & O’Leary, D. P. (1999). Tikhonov regularization and total least squares. SIAM Journal on Matrix Analysis and Applications, 21(1), 185–194. https://doi.org/10.1137/S0895479897326432

Greenbaum, A. (1997). Iterative methods for solving linear systems. Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9781611970937

Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7

Hestenes, M. R., & Stiefel, E. (1952). Methods of conjugate gradients for solving linear systems. Journal of Research of the National Bureau of Standards, 49(6), 409–436. https://doi.org/10.6028/jres.049.044

Hoerl, A. E., & Kennard, R. W. (1970). Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1), 55–67. https://doi.org/10.1080/00401706.1970.10488634

Horn, R. A., & Johnson, C. R. (2012). Matrix analysis (2nd ed.). Cambridge University Press. https://doi.org/10.1017/CBO9781139020411

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. JMLR article

Qin, S. J., Liu, Y., & Tang, S. (2023). Partial least squares, steepest descent, and conjugate gradient for regularized predictive modeling. AIChE Journal, 69(4), e17992. https://doi.org/10.1002/aic.17992

Rifkin, R. M., & Klautau, A. (2004). In defense of one-vs-all classification. Journal of Machine Learning Research, 5, 101–141. JMLR article

Rifkin, R. M., Yeo, G. W., & Poggio, T. (2003). Regularized least-squares classification. In J. A. K. Suykens, G. Horvath, S. Basu, C. Micchelli, & J. Vandewalle (Eds.), Advances in learning theory: Methods, models and applications (pp. 131–154). IOS Press.

Saad, Y. (2003). Iterative methods for sparse linear systems (2nd ed.). Society for Industrial and Applied Mathematics. https://doi.org/10.1137/1.9780898718003

Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427–437. https://doi.org/10.1016/j.ipm.2009.03.002

Suykens, J. A. K., & Vandewalle, J. (1999). Least squares support vector machine classifiers. Neural Processing Letters, 9(3), 293–300. https://doi.org/10.1023/A:1018628609742

Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., . . . van Mulbregt, P. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17(3), 261–272. https://doi.org/10.1038/s41592-019-0686-2

Downloads

Published

2026-09-06

How to Cite

Aguilar-Ortiz, J. (2026). Cayley–Hamilton-Guided Krylov Regularization for Machine Learning: Theory, Computational Architecture, and Empirical Evidence. International Journal of Combinatorial Optimization Problems and Informatics, 17(5), 61–79. https://doi.org/10.61467/2007.1558.2026.v17i5.1420

Issue

Section

Articles

Most read articles by the same author(s)