Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

Machine Learning for Predicting Job States and CPU Power on a Supercomputer

  • National High Technology Center

Producción científica: Capítulo del libro/informe/acta de congresoContribución a la conferenciarevisión exhaustiva

Resumen

Efficient resource management in high-performance computing (HPC) is essential for optimizing costs, reducing energy consumption, and improving system productivity. However, job variability and failures introduce uncertainties that complicate scheduling and resource allocation. Accurately predicting job failures and estimating energy consumption can enhance planning and operational efficiency. This study analyzes data from the Simple Linux Utility for Resource Management (SLURM) on the Kabré supercomputer at Costa Rica’s National High Technology Center (CeNAT). After selecting and preprocessing relevant variables, a dataset was created to train a two-stage machine learning model comprising a binary classifier and a regression model. Using 10-fold cross-validation, multiple models were evaluated, with Random Forest emerging as the best performer in both stages. The classification model was assessed using the confusion matrix and ROC curve, while the regression model was evaluated through residual analysis and metrics such as Root Mean Square Error (RMSE) and the Coefficient of Determination (R2). This approach can support users and administrators by improving job scheduling decisions and reducing energy waste in HPC systems.

Idioma originalInglés
Título de la publicación alojadaHigh Performance Computing - 12th Latin American High Performance Computing Conference, CARLA 2025, Proceedings
EditoresKevin Brown, Kyle Felker, Esteban Meneses, Antônio Tadeu Azevedo Gomes, José Manuel Monsalve Diaz, Katherine Rasmussen
EditorialSpringer Science and Business Media Deutschland GmbH
Páginas109-124
Número de páginas16
ISBN (versión impresa)9783032249227
DOI
EstadoPublicada - 2026
Evento12th Latin American Conference on High Performance Computing, CARLA 2025 - Kingston, Jamaica
Duración: 22 sept 202526 sept 2025

Serie de la publicación

NombreCommunications in Computer and Information Science
Volumen2750 CCIS
ISSN (versión impresa)1865-0929
ISSN (versión digital)1865-0937

Conferencia

Conferencia12th Latin American Conference on High Performance Computing, CARLA 2025
País/TerritorioJamaica
CiudadKingston
Período22/09/2526/09/25

ODS de las Naciones Unidas

Este resultado contribuye a los siguientes Objetivos de Desarrollo Sostenible

  1. ODS 7: Energía asequible y no contaminante
    ODS 7: Energía asequible y no contaminante

Huella

Profundice en los temas de investigación de 'Machine Learning for Predicting Job States and CPU Power on a Supercomputer'. En conjunto forman una huella única.

Citar esto