Predictive maintenance using administrative work orders: A machine learning framework for failure criticality classification.
| dc.contributor.advisor | García Rodriguez, Alejandro | |
| dc.contributor.author | Cruz Martínez, Juan Andrés | |
| dc.contributor.corporatename | Universidad Santo Tomas | |
| dc.contributor.orcid | 0009-0008-8952-0198 | |
| dc.date.accessioned | 2026-07-30T17:59:55Z | |
| dc.date.available | 2026-07-30T17:59:55Z | |
| dc.date.issued | 2026-07-21 | |
| dc.description | Este estudio evalúa la viabilidad del mantenimiento predictivo utilizando únicamente registros históricos de mantenimiento correctivo de una planta de la industria alimentaria, sin datos de sensores. El conjunto de datos comprendió 6731 eventos de falla en 664 máquinas, con nueve variables operacionales. La criticidad de la falla se clasificó en baja (≤23 min), media (24–44 min) y alta (>44 min) según los percentiles de tiempo de inactividad. Se desarrollaron tres modelos basados en árboles (Random Forest, XGBoost, LightGBM) y un perceptrón multicapa. La optimización de hiperparámetros se realizó utilizando RandomizedSearchCV y Optuna, este último con una función objetivo personalizada penalizada por brecha para mitigar el sobreajuste. Una división cronológica (70/15/15) respetó el orden temporal. El mejor modelo, XGBoost con Optuna penalizada por brecha, alcanzó un F1 macro de prueba de 0,3829 y un F1 ponderado de 0,4871, con una brecha de generalización mínima. La red neuronal obtuvo el F1 ponderado más alto (0,5512), pero falló en la clase de criticidad media. Los resultados confirman que los registros administrativos por sí solos tienen un techo predictivo estructural, y que la penalización explícita de brechas es efectiva para potenciar los modelos. El marco propuesto permite la comparación sistemática de modelos neuronales y basados en árboles para la clasificación de criticidad de fallas utilizando únicamente registros de órdenes de trabajo. Si bien los métodos estadísticos como ANOVA pueden establecer si existen diferencias significativas entre grupos operativos, no proporcionan un marco predictivo capaz de clasificar nuevos eventos de falla en el momento en que ocurren. Por el contrario, los modelos de aprendizaje automático están diseñados para identificar y generalizar patrones a partir de datos históricos, pero no explican inherentemente la estructura estadística de las variables subyacentes. Este estudio aborda ambas dimensiones de manera complementaria: se utilizan ANOVA y pruebas post-hoc Tukey HSD para validar estadísticamente la relevancia de la identidad de la máquina, la interacción del turno de trabajo y el número de técnicos en el tiempo de inactividad del equipo, mientras que posteriormente se desarrollan modelos de aprendizaje automático para evaluar si esos patrones son aprendibles y generalizables a eventos de falla no vistos. Este enfoque dual permite una caracterización más completa de la dinámica de fallas industriales que cualquiera de los métodos por separado. | |
| dc.description.abstract | This study evaluates the feasibility of predictive maintenance using only historical correc-tive maintenance records from a food industry plant, without sensor data. The dataset comprised 6,731 failure events across 664 machines, with nine operational variables. Failure criticality was classified into low (≤23 min), medium (24–44 min), and high (>44 min) based on downtime percentiles. Three tree based models (Random Forest, XGBoost, LightGBM) and a multilayer perceptron were developed. Hyperparameter optimization was performed using RandomizedSearchCV and Optuna, the latter with a custom gap penalized objective function to mitigate overfitting. A chronological split (70/15/15) respected temporal order. The best model, XGBoost with gap penalized Optuna, achieved a test Macro F1 of 0.3829 and a weighted F1 of 0.4871, with a minimal generalization gap. The neural network obtained the highest weighted F1 (0.5512) but failed on the medium criticality class. Results confirm that administrative records alone have a structural pre-dictive ceiling, and that explicit gap penalization is effective for boosting models. The proposed framework enables systematic comparison of tree based and neural models for failure criticality classification using only work order logs. While statistical methods such as ANOVA can establish whether significant differences exist between operational groups, they do not provide a predictive framework capable of classifying new failure events at the moment they occur. Conversely, machine learning models are designed to identify and generalize patterns from historical data, but do not inherently explain the statistical structure of the underlying variables. This study addresses both dimensions in a complementary fashion: ANOVA and post-hoc Tukey HSD tests are used to statistically validate the relevance of machine identity, work shift interaction, and technician count on equipment downtime, while machine learning models are subsequently developed to as-sess whether those patterns are learnable and generalizable to unseen failure events. This dual approach enables a more complete characterization of industrial failure dynamics than either method alone. | |
| dc.description.degreelevel | Pregrado | spa |
| dc.description.degreename | Ingeniero Mecánico | spa |
| dc.format.mimetype | application/pdf | |
| dc.identifier.citation | Cruz Martinez, J.M(2026). Predictive Maintenance Using Administrative Work Orders: A Machine Learning Framework For Failure Criticality Classification.[Trabajo de grado Pregrado, Universidad Santo Tomás]. Reposito institucional | |
| dc.identifier.instname | instname:Universidad Santo Tomás | spa |
| dc.identifier.reponame | reponame:Repositorio Institucional Universidad Santo Tomás | spa |
| dc.identifier.repourl | repourl:https://repository.usta.edu.co | spa |
| dc.identifier.uri | http://hdl.handle.net/11634/73729 | |
| dc.language.iso | spa | |
| dc.publisher | Universidad Santo Tomás | spa |
| dc.publisher.branch | CRAI-USTA Bogotá | |
| dc.publisher.faculty | Facultad de Ingeniería Mecánica | spa |
| dc.publisher.program | Pregrado Ingeniería Mecánica | spa |
| dc.relation.references | [1] A. A. Sisode and M. Devare, “A Review on Machine Learning Techniques for Predictive Maintenance in Industry 4.0,” International Research Journal of Engineering and Technology (IRJET), vol. 9, no. 8, pp. 1169–1173, Aug. 2022. | |
| dc.relation.references | [2] E. Samatas, S. Siatras, and I. Chatzigiannakis, “Predictive Maintenance: Bridging Artificial Intelligence and IoT,” in Pro-ceedings of the 2021 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA), Dalian, China, Jun. 2021, pp. 464–471. | |
| dc.relation.references | [3] P. Susto, A. Schirru, S. Pampuri, S. McLoone, and A. Beghi, “Machine Learning for Predictive Maintenance: A Multiple Classifier Approach,” IEEE Transactions on Industrial Informatics, vol. 11, no. 3, pp. 812–820, Jun. 2015. | |
| dc.relation.references | [4] C. Campos-Olivares, D. Hernández-Gress, R. Rodríguez-Molina, and E. Olivares-Benitez, “A Systematic Mapping of the Advancing Use of Machine Learning Techniques for Predictive Maintenance,” DYNA, vol. 91, no. 1, pp. 73–81, Jan. 2024. | |
| dc.relation.references | [5] I. Hector and R. Panjanathan, “Predictive maintenance in Industry 4.0: A survey of planning models and machine learning techniques,” PeerJ Computer Science, vol. 10, p. e2016, 2024, doi: 10.7717/peerj-cs.2016. | |
| dc.relation.references | [6] B. van Oudenhoven, P. Van de Calseyde, R. Basten, and E. Demerouti, “Predictive maintenance for Industry 5.0: Behavioural inquiries from a work system perspective,” International Journal of Production Research, vol. 61, no. 22, pp. 7846–7865, 2023, doi: 10.1080/00207543.2022.2154403. | |
| dc.relation.references | [7] P. Karrupusamy, “Machine learning approach to predictive maintenance in manufacturing industry – A comparative study,” Journal of Soft Computing Paradigm (JSCP), vol. 2, no. 4, pp. 246–255, 2020, doi: 10.36548/jscp.2020.4.006. | |
| dc.relation.references | [8] D. C. Montgomery, Design and Analysis of Experiments, 9th ed. Wiley, 2017. | |
| dc.relation.references | [9] J. Han, M. Kamber, y J. Pei, Data Mining: Concepts and Techniques, 3ra ed. Morgan Kaufmann, 2011 | |
| dc.relation.references | [10] T. Wireman, Developing Performance Indicators for Managing Maintenance. Morgan Kaufmann, 2004. | |
| dc.relation.references | [11] B. W. Silverman, Density Estimation for Statistics and Data Analysis. Chapman and Hall, 1986. | |
| dc.relation.references | [12] S. Sundaram and A. Zeid, "Technical language processing for Prognostics and Health Management: applying text simi-larity and topic modeling to maintenance work orders," Journal of Intelligent Manufacturing, Feb. 2024. | |
| dc.relation.references | [13] S. Nakajima, Introduction to TPM: Total Productive Maintenance. Productivity Press, Cambridge, MA, 1988. | |
| dc.relation.references | [14] U. M. Fayyad and K. B. Irani, "Multi-interval discretization of continuous-valued attributes for classification learning," in Proc. 13th IJCAI, Chambéry, France, 1993, pp. 1022–1027. | |
| dc.relation.references | [15] M. Sivakumar, S. Parthasarathy, and T. Padmapriya, "Trade-off between training and testing ratio in machine learning for medical image processing," PeerJ Computer Science, vol. 10, p. e2245, 2024. doi: 10.7717/peerj-cs.2245 | |
| dc.relation.references | [16] Y. Xu and R. Goodacre, "On Splitting Training and Validation Set: A Comparative Study of Cross-Validation, Bootstrap and Systematic Sampling for Estimating the Generalization Performance of Supervised Learning," Journal of Analysis and Testing, vol. 2, pp. 249–262, 2018. doi: 10.1007/s41664-018-0068-2 | |
| dc.relation.references | [17] R. J. Hyndman and G. Athanasopoulos, Forecasting: Principles and Practice, 3rd ed. Melbourne, Australia | |
| dc.relation.references | [18] H. R. Friesacher, E. Svensson, S. Winiwarter, L. Mervin, A. Arany, and O. Engkvist, “Temporal distribution shift in re-al-world pharmaceutical data: Implications for uncertainty quantification in QSAR models,” Artificial Intelligence in the Life Sciences, vol. 8, p. 100132, 2025, doi: 10.1016/j.ailsci.2025.100132. | |
| dc.relation.references | [19] J. Bergstra and Y. Bengio, “Random search for Hyper-Parameter Optimization,” 2012. https://jmlr.org/beta/papers/v13/bergstra12a.html | |
| dc.relation.references | [20] T. Chen and C. Guestrin, "XGBoost: A Scalable Tree Boosting System," in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, San Francisco, CA, 2016, pp. 785–794. doi: 10.1145/2939672.2939785 | |
| dc.relation.references | [21] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, "LightGBM: A Highly Efficient Gradient Boosting Decision Tree," in Proc. 31st Conf. Neural Information Processing Systems (NIPS 2017), Long Beach, CA, 2017. | |
| dc.relation.references | [22] R. Ballester, X. Arnal Clemente, C. Casacuberta, M. Madadi, C. A. Corneanu, and S. Escalera, “Predicting the generalization gap in neural networks using topological data analysis,” arXiv preprint arXiv:2203.12330v2, 2023. | |
| dc.relation.references | [23] L. Breiman, “Random Forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001. | |
| dc.relation.references | [24] P. Probst, M. N. Wright, and A. L. Boulesteix, “Hyperparameters and tuning strategies for random forest,” Wiley Inter-disciplinary Reviews: Data Mining and Knowledge Discovery, vol. 9, no. 3, e1301, 2019. | |
| dc.relation.references | [25] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, "Optuna: A Next-generation Hyperparameter Optimization Framework," in Proc. 25th ACM SIGKDD Int. Conf. Knowledge Discovery & Data Mining, Anchorage, AK, 2019, pp. 2623–2631. doi: 10.1145/3292500.3330701 | |
| dc.rights | Attribution-NonCommercial-NoDerivs 2.5 Colombia | en |
| dc.rights.accessrights | info:eu-repo/semantics/openAccess | |
| dc.rights.coar | http://purl.org/coar/access_right/c_abf2 | |
| dc.rights.local | Abierto (Texto Completo) | spa |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/2.5/co/ | |
| dc.subject.keyword | Machine learning | |
| dc.subject.keyword | Predictive maintenance | |
| dc.subject.keyword | Random Forest | |
| dc.subject.keyword | XGBoost | |
| dc.subject.keyword | Classification models | |
| dc.subject.lemb | Ingenieria Mecanica | |
| dc.subject.lemb | Bosque aleatorio | |
| dc.subject.lemb | Modelos de clasificación | |
| dc.subject.lemb | XGBoost | |
| dc.subject.proposal | Aprendizaje automático | |
| dc.subject.proposal | Mantenimiento predictivo | |
| dc.subject.proposal | Bosque aleatorio | |
| dc.subject.proposal | XGBoost | |
| dc.subject.proposal | Modelos de clasificación | |
| dc.title | Predictive maintenance using administrative work orders: A machine learning framework for failure criticality classification. | |
| dc.type | bachelor thesis | |
| dc.type.coar | http://purl.org/coar/resource_type/c_7a1f | |
| dc.type.coarversion | http://purl.org/coar/version/c_ab4af688f83e57aa | |
| dc.type.drive | info:eu-repo/semantics/bachelorThesis | |
| dc.type.local | Trabajo de grado | spa |
| dc.type.version | info:eu-repo/semantics/acceptedVersion |
Archivos
Bloque original
1 - 1 de 1
Bloque de licencias
1 - 3 de 3
Cargando...
- Nombre:
- license.txt
- Tamaño:
- 807 B
- Formato:
- Item-specific license agreed upon to submission
- Descripción:
Cargando...
- Nombre:
- 2026cartaderechosdeautor.pdf
- Tamaño:
- 884.6 KB
- Formato:
- Adobe Portable Document Format
- Descripción:
- Carta derechos de autor
Cargando...
- Nombre:
- 2026cartadefacultad.pdf
- Tamaño:
- 156.72 KB
- Formato:
- Adobe Portable Document Format
- Descripción:
- Carta de Facultad

