Información Clínica en Español: Cómo Aprovecharla sin Poner en Riesgo la Privacidad

dc.contributor.advisorPiracoca Arcos, Jhon Alexis
dc.contributor.authorOspina Leguizamón, Daniel Mauricio
dc.contributor.corporatenameUniversidad Santo Tomás
dc.contributor.cvlachttps://scienti.minciencias.gov.co/cvlac/visualizador/generarCurriculoCv.do?cod_rh=0002450348
dc.date.accessioned2026-10-01T19:46:35Z
dc.date.available2026-10-01T19:46:35Z
dc.date.issued2026-09-30
dc.descriptionEl crecimiento de la información clínica en formato no estructurado, como las historias clínicas, las notas médicas y los reportes narrativos, deja una gran tarea al tener como principal reto su análisis y aprovechamiento. En este caso, el procesamiento de lenguaje natural (NLP) y los modelos de lenguaje pueden comprenderse como herramientas clave y de gran repercusión para poder extraer y tener una comprensión más directa y exacta de su contenido sin dejar de lado la protección de información relevante expuesta en los mismos. Se desarrolló este artículo de divulgación mediante una revisión de literatura descriptiva y cualitativa de cincuenta fuentes académicas, técnicas y normativas, con el objetivo de exponer la importancia de los datos no estructurados en salud, describir las principales aplicaciones del NLP en este campo y analizar sus desafíos más relevantes, como la precisión de los modelos, la privacidad de los pacientes y la desidentificación de información sensible. Los hallazgos muestran que gran parte de la información clínica más valiosa reside en el texto libre, que las técnicas basadas en transformadores y modelos de lenguaje permiten extraerla con precisión creciente también en español y que la desidentificación y la interoperabilidad son condiciones indispensables para su uso ético. Se concluye que los datos no estructurados, manejados de forma segura, representan una oportunidad real para fortalecer la investigación y con ellas dar un aporte significativo en decisiones clínicas, como también dejar una base clara de cómo avanzar en la gestión de la información de diferentes formatos.
dc.description.abstractThe growth of unstructured clinical information, such as medical records, clinical notes, and narrative reports, has created new challenges for its organization, analysis, and use in the healthcare sector. In this context, natural language processing (NLP) and language models have become key tools for extracting, structuring, and protecting the relevant information contained in medical texts. This popular science article was developed through a descriptive, qualitative literature review of fifty academic, technical, and regulatory sources, with the aim of explaining the importance of unstructured data in healthcare, describing the main applications of NLP in this field, and analyzing its most significant challenges, such as model accuracy, patient privacy, and the de-identification of sensitive information. The findings show that much of the most valuable clinical information resides in free text, that transformer-based techniques and language models can extract it with growing accuracy also in Spanish and that de-identification and interoperability are indispensable conditions for its ethical use. It is concluded that unstructured data, when handled securely, represents a real opportunity to strengthen research, support clinical decision-making, and modernize health information systems.
dc.description.degreelevelPregradospa
dc.description.degreenameIngeniero Informáticospa
dc.description.domainhttp://www.ustatunja.edu.co/investigacion
dc.format.mimetypeapplication/pdf
dc.identifier.citationOspina Leguizamón, D. M. (2026). Información Clínica en Español: Cómo Aprovecharla sin Poner en Riesgo la Privacidad [Trabajo de Grado, Universidad Santo Tomás].Repositorio Institucional
dc.identifier.instnameinstname:Universidad Santo Tomásspa
dc.identifier.reponamereponame:Repositorio Institucional Universidad Santo Tomásspa
dc.identifier.repourlrepourl:https://repository.usta.edu.cospa
dc.identifier.urihttp://hdl.handle.net/11634/74429
dc.language.isospa
dc.publisherUniversidad Santo Tomásspa
dc.publisher.branchCRAI-USTA Tunja
dc.publisher.facultyFacultad de Ingeniería de Sistemasspa
dc.publisher.programIngeniería Informáticaspa
dc.relation.referencesAlsentzer, E., Murphy, J. R., Boag, W., Weng, W.-H., Jin, D., Naumann, T., & McDermott, M. B. A. (2019). Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop (pp. 72–78). Association for Computational Linguistics. https://doi.org/10.18653/v1/W19-1909
dc.relation.referencesBalasubramanian, J. B., Adams, D., Roxanis, I., Berrington de Gonzalez, A., Coulson, P., Almeida, J. S., & García-Closas, M. (2025). Leveraging large language models for structured information extraction from pathology reports. Journal of Pathology Informatics, 19, 100521. https://doi.org/10.1016/j.jpi.2025.100521
dc.relation.referencesBáez, P., Arancibia, A. P., Chaparro, M. I., Bucarey, T., Núñez, F., & Dunstan, J. (2022). Procesamiento de lenguaje natural para texto clínico en español: el caso de las listas de espera en Chile. Revista Médica Clínica Las Condes, 33(6), 576–582. https://doi.org/10.1016/j.rmclc.2022.10.002
dc.relation.referencesBazoge, A., Wargny, M., Constant dit Beaufils, P., Morin, E., Daille, B., Gourraud, P.-A., & Hadjadj, S. (2025). Assessing large language models for acute heart failure classification and information extraction from French clinical notes. Computers in Biology and Medicine, 195, 110609. https://doi.org/10.1016/j.compbiomed.2025.110609
dc.relation.referencesBeltagy, I., Peters, M. E., & Cohan, A. (2020). Longformer: The long-document transformer [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2004.05150
dc.relation.referencesCongreso de Colombia. (1981, 18 de febrero). Ley 23 de 1981. Por la cual se dictan normas en materia de ética médica. https://www.funcionpublica.gov.co/eva/gestornormativo/norma.php?i=68760
dc.relation.referencesCongreso de Colombia. (2012, 17 de octubre). Ley 1581 de 2012. Por la cual se dictan disposiciones generales para la protección de datos personales. https://www.funcionpublica.gov.co/eva/gestornormativo/norma.php?i=49981
dc.relation.referencesCongreso de Colombia. (2020, 31 de enero). Ley 2015 de 2020. Por medio de la cual se crea la historia clínica electrónica interoperable y se dictan otras disposiciones. https://www.minsalud.gov.co/Normatividad_Nuevo/Ley%202015%202020.pdf
dc.relation.referencesDash, S., Shakyawar, S. K., Sharma, M., & Kaushik, S. (2019). Big data in healthcare: Management, analysis and future prospects. Journal of Big Data, 6, Article 54. https://doi.org/10.1186/s40537-019-0217-0
dc.relation.referencesDaskalo, C., Abu-Ashour, W., Tshimula, J. M., Amoei, M., Guadagno, E., & Poenaru, D. (2026). Large language models for electronic health records in pediatric and surgical care: A systematic review. Journal of Pediatric Surgery, 162956. https://doi.org/10.1016/j.jpedsurg.2026.162956
dc.relation.referencesDepartamento Administrativo Nacional de Estadística [DANE]. (2024). Guía para la anonimización de datos estructurados. https://www.dane.gov.co/files/sen/registros-administrativos/guia-anonimizacion-datos2024.pdf
dc.relation.referencesDevlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
dc.relation.referencesDorémus, O., Russon, D., Contrand, B., Guerra-Adames, A., Avalos-Fernandez, M., Gil-Jardiné, C., & Lagarde, E. (2025). Harnessing moderate-sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study. JMIR AI, 4. https://doi.org/10.2196/57828
dc.relation.referencesFaustini, P., McIver, A., Sullivan, R., & Dras, M. (2026). De-identification of clinical data: A systematic review of free text, image and tabular data approaches. International Journal of Medical Informatics, 208, 106225. https://doi.org/10.1016/j.ijmedinf.2025.106225
dc.relation.referencesGarcia-Carmona, A. M., Prieto, M.-L., Puertas, E., & Beunza, J.-J. (2025). Leveraging large language models for accurate retrieval of patient information from medical reports: Systematic evaluation study. JMIR AI, 4. https://doi.org/10.2196/68776
dc.relation.referencesGarcía Subies, G., Barbero Jiménez, Á., & Martínez Fernández, P. (2024). A comparative analysis of Spanish clinical encoder-based models on NER and classification tasks. Journal of the American Medical Informatics Association, 31(9), 2137–2146. https://doi.org/10.1093/jamia/ocae054
dc.relation.referencesGonzález-Castro, L., Cal-González, V. M., Del Fiol, G., & López-Nores, M. (2021). CASIDE: A data model for interoperable cancer survivorship information based on FHIR. Journal of Biomedical Informatics, 124, 103953. https://doi.org/10.1016/j.jbi.2021.103953
dc.relation.referencesGuan, H., Novoa-Laurentiev, J., & Zhou, L. (2025). CD-Tron: Leveraging large clinical language model for early detection of cognitive decline from electronic health records. Journal of Biomedical Informatics, 166, 104830. https://doi.org/10.1016/j.jbi.2025.104830
dc.relation.referencesHossain, E., Rana, R. K., Higgins, N. S., Soar, J., Barua, P. D., Pisani, A. R., & Turner, K. (2023). Natural language processing in electronic health records in relation to healthcare decision-making: A systematic review. Computers in Biology and Medicine, 155, 106649. https://doi.org/10.1016/j.compbiomed.2023.106649
dc.relation.referencesHuang, K., Altosaar, J., & Ranganath, R. (2019). ClinicalBERT: Modeling clinical notes and predicting hospital readmission [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1904.05342
dc.relation.referencesHurtado, L.-F., Marco-Ruiz, L., Segarra, E., Castro-Bleda, M. J., Bustos-Moreno, A., de la Iglesia-Vayá, M., & Vallalta-Rueda, J. F. (2025). Leveraging transformers-based models and linked data for deep phenotyping in radiology. Computer Methods and Programs in Biomedicine, 260, 108567. https://doi.org/10.1016/j.cmpb.2024.108567
dc.relation.referencesJerfy, A., Selden, O., & Balkrishnan, R. (2024). The growing impact of natural language processing in healthcare and public health. Inquiry: The Journal of Health Care Organization, Provision, and Financing, 61, 469580241290095. https://doi.org/10.1177/00469580241290095
dc.relation.referencesJia, J., & Nishi, H. (2025). A flexible two-stage anonymization framework for narrative medical records adapting to various language models. Computers in Biology and Medicine, 195, 110624. https://doi.org/10.1016/j.compbiomed.2025.110624
dc.relation.referencesKim, M. K., Rouphael, C., McMichael, J., Welch, N., & Dasarathy, S. (2024). Challenges in and opportunities for electronic health record-based data analysis and interpretation. Gut and Liver, 18(2), 201–208. https://doi.org/10.5009/gnl230272
dc.relation.referencesKlug, K., Beckh, K., Antweiler, D., Chakraborty, N., Baldini, G., Laue, K., Hosch, R., Nensa, F., Schuler, M., & Giesselbach, S. (2024). From admission to discharge: A systematic review of clinical natural language processing along the patient journey. BMC Medical Informatics and Decision Making, 24, Article 238. https://doi.org/10.1186/s12911-024-02641-w
dc.relation.referencesKoleck, T. A., Dreisbach, C., Bourne, P. E., & Bakken, S. (2019). Natural language processing of symptoms documented in free-text narratives of electronic health records: A systematic review. Journal of the American Medical Informatics Association, 26(4), 364–379. https://doi.org/10.1093/jamia/ocy173
dc.relation.referencesKreimeyer, K., Foster, M., Pandey, A., Arya, N., Halford, G., Jones, S. F., Forshee, R., Walderhaug, M., & Botsis, T. (2017). Natural language processing systems for capturing and standardizing unstructured clinical information: A systematic review. Journal of Biomedical Informatics, 73, 14–29. https://doi.org/10.1016/j.jbi.2017.07.012
dc.relation.referencesLee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240. https://doi.org/10.1093/bioinformatics/btz682
dc.relation.referencesLi, I., Pan, J., Goldwasser, J., Verma, N., Wong, W. P., Nuzumlalı, M. Y., Rosand, B., Li, Y., Zhang, M., Chang, D., Taylor, R. A., Krumholz, H. M., & Radev, D. (2022). Neural natural language processing for unstructured data in electronic health records: A review. Computer Science Review, 46, 100511. https://doi.org/10.1016/j.cosrev.2022.100511
dc.relation.referencesLiu, S., Wang, Y., Wen, A., Wang, L., Hong, N., Shen, F., Bedrick, S., Hersh, W., & Liu, H. (2020). Implementation of a cohort retrieval system for clinical data repositories using the Observational Medical Outcomes Partnership Common Data Model: Proof-of-concept system validation. JMIR Medical Informatics, 8(10). https://doi.org/10.2196/17376
dc.relation.referencesLópez-Úbeda, P., Martín-Noguerol, T., & Luna, A. (2026). Integrating semantic retrieval and chain-of-thought reasoning in small language models for SNOMED CT normalization. International Journal of Medical Informatics, 211, 106340. https://doi.org/10.1016/j.ijmedinf.2026.106340
dc.relation.referencesMata, J., Pachón, V., Manovel, A., Maña, M. J., & de la Villa, M. (2025). Multicriteria optimization of language models for heart failure with preserved ejection fraction symptom detection in Spanish electronic health records: Comparative modeling study. Journal of Medical Internet Research, 27. https://doi.org/10.2196/76433
dc.relation.referencesMinisterio de Ciencia, Tecnología e Innovación. (2022, 9 de julio). Definiciones y conceptos básicos. Gestión de datos de investigación. https://red-documentacion.minciencias.gov.co/Gestion_Datos_Investigacion/gestion-datos
dc.relation.referencesMinisterio de Salud. (1999, 8 de julio). Resolución 1995 de 1999. Por la cual se establecen normas para el manejo de la historia clínica. https://www.minsalud.gov.co/normatividad_nuevo/resoluci%C3%93n%201995%20de%201999.pdf
dc.relation.referencesMinisterio de Salud y Protección Social. (2021, 25 de junio). Resolución 866 de 2021. Por la cual se reglamenta el conjunto de elementos de datos clínicos relevantes para la interoperabilidad de la historia clínica en el país y se dictan otras disposiciones. https://www.minsalud.gov.co/sites/rid/Lists/BibliotecaDigital/RIDE/DE/DIJ/resolucion-866-de-2021.pdf
dc.relation.referencesMinisterio de Salud y Protección Social. (2025, 15 de septiembre). Resolución 1888 de 2025. Por medio de la cual se adopta el Resumen Digital de Atención en Salud (RDA) en el marco de la Interoperabilidad de la Historia Clínica Electrónica (IHCE), se establece el mecanismo para su implementación a nivel nacional y se dictan otras disposiciones. https://www.minsalud.gov.co/Normatividad_Nuevo/Resolucion%20No%201888%20de%202025.pdf
dc.relation.referencesMoreno-Barea, F. J., López-García, G., Mesa, H., Ribelles, N., Alba, E., Jerez, J. M., & Veredas, F. J. (2025). Named entity recognition for de-identifying Spanish electronic health records. Computers in Biology and Medicine, 185, 109576. https://doi.org/10.1016/j.compbiomed.2024.109576
dc.relation.referencesNajafabadipour, M., Zanin, M., Rodríguez-González, A., Torrente, M., Nuñez García, B., Cruz Bermudez, J. L., Provencio, M., & Menasalvas, E. (2020). Reconstructing the patient's natural history from electronic health records. Artificial Intelligence in Medicine, 105, 101860. https://doi.org/10.1016/j.artmed.2020.101860
dc.relation.referencesNegash, B., Katz, A., Neilson, C. J., Moni, M., Nesca, M., Singer, A., & Enns, J. E. (2023). De-identification of free text data containing personal health information: A scoping review of reviews. International Journal of Population Data Science, 8(1), Article 2153. https://doi.org/10.23889/ijpds.v8i1.2153
dc.relation.referencesPinheiro da Silva, D., da Rosa Fröhlich, W., de Mello, B. H., Vieira, R., & Rigo, S. J. (2023). Exploring named entity recognition and relation extraction for ontology and medical records integration. Informatics in Medicine Unlocked, 43, 101381. https://doi.org/10.1016/j.imu.2023.101381
dc.relation.referencesSeinen, T. M., Kors, J. A., van Mulligen, E. M., & Rijnbeek, P. R. (2025). Using structured codes and free-text notes to measure information complementarity in electronic health records: Feasibility and validation study. Journal of Medical Internet Research, 27, e66910. https://doi.org/10.2196/66910
dc.relation.referencesSiepmann, R. M., Baldini, G., Schmidt, C. S., Truhn, D., Müller-Franzes, G. A., Dada, A., Kleesiek, J., Nensa, F., & Hosch, R. (2025). An automated information extraction model for unstructured discharge letters using large language models and GPT-4. Healthcare Analytics, 7, 100378. https://doi.org/10.1016/j.health.2024.100378
dc.relation.referencesTabari, P., Costagliola, G., De Rosa, M., & Boeker, M. (2024). State-of-the-art Fast Healthcare Interoperability Resources (FHIR)–based data model and structure implementations: Systematic scoping review. JMIR Medical Informatics, 12. https://doi.org/10.2196/58445
dc.relation.referencesTaseh, A., Moradian, A. D., Chan, M., Sirls, E., Nazarian, A., Batmanghelich, K., & Bean, J. F. (2025). Performance of natural language processing versus International Classification of Diseases codes in building registries for patients with fall injury: Retrospective analysis. JMIR Medical Informatics, 13. https://doi.org/10.2196/66973
dc.relation.referencesTavabi, N., Pruneski, J., Golchin, S., Singh, M., Sanborn, R., Heyworth, B., Landschaft, A., Kimia, A., & Kiapour, A. (2024). Building large-scale registries from unstructured clinical notes using a low-resource natural language processing pipeline. Artificial Intelligence in Medicine, 151, 102847. https://doi.org/10.1016/j.artmed.2024.102847
dc.relation.referencesVithanage, D., Yu, P., Xie, Q., Xu, H., Wang, L., & Deng, C. (2025). A comprehensive evaluation of large language models for information extraction from unstructured electronic health records in residential aged care. Computers in Biology and Medicine, 197, 111013. https://doi.org/10.1016/j.compbiomed.2025.111013
dc.relation.referencesWoo, B. F. Y., Cato, K., Cho, H., You, S. B., & Song, J. (2026). The use of large language models in clinical documentation: A scoping review. International Journal of Nursing Studies, 176, 105322. https://doi.org/10.1016/j.ijnurstu.2025.105322
dc.relation.referencesZeinali, N., Albashayreh, A., Fan, W., & Gilbertson White, S. (2024). Symptom-BERT: Enhancing cancer symptom detection in EHR clinical notes. Journal of Pain and Symptom Management, 68(2), 190–198.e1. https://doi.org/10.1016/j.jpainsymman.2024.05.015
dc.relation.referencesZhong, X., Li, S., Chen, Z., Ge, L., Yu, D., Wang, S., You, L., & Shang, H. (2025). Considerations for patient privacy of large language models in health care: Scoping review. Journal of Medical Internet Research, 27. https://doi.org/10.2196/76571
dc.relation.referencesZotova, E., Cuadros, M., & Rigau, G. (2026). Generative models for clinical entity linking in Spanish. Array, 30, 100805. https://doi.org/10.1016/j.array.2026.100805
dc.rightsAttribution-NonCommercial-NoDerivs 2.5 Colombiaen
dc.rights.accessrightsinfo:eu-repo/semantics/openAccess
dc.rights.coarhttp://purl.org/coar/access_right/c_abf2
dc.rights.localAbierto (Texto Completo)spa
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/2.5/co/
dc.subject.keywordNatural language processing
dc.subject.keywordUnstructured data
dc.subject.keywordElectronic health record
dc.subject.keywordDe-identification
dc.subject.keywordInteroperability
dc.subject.proposalProcesamiento de lenguaje natural
dc.subject.proposalDatos no estructurados
dc.subject.proposalHistoria clínica electrónica
dc.subject.proposalDesidentificación
dc.subject.proposalInteroperabilidad
dc.titleInformación Clínica en Español: Cómo Aprovecharla sin Poner en Riesgo la Privacidad
dc.typebachelor thesis
dc.type.coarhttp://purl.org/coar/resource_type/c_7a1f
dc.type.coarversionhttp://purl.org/coar/version/c_ab4af688f83e57aa
dc.type.driveinfo:eu-repo/semantics/bachelorThesis
dc.type.localTrabajo de gradospa
dc.type.versioninfo:eu-repo/semantics/acceptedVersion

Archivos

Bloque original

Mostrando 1 - 3 de 3
Cargando...
Miniatura
Nombre:
2026DanielOspina
Tamaño:
925.41 KB
Formato:
Adobe Portable Document Format
Cargando...
Miniatura
Nombre:
Autorización estudiante
Tamaño:
450.51 KB
Formato:
Adobe Portable Document Format
Cargando...
Miniatura
Nombre:
Autorización facultad
Tamaño:
397.46 KB
Formato:
Adobe Portable Document Format

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
license.txt
Tamaño:
807 B
Formato:
Item-specific license agreed upon to submission
Descripción: