Información Clínica en Español: Cómo Aprovecharla sin Poner en Riesgo la Privacidad
| dc.contributor.advisor | Piracoca Arcos, Jhon Alexis | |
| dc.contributor.author | Ospina Leguizamón, Daniel Mauricio | |
| dc.contributor.corporatename | Universidad Santo Tomás | |
| dc.contributor.cvlac | https://scienti.minciencias.gov.co/cvlac/visualizador/generarCurriculoCv.do?cod_rh=0002450348 | |
| dc.date.accessioned | 2026-10-01T19:46:35Z | |
| dc.date.available | 2026-10-01T19:46:35Z | |
| dc.date.issued | 2026-09-30 | |
| dc.description | El crecimiento de la información clínica en formato no estructurado, como las historias clínicas, las notas médicas y los reportes narrativos, deja una gran tarea al tener como principal reto su análisis y aprovechamiento. En este caso, el procesamiento de lenguaje natural (NLP) y los modelos de lenguaje pueden comprenderse como herramientas clave y de gran repercusión para poder extraer y tener una comprensión más directa y exacta de su contenido sin dejar de lado la protección de información relevante expuesta en los mismos. Se desarrolló este artículo de divulgación mediante una revisión de literatura descriptiva y cualitativa de cincuenta fuentes académicas, técnicas y normativas, con el objetivo de exponer la importancia de los datos no estructurados en salud, describir las principales aplicaciones del NLP en este campo y analizar sus desafíos más relevantes, como la precisión de los modelos, la privacidad de los pacientes y la desidentificación de información sensible. Los hallazgos muestran que gran parte de la información clínica más valiosa reside en el texto libre, que las técnicas basadas en transformadores y modelos de lenguaje permiten extraerla con precisión creciente también en español y que la desidentificación y la interoperabilidad son condiciones indispensables para su uso ético. Se concluye que los datos no estructurados, manejados de forma segura, representan una oportunidad real para fortalecer la investigación y con ellas dar un aporte significativo en decisiones clínicas, como también dejar una base clara de cómo avanzar en la gestión de la información de diferentes formatos. | |
| dc.description.abstract | The growth of unstructured clinical information, such as medical records, clinical notes, and narrative reports, has created new challenges for its organization, analysis, and use in the healthcare sector. In this context, natural language processing (NLP) and language models have become key tools for extracting, structuring, and protecting the relevant information contained in medical texts. This popular science article was developed through a descriptive, qualitative literature review of fifty academic, technical, and regulatory sources, with the aim of explaining the importance of unstructured data in healthcare, describing the main applications of NLP in this field, and analyzing its most significant challenges, such as model accuracy, patient privacy, and the de-identification of sensitive information. The findings show that much of the most valuable clinical information resides in free text, that transformer-based techniques and language models can extract it with growing accuracy also in Spanish and that de-identification and interoperability are indispensable conditions for its ethical use. It is concluded that unstructured data, when handled securely, represents a real opportunity to strengthen research, support clinical decision-making, and modernize health information systems. | |
| dc.description.degreelevel | Pregrado | spa |
| dc.description.degreename | Ingeniero Informático | spa |
| dc.description.domain | http://www.ustatunja.edu.co/investigacion | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.citation | Ospina Leguizamón, D. M. (2026). Información Clínica en Español: Cómo Aprovecharla sin Poner en Riesgo la Privacidad [Trabajo de Grado, Universidad Santo Tomás].Repositorio Institucional | |
| dc.identifier.instname | instname:Universidad Santo Tomás | spa |
| dc.identifier.reponame | reponame:Repositorio Institucional Universidad Santo Tomás | spa |
| dc.identifier.repourl | repourl:https://repository.usta.edu.co | spa |
| dc.identifier.uri | http://hdl.handle.net/11634/74429 | |
| dc.language.iso | spa | |
| dc.publisher | Universidad Santo Tomás | spa |
| dc.publisher.branch | CRAI-USTA Tunja | |
| dc.publisher.faculty | Facultad de Ingeniería de Sistemas | spa |
| dc.publisher.program | Ingeniería Informática | spa |
| dc.relation.references | Alsentzer, E., Murphy, J. R., Boag, W., Weng, W.-H., Jin, D., Naumann, T., & McDermott, M. B. A. (2019). Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop (pp. 72–78). Association for Computational Linguistics. https://doi.org/10.18653/v1/W19-1909 | |
| dc.relation.references | Balasubramanian, J. B., Adams, D., Roxanis, I., Berrington de Gonzalez, A., Coulson, P., Almeida, J. S., & García-Closas, M. (2025). Leveraging large language models for structured information extraction from pathology reports. Journal of Pathology Informatics, 19, 100521. https://doi.org/10.1016/j.jpi.2025.100521 | |
| dc.relation.references | Báez, P., Arancibia, A. P., Chaparro, M. I., Bucarey, T., Núñez, F., & Dunstan, J. (2022). Procesamiento de lenguaje natural para texto clínico en español: el caso de las listas de espera en Chile. Revista Médica Clínica Las Condes, 33(6), 576–582. https://doi.org/10.1016/j.rmclc.2022.10.002 | |
| dc.relation.references | Bazoge, A., Wargny, M., Constant dit Beaufils, P., Morin, E., Daille, B., Gourraud, P.-A., & Hadjadj, S. (2025). Assessing large language models for acute heart failure classification and information extraction from French clinical notes. Computers in Biology and Medicine, 195, 110609. https://doi.org/10.1016/j.compbiomed.2025.110609 | |
| dc.relation.references | Beltagy, I., Peters, M. E., & Cohan, A. (2020). Longformer: The long-document transformer [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2004.05150 | |
| dc.relation.references | Congreso de Colombia. (1981, 18 de febrero). Ley 23 de 1981. Por la cual se dictan normas en materia de ética médica. https://www.funcionpublica.gov.co/eva/gestornormativo/norma.php?i=68760 | |
| dc.relation.references | Congreso de Colombia. (2012, 17 de octubre). Ley 1581 de 2012. Por la cual se dictan disposiciones generales para la protección de datos personales. https://www.funcionpublica.gov.co/eva/gestornormativo/norma.php?i=49981 | |
| dc.relation.references | Congreso de Colombia. (2020, 31 de enero). Ley 2015 de 2020. Por medio de la cual se crea la historia clínica electrónica interoperable y se dictan otras disposiciones. https://www.minsalud.gov.co/Normatividad_Nuevo/Ley%202015%202020.pdf | |
| dc.relation.references | Dash, S., Shakyawar, S. K., Sharma, M., & Kaushik, S. (2019). Big data in healthcare: Management, analysis and future prospects. Journal of Big Data, 6, Article 54. https://doi.org/10.1186/s40537-019-0217-0 | |
| dc.relation.references | Daskalo, C., Abu-Ashour, W., Tshimula, J. M., Amoei, M., Guadagno, E., & Poenaru, D. (2026). Large language models for electronic health records in pediatric and surgical care: A systematic review. Journal of Pediatric Surgery, 162956. https://doi.org/10.1016/j.jpedsurg.2026.162956 | |
| dc.relation.references | Departamento Administrativo Nacional de Estadística [DANE]. (2024). Guía para la anonimización de datos estructurados. https://www.dane.gov.co/files/sen/registros-administrativos/guia-anonimizacion-datos2024.pdf | |
| dc.relation.references | Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423 | |
| dc.relation.references | Dorémus, O., Russon, D., Contrand, B., Guerra-Adames, A., Avalos-Fernandez, M., Gil-Jardiné, C., & Lagarde, E. (2025). Harnessing moderate-sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study. JMIR AI, 4. https://doi.org/10.2196/57828 | |
| dc.relation.references | Faustini, P., McIver, A., Sullivan, R., & Dras, M. (2026). De-identification of clinical data: A systematic review of free text, image and tabular data approaches. International Journal of Medical Informatics, 208, 106225. https://doi.org/10.1016/j.ijmedinf.2025.106225 | |
| dc.relation.references | Garcia-Carmona, A. M., Prieto, M.-L., Puertas, E., & Beunza, J.-J. (2025). Leveraging large language models for accurate retrieval of patient information from medical reports: Systematic evaluation study. JMIR AI, 4. https://doi.org/10.2196/68776 | |
| dc.relation.references | García Subies, G., Barbero Jiménez, Á., & Martínez Fernández, P. (2024). A comparative analysis of Spanish clinical encoder-based models on NER and classification tasks. Journal of the American Medical Informatics Association, 31(9), 2137–2146. https://doi.org/10.1093/jamia/ocae054 | |
| dc.relation.references | González-Castro, L., Cal-González, V. M., Del Fiol, G., & López-Nores, M. (2021). CASIDE: A data model for interoperable cancer survivorship information based on FHIR. Journal of Biomedical Informatics, 124, 103953. https://doi.org/10.1016/j.jbi.2021.103953 | |
| dc.relation.references | Guan, H., Novoa-Laurentiev, J., & Zhou, L. (2025). CD-Tron: Leveraging large clinical language model for early detection of cognitive decline from electronic health records. Journal of Biomedical Informatics, 166, 104830. https://doi.org/10.1016/j.jbi.2025.104830 | |
| dc.relation.references | Hossain, E., Rana, R. K., Higgins, N. S., Soar, J., Barua, P. D., Pisani, A. R., & Turner, K. (2023). Natural language processing in electronic health records in relation to healthcare decision-making: A systematic review. Computers in Biology and Medicine, 155, 106649. https://doi.org/10.1016/j.compbiomed.2023.106649 | |
| dc.relation.references | Huang, K., Altosaar, J., & Ranganath, R. (2019). ClinicalBERT: Modeling clinical notes and predicting hospital readmission [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1904.05342 | |
| dc.relation.references | Hurtado, L.-F., Marco-Ruiz, L., Segarra, E., Castro-Bleda, M. J., Bustos-Moreno, A., de la Iglesia-Vayá, M., & Vallalta-Rueda, J. F. (2025). Leveraging transformers-based models and linked data for deep phenotyping in radiology. Computer Methods and Programs in Biomedicine, 260, 108567. https://doi.org/10.1016/j.cmpb.2024.108567 | |
| dc.relation.references | Jerfy, A., Selden, O., & Balkrishnan, R. (2024). The growing impact of natural language processing in healthcare and public health. Inquiry: The Journal of Health Care Organization, Provision, and Financing, 61, 469580241290095. https://doi.org/10.1177/00469580241290095 | |
| dc.relation.references | Jia, J., & Nishi, H. (2025). A flexible two-stage anonymization framework for narrative medical records adapting to various language models. Computers in Biology and Medicine, 195, 110624. https://doi.org/10.1016/j.compbiomed.2025.110624 | |
| dc.relation.references | Kim, M. K., Rouphael, C., McMichael, J., Welch, N., & Dasarathy, S. (2024). Challenges in and opportunities for electronic health record-based data analysis and interpretation. Gut and Liver, 18(2), 201–208. https://doi.org/10.5009/gnl230272 | |
| dc.relation.references | Klug, K., Beckh, K., Antweiler, D., Chakraborty, N., Baldini, G., Laue, K., Hosch, R., Nensa, F., Schuler, M., & Giesselbach, S. (2024). From admission to discharge: A systematic review of clinical natural language processing along the patient journey. BMC Medical Informatics and Decision Making, 24, Article 238. https://doi.org/10.1186/s12911-024-02641-w | |
| dc.relation.references | Koleck, T. A., Dreisbach, C., Bourne, P. E., & Bakken, S. (2019). Natural language processing of symptoms documented in free-text narratives of electronic health records: A systematic review. Journal of the American Medical Informatics Association, 26(4), 364–379. https://doi.org/10.1093/jamia/ocy173 | |
| dc.relation.references | Kreimeyer, K., Foster, M., Pandey, A., Arya, N., Halford, G., Jones, S. F., Forshee, R., Walderhaug, M., & Botsis, T. (2017). Natural language processing systems for capturing and standardizing unstructured clinical information: A systematic review. Journal of Biomedical Informatics, 73, 14–29. https://doi.org/10.1016/j.jbi.2017.07.012 | |
| dc.relation.references | Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240. https://doi.org/10.1093/bioinformatics/btz682 | |
| dc.relation.references | Li, I., Pan, J., Goldwasser, J., Verma, N., Wong, W. P., Nuzumlalı, M. Y., Rosand, B., Li, Y., Zhang, M., Chang, D., Taylor, R. A., Krumholz, H. M., & Radev, D. (2022). Neural natural language processing for unstructured data in electronic health records: A review. Computer Science Review, 46, 100511. https://doi.org/10.1016/j.cosrev.2022.100511 | |
| dc.relation.references | Liu, S., Wang, Y., Wen, A., Wang, L., Hong, N., Shen, F., Bedrick, S., Hersh, W., & Liu, H. (2020). Implementation of a cohort retrieval system for clinical data repositories using the Observational Medical Outcomes Partnership Common Data Model: Proof-of-concept system validation. JMIR Medical Informatics, 8(10). https://doi.org/10.2196/17376 | |
| dc.relation.references | López-Úbeda, P., Martín-Noguerol, T., & Luna, A. (2026). Integrating semantic retrieval and chain-of-thought reasoning in small language models for SNOMED CT normalization. International Journal of Medical Informatics, 211, 106340. https://doi.org/10.1016/j.ijmedinf.2026.106340 | |
| dc.relation.references | Mata, J., Pachón, V., Manovel, A., Maña, M. J., & de la Villa, M. (2025). Multicriteria optimization of language models for heart failure with preserved ejection fraction symptom detection in Spanish electronic health records: Comparative modeling study. Journal of Medical Internet Research, 27. https://doi.org/10.2196/76433 | |
| dc.relation.references | Ministerio de Ciencia, Tecnología e Innovación. (2022, 9 de julio). Definiciones y conceptos básicos. Gestión de datos de investigación. https://red-documentacion.minciencias.gov.co/Gestion_Datos_Investigacion/gestion-datos | |
| dc.relation.references | Ministerio de Salud. (1999, 8 de julio). Resolución 1995 de 1999. Por la cual se establecen normas para el manejo de la historia clínica. https://www.minsalud.gov.co/normatividad_nuevo/resoluci%C3%93n%201995%20de%201999.pdf | |
| dc.relation.references | Ministerio de Salud y Protección Social. (2021, 25 de junio). Resolución 866 de 2021. Por la cual se reglamenta el conjunto de elementos de datos clínicos relevantes para la interoperabilidad de la historia clínica en el país y se dictan otras disposiciones. https://www.minsalud.gov.co/sites/rid/Lists/BibliotecaDigital/RIDE/DE/DIJ/resolucion-866-de-2021.pdf | |
| dc.relation.references | Ministerio de Salud y Protección Social. (2025, 15 de septiembre). Resolución 1888 de 2025. Por medio de la cual se adopta el Resumen Digital de Atención en Salud (RDA) en el marco de la Interoperabilidad de la Historia Clínica Electrónica (IHCE), se establece el mecanismo para su implementación a nivel nacional y se dictan otras disposiciones. https://www.minsalud.gov.co/Normatividad_Nuevo/Resolucion%20No%201888%20de%202025.pdf | |
| dc.relation.references | Moreno-Barea, F. J., López-García, G., Mesa, H., Ribelles, N., Alba, E., Jerez, J. M., & Veredas, F. J. (2025). Named entity recognition for de-identifying Spanish electronic health records. Computers in Biology and Medicine, 185, 109576. https://doi.org/10.1016/j.compbiomed.2024.109576 | |
| dc.relation.references | Najafabadipour, M., Zanin, M., Rodríguez-González, A., Torrente, M., Nuñez García, B., Cruz Bermudez, J. L., Provencio, M., & Menasalvas, E. (2020). Reconstructing the patient's natural history from electronic health records. Artificial Intelligence in Medicine, 105, 101860. https://doi.org/10.1016/j.artmed.2020.101860 | |
| dc.relation.references | Negash, B., Katz, A., Neilson, C. J., Moni, M., Nesca, M., Singer, A., & Enns, J. E. (2023). De-identification of free text data containing personal health information: A scoping review of reviews. International Journal of Population Data Science, 8(1), Article 2153. https://doi.org/10.23889/ijpds.v8i1.2153 | |
| dc.relation.references | Pinheiro da Silva, D., da Rosa Fröhlich, W., de Mello, B. H., Vieira, R., & Rigo, S. J. (2023). Exploring named entity recognition and relation extraction for ontology and medical records integration. Informatics in Medicine Unlocked, 43, 101381. https://doi.org/10.1016/j.imu.2023.101381 | |
| dc.relation.references | Seinen, T. M., Kors, J. A., van Mulligen, E. M., & Rijnbeek, P. R. (2025). Using structured codes and free-text notes to measure information complementarity in electronic health records: Feasibility and validation study. Journal of Medical Internet Research, 27, e66910. https://doi.org/10.2196/66910 | |
| dc.relation.references | Siepmann, R. M., Baldini, G., Schmidt, C. S., Truhn, D., Müller-Franzes, G. A., Dada, A., Kleesiek, J., Nensa, F., & Hosch, R. (2025). An automated information extraction model for unstructured discharge letters using large language models and GPT-4. Healthcare Analytics, 7, 100378. https://doi.org/10.1016/j.health.2024.100378 | |
| dc.relation.references | Tabari, P., Costagliola, G., De Rosa, M., & Boeker, M. (2024). State-of-the-art Fast Healthcare Interoperability Resources (FHIR)–based data model and structure implementations: Systematic scoping review. JMIR Medical Informatics, 12. https://doi.org/10.2196/58445 | |
| dc.relation.references | Taseh, A., Moradian, A. D., Chan, M., Sirls, E., Nazarian, A., Batmanghelich, K., & Bean, J. F. (2025). Performance of natural language processing versus International Classification of Diseases codes in building registries for patients with fall injury: Retrospective analysis. JMIR Medical Informatics, 13. https://doi.org/10.2196/66973 | |
| dc.relation.references | Tavabi, N., Pruneski, J., Golchin, S., Singh, M., Sanborn, R., Heyworth, B., Landschaft, A., Kimia, A., & Kiapour, A. (2024). Building large-scale registries from unstructured clinical notes using a low-resource natural language processing pipeline. Artificial Intelligence in Medicine, 151, 102847. https://doi.org/10.1016/j.artmed.2024.102847 | |
| dc.relation.references | Vithanage, D., Yu, P., Xie, Q., Xu, H., Wang, L., & Deng, C. (2025). A comprehensive evaluation of large language models for information extraction from unstructured electronic health records in residential aged care. Computers in Biology and Medicine, 197, 111013. https://doi.org/10.1016/j.compbiomed.2025.111013 | |
| dc.relation.references | Woo, B. F. Y., Cato, K., Cho, H., You, S. B., & Song, J. (2026). The use of large language models in clinical documentation: A scoping review. International Journal of Nursing Studies, 176, 105322. https://doi.org/10.1016/j.ijnurstu.2025.105322 | |
| dc.relation.references | Zeinali, N., Albashayreh, A., Fan, W., & Gilbertson White, S. (2024). Symptom-BERT: Enhancing cancer symptom detection in EHR clinical notes. Journal of Pain and Symptom Management, 68(2), 190–198.e1. https://doi.org/10.1016/j.jpainsymman.2024.05.015 | |
| dc.relation.references | Zhong, X., Li, S., Chen, Z., Ge, L., Yu, D., Wang, S., You, L., & Shang, H. (2025). Considerations for patient privacy of large language models in health care: Scoping review. Journal of Medical Internet Research, 27. https://doi.org/10.2196/76571 | |
| dc.relation.references | Zotova, E., Cuadros, M., & Rigau, G. (2026). Generative models for clinical entity linking in Spanish. Array, 30, 100805. https://doi.org/10.1016/j.array.2026.100805 | |
| dc.rights | Attribution-NonCommercial-NoDerivs 2.5 Colombia | en |
| dc.rights.accessrights | info:eu-repo/semantics/openAccess | |
| dc.rights.coar | http://purl.org/coar/access_right/c_abf2 | |
| dc.rights.local | Abierto (Texto Completo) | spa |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/2.5/co/ | |
| dc.subject.keyword | Natural language processing | |
| dc.subject.keyword | Unstructured data | |
| dc.subject.keyword | Electronic health record | |
| dc.subject.keyword | De-identification | |
| dc.subject.keyword | Interoperability | |
| dc.subject.proposal | Procesamiento de lenguaje natural | |
| dc.subject.proposal | Datos no estructurados | |
| dc.subject.proposal | Historia clínica electrónica | |
| dc.subject.proposal | Desidentificación | |
| dc.subject.proposal | Interoperabilidad | |
| dc.title | Información Clínica en Español: Cómo Aprovecharla sin Poner en Riesgo la Privacidad | |
| dc.type | bachelor thesis | |
| dc.type.coar | http://purl.org/coar/resource_type/c_7a1f | |
| dc.type.coarversion | http://purl.org/coar/version/c_ab4af688f83e57aa | |
| dc.type.drive | info:eu-repo/semantics/bachelorThesis | |
| dc.type.local | Trabajo de grado | spa |
| dc.type.version | info:eu-repo/semantics/acceptedVersion |
Archivos
Bloque original
1 - 3 de 3
Cargando...
- Nombre:
- Autorización estudiante
- Tamaño:
- 450.51 KB
- Formato:
- Adobe Portable Document Format
Cargando...
- Nombre:
- Autorización facultad
- Tamaño:
- 397.46 KB
- Formato:
- Adobe Portable Document Format
Bloque de licencias
1 - 1 de 1
Cargando...
- Nombre:
- license.txt
- Tamaño:
- 807 B
- Formato:
- Item-specific license agreed upon to submission
- Descripción:

