De-identification of patient notes with recurrent neural networks

255Citations
Citations of this article
380Readers
Mendeley users who have this article in their library.

Abstract

Objective: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the confidentiality of patients. In the United States, the Health Insurance Portability and Accountability Act (HIPAA) defines 18 types of protected health information that needs to be removed to de-identify patient notes. Manual de-identification is impractical given the size of electronic health record databases, the limited number of researchers with access to non-de-identified notes, and the frequent mistakes of human annotators. A reliable automated de-identification system would consequently be of high value. Materials and Methods: We introduce the first de-identification system based on artificial neural networks (ANNs), which requires no handcrafted features or rules, unlike existing systems. We compare the performance of the system with state-of-the-art systems on two datasets: the i2b2 2014 de-identification challenge dataset, which is the largest publicly available de-identification dataset, and the MIMIC de-identification dataset, which we assembled and is twice as large as the i2b2 2014 dataset. Results: Our ANN model outperforms the state-of-the-art systems. It yields an F1-score of 97.85 on the i2b2 2014 dataset, with a recall of 97.38 and a precision of 98.32, and an F1-score of 99.23 on the MIMIC deidentification dataset, with a recall of 99.25 and a precision of 99.21. Conclusion: Our findings support the use of ANNs for de-identification of patient notes, as they show better performance than previously published systems while requiring no manual feature engineering.

References Powered by Scopus

78478Citations
26334Readers
Get full text

GloVe: Global vectors for word representation

27202Citations
11238Readers

Convolutional neural networks for sentence classification

8128Citations
7491Readers

Cited by Powered by Scopus

1931Citations
3057Readers

This article is free to access.

Get full text

This article is free to access.

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Cite

CITATION STYLE

APA

Dernoncourt, F., Lee, J. Y., Uzuner, O., & Szolovits, P. (2017). De-identification of patient notes with recurrent neural networks. Journal of the American Medical Informatics Association, 24(3), 596–606. https://doi.org/10.1093/jamia/ocw156

Readers over time

‘16‘17‘18‘19‘20‘21‘22‘23‘24‘250306090120

Readers' Seniority

Tooltip

PhD / Post grad / Masters / Doc 144

59%

Researcher 71

29%

Professor / Associate Prof. 20

8%

Lecturer / Post doc 9

4%

Readers' Discipline

Tooltip

Computer Science 161

69%

Medicine and Dentistry 34

15%

Engineering 30

13%

Business, Management and Accounting 9

4%

Save time finding and organizing research with Mendeley

Sign up for free
0