Data from: Detecting diabetic retinopathy through machine learning on electronic health record data from an urban, safety net healthcare system
Data files
Sep 21, 2026 version files 5.17 MB
-
DRRisk_external_validation_imputed.csv
500.67 KB
-
DRRisk_external_validation_original.csv
480.47 KB
-
DRRisk_test_imputed.csv
720.72 KB
-
DRRisk_test_original.csv
656.52 KB
-
DRRisk_training_imputed.csv
1.46 MB
-
DRRisk_training_original.csv
1.33 MB
-
README.md
17.03 KB
Abstract
Diabetic retinopathy is a major diabetes-related eye complication and an important cause of preventable vision loss. Timely identification of patients at higher risk can help healthcare systems prioritize screening and follow-up care. This dataset contains de-identified electronic health record data used to develop, test, and externally validate machine learning models for predicting the presence of diabetic retinopathy among patients with diabetes in a large urban public safety net healthcare system. The data include clinical and demographic predictors such as duration of diabetes, hemoglobin A1C, blood urea nitrogen, age, systolic and diastolic blood pressure, hemoglobin, sex, ethnicity, insulin dependence, nephropathy, neuropathy, stroke, and triglycerides.
The dataset is organized into three analytic cohorts. The training and internal test cohorts were derived from 40,631 unique patients with diabetes seen between January 1, 2015, and December 31, 2017. A random 67% of this dataset, consisting of 27,223 cases, was used for model training and cross-validation, while the remaining 33%, consisting of 13,408 cases, was reserved as an internal test set. A temporally distinct cohort of 9,300 patients seen between January 1, 2018, and December 31, 2018, was used as an external validation set. For each cohort, this submission provides two versions of the data: original de-identified data with missing values retained where applicable and imputed data used for machine learning analyses. These files support reproducibility of the associated machine learning analyses and further development of diabetic retinopathy risk prediction approaches using structured EHR data. The data support the associated JAMIA Open publication, the DRRisk web-based prediction tool, and its publication. All protected health information was removed, and the study was approved by the Charles R. Drew University of Medicine and Science Institutional Review Board.
Associated Publications
Ogunyemi OI, Gandhi M, Lee M, Teklehaimanot S, Daskivich LP, Hindman D, Lopez K, Taira RK. Detecting diabetic retinopathy through machine learning on electronic health record data from an urban, safety net healthcare system. JAMIA Open. 2021;4(3):ooab066.
DOI: https://doi.org/10.1093/jamiaopen/ooab066
Gandhi M, Daskivich LP, Ogunyemi OI. DRRisk: A Web-based tool to Assess the Risk of Diabetic Retinopathy through Machine Learning on Electronic Health Records. AMIA Annual Symposium Proceedings. 2022;2022:452-460.
PMID: 37128428; PMCID: PMC10148369.
https://pmc.ncbi.nlm.nih.gov/articles/PMC10148369/
Description
This dataset supports the development, testing, and external validation of machine learning models for predicting diabetic retinopathy (DR) from structured electronic health record (EHR) data. The data were derived from records of adult patients with diabetes treated within a large urban public safety net system.
The dataset is organized into three analytic cohorts:
- Training set — used to build, train, tune, and select machine learning models.
- Test set — an internal evaluation cohort randomly held out from the 2015-2017 dataset and not used for model training.
- External validation set — a temporally distinct 2018 cohort not used to train or internally test the models.
For each cohort, two versions of the data are provided:
- Original de-identified data — cleaned analytic data with missing values retained where applicable.
- Imputed data — corresponding data after missing values were imputed for machine learning analyses.
The demographic transformations described in this README were applied to both the original de-identified and imputed files. Direct identifiers and protected health information (PHI) were removed before public sharing.
The study was approved by the Charles R. Drew University of Medicine and Science Institutional Review Board under approval number 16-10-2491-03.
Data Collection
- Source: Structured EHR data from a large urban public healthcare system.
- Study population: Adult patients with Type 1 or Type 2 diabetes who received an eye examination.
- Training and test timeframe: January 1, 2015 through December 31, 2017.
- External validation timeframe: January 1, 2018 through December 31, 2018.
Cohort Sizes
| Cohort | Timeframe | Description | Number of cases |
|---|---|---|---|
| Training set | 2015-2017 | Random 67% of the 2015-2017 dataset used for model development | 27,223 |
| Test set | 2015-2017 | Random 33% of the 2015-2017 dataset held out for internal testing | 13,408 |
| External validation set | 2018 | Temporally distinct validation cohort | 9,300 |
The combined 2015-2017 dataset included 40,631 unique patients with diabetes. The external validation dataset included an additional distinct set of 9,300 patients with diabetes seen during 2018.
File Inventory
The Dryad submission includes two versions of each analytic cohort: original de-identified data and imputed data.
| Data version | Filename | Cohort | Rows | Description |
|---|---|---|---|---|
| Original de-identified data | DRRisk_training_original.csv |
Training | 27,223 | De-identified training cohort with missing values retained where applicable |
| Original de-identified data | DRRisk_test_original.csv |
Test | 13,408 | De-identified internal test cohort with missing values retained where applicable |
| Original de-identified data | DRRisk_external_validation_original.csv |
External validation | 9,300 | De-identified external validation cohort with missing values retained where applicable |
| Imputed data | DRRisk_training_imputed.csv |
Training | 27,223 | Training cohort after missing value imputation |
| Imputed data | DRRisk_test_imputed.csv |
Test | 13,408 | Test cohort after applying the imputation approach |
| Imputed data | DRRisk_external_validation_imputed.csv |
External validation | 9,300 | External validation cohort after applying the imputation approach |
All files are comma-separated values (.csv) files.
Variable Description
The DRRisk prediction framework uses the following predictors to assess current diabetic retinopathy risk.
| Variable | Display label | Description | Unit or coding |
|---|---|---|---|
DURATION_OF_DIABETES_YRS |
Duration of Diabetes | Duration of diabetes at the time of assessment | Years |
HEMOGLOBIN_A1C |
Hemoglobin A1C | Most recent hemoglobin A1C value | Percent (%) |
BLOOD_UREA_NITROGEN |
Blood Urea Nitrogen | Most recent blood urea nitrogen laboratory value | mg/dL |
AGE |
Age Group | Patient age grouped into 10-year ranges | 18-29, 30-39, 40-49, 50-59, 60-69, 70-79, or 80+ |
SYSTOLIC_BLOOD_PRESSURE |
Systolic Blood Pressure | Most recent systolic blood pressure value | mmHg |
DIASTOLIC_BLOOD_PRESSURE |
Diastolic Blood Pressure | Most recent diastolic blood pressure value | mmHg |
HEMOGLOBIN |
Hemoglobin | Most recent hemoglobin laboratory value | g/dL |
SEX |
Sex | Binary-coded patient sex | 0 or 1; correspondence with the original categories is not publicly disclosed |
ETHNICITY |
Ethnicity | Binary-coded patient ethnicity | 0 or 1; correspondence with the original categories is not publicly disclosed |
INSULIN_DEPENDENCE |
Insulin Dependence | Whether the patient was insulin dependent | 1 or 2; correspondence with the original categories is not publicly disclosed |
NEPHROPATHY |
Nephropathy | History or diagnosis of nephropathy | 1 or 2; correspondence with the original categories is not publicly disclosed |
NEUROPATHY |
Neuropathy | History or diagnosis of neuropathy | 1 or 2; correspondence with the original categories is not publicly disclosed |
STROKE |
Stroke | History or diagnosis of stroke | 1 or 2; correspondence with the original categories is not publicly disclosed |
TRIGLYCERIDES |
Triglycerides | Most recent triglyceride laboratory value | mg/dL |
RETINOPATHY |
Diabetic Retinopathy | Outcome indicating the presence or absence of diabetic retinopathy | Y = Present, N = Absent |
Additional information about the DRRisk predictors is available at:
https://drandml.cdrewu.edu/
Missing Data
The original de-identified files intentionally retain missing values from the source EHR-derived analytic data. These files are provided so that users can examine missingness patterns and apply alternative preprocessing or imputation methods.
Missing values are represented as NA or blank cells, depending on the source file.
The corresponding imputed files were created using k-nearest neighbor imputation with k = 9, as described in the associated JAMIA Open publication.
Data Preparation and Cleaning
- The data were derived from retrospective electronic health records.
- Patients younger than 18 years were excluded.
- Direct identifiers and PHI were removed before public sharing.
- The correspondence between the binary codes and the original categories is not publicly disclosed.
- Missing values were retained in the original de-identified files.
- Missing values were imputed using k-nearest neighbor imputation with
k = 9for the imputed files.
De-identification
To protect patient privacy and comply with data-sharing requirements, the following de-identification measures were applied.
1. Exclusion of Direct Identifiers
- No direct patient identifiers, including names, addresses, medical record numbers, telephone numbers, email addresses, or exact patient-level dates, are included in the dataset.
- All demographic and clinical information is provided in a de-identified format.
2. Temporal De-identification
- Exact patient-level dates are not included in the shared files.
- Time-related information needed for analysis is represented using derived measures, such as duration of diabetes in years, rather than exact diagnosis or encounter dates.
3. Age Censoring
- Patient ages were categorized into predefined 10-year age groups:
18-29,30-39,40-49,50-59,60-69,70-79, and80+. - All patients aged 80 years or older were grouped into a single
80+category. - This approach reduces the precision of age information and limits disclosure based on exact age or age outliers.
4. Sex Encoding
- Sex is recorded as binary values of
0or1without defining which value corresponds to each original sex category. - This approach reduces the interpretability of individual demographic values while allowing the variable to remain available for analytical use.
5. Ethnicity Encoding
- Ethnicity is recorded as binary values of
0or1without defining which value corresponds to each original ethnicity category. - The correspondence between the numerical values and the original categories is not included in the public dataset or README.
- This approach provides an additional privacy measure while retaining the variable for analytical use.
6. Clinical Data
- Clinical predictors, including laboratory measurements and medical-condition indicators, are provided without direct patient identifiers.
- The diabetic retinopathy outcome is retained as a binary
Y/Nvariable because it is the primary outcome used in the associated machine learning analyses.
7. Comorbid Conditions Encoding
- Comorbid conditions - Insulin dependence, Stroke, Neuropathy, and Nephropathy are encoded as binary values of
1or2without defining which value corresponds to the original category. - This approach provides an additional privacy measure while retaining the variable for analytical use.
8. Data-Source Authorization
- The data-providing healthcare system authorized sharing of the de-identified dataset.
- The shared files were prepared to exclude direct identifiers and PHI before public release.
Software and Code
Analyses in the associated publications used both R and Python.
- R: Used for machine learning workflows, training, cross-validation, preprocessing, and model assessment.
- R packages:
caret,VIM, and supporting packages. - Python: Used for deep neural network modeling.
- Python packages:
scikit-learn,pandas,numpy, and supporting packages. - Machine learning methods: Random forest, support vector machine, extreme gradient boosting, ensemble models, and deep neural networks.
- Sampling methods: Majority class undersampling and synthetic minority over-sampling technique (SMOTE).
- Model evaluation: 10-fold cross-validation on the training set, internal testing on the held-out test set, and external validation using the 2018 validation cohort.
Ethics, Participant Consent, and Data Sharing
This retrospective study was approved by the Charles R. Drew University of Medicine and Science Institutional Review Board under approval number 16-10-2491-03.
The Institutional Review Board determined that the study involved no more than minimal risk and granted a waiver of informed consent because:
- The research involved no more than minimal risk to participants.
- The waiver did not adversely affect the rights or welfare of participants.
- The information had already been collected as part of routine clinical care.
Participants did not provide explicit consent for public release of their de-identified data. The data-providing healthcare system authorized sharing of the de-identified dataset.
Before public sharing, direct identifiers and PHI were removed, exact age was generalized into ranges, and selected demographic variables were converted into binary-coded values.
No direct identifiers or PHI are included in the Dryad data files.
Funding
This work was funded by the National Library of Medicine under grant 1 R01 LM012309.
Data acquisition was supported by the UCLA and USC Clinical and Translational Science Institutes under National Center for Advancing Translational Sciences grants UL1TR001881 and UL1TR001855.
The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Citation
If you use these data, please cite the Dryad dataset and the associated publications.
Ogunyemi OI, Gandhi M, Lee M, Teklehaimanot S, Daskivich LP, Hindman D, Lopez K, Taira RK. Detecting diabetic retinopathy through machine learning on electronic health record data from an urban, safety net healthcare system. JAMIA Open. 2021;4(3):ooab066.
DOI: https://doi.org/10.1093/jamiaopen/ooab066
Gandhi M, Daskivich LP, Ogunyemi OI. DRRisk: A Web-based tool to Assess the Risk of Diabetic Retinopathy through Machine Learning on Electronic Health Records. AMIA Annual Symposium Proceedings. 2022;2022:452-460.
PMID: 37128428; PMCID: PMC10148369.
Contact
For questions about this dataset, contact:
Omolola Ogunyemi, PhD
Center for Biomedical Informatics
Charles R. Drew University of Medicine and Science
Email: lolaogunyemi@cdrewu.edu
Lauren Daskivich, MD
Los Angeles County Department of Health Services
Email: lpdaskivich@dhs.lacounty.gov
Human subjects data
This retrospective study was approved by the Charles R. Drew University of Medicine and Science Institutional Review Board under approval number 16-10-2491-03.
The Institutional Review Board determined that the study involved no more than minimal risk and granted a waiver of informed consent because:
- The research involved no more than minimal risk to participants.
- The waiver did not adversely affect the rights or welfare of participants.
- The information had already been collected as part of routine clinical care.
Participants did not provide explicit consent for public release of their de-identified data. The data-providing healthcare system authorized sharing of the de-identified dataset.
Before public sharing, direct identifiers and PHI were removed, exact age was generalized into ranges, and selected demographic variables were converted into binary coded values.
No direct identifiers or PHI are included in the Dryad data files.
