Data from: Blood-derived DNA methylation biomarkers predict diabetic kidney disease
Data files
Jul 24, 2026 version files 801.10 MB
-
README.md
5.16 KB
-
Sample_Information_Manuscript.tsv
7.70 KB
-
Steno_T1D_MBDseq_Count_Matrix_Manuscript.tsv
801.09 MB
Abstract
Diabetic kidney disease (DKD) is a leading cause of end-stage kidney disease in people with type 1 diabetes (T1D), yet current markers based on albuminuria and estimated glomerular filtration rate (eGFR) lack precision for early prognostication. We aimed to identify blood-derived DNA methylation biomarkers that predict DKD progression and to determine whether these systemic signals reflect epigenomic remodeling in renal cells.
We studied the PROFIL cohort, performing high-depth genome-wide methylation sequencing of leukocyte DNA from 101 adults with T1D and 20 non-diabetic controls stratified by KDIGO risk and followed for a median of 5.6 years. Differentially methylated regions associated with eGFR decline and rising albumin excretion were prioritized, and a four-locus Methylation Risk Score (MRS; SLC4A4, FSTL4, ICA1, PTK2) was derived and compared with a clinical risk model. The MRS was validated using targeted assays in an independent subset of 330 T1D participants with low or moderate KDIGO risk at baseline. In parallel, differentiated human podocytes were exposed to normal or high glucose for integrative methyl-seq, ChIP-seq and RNA-seq profiling.
T1D was characterized by loss of 5-methylcytosine with hypomethylated regions enriched at regulatory sites. The four-gene MRS predicted eGFR decline and albuminuria more accurately than clinical risk factors (AUC 0.88 vs 0.69), and the combined MRS+clinical model reached AUCs of 0.91 and 0.86 for early and advanced progressors. In the validation cohort, the MRS predicted transitions from low to higher KDIGO risk and from moderate to higher risk (AUCs 0.73–0.92), outperforming clinical scores. In podocytes, high glucose induced concordant hypomethylation, altered chromatin binding and upregulation of the MRS genes.
Blood-derived DNA methylation biomarkers, summarized in a four-locus MRS, improve early prediction and KDIGO-based risk stratification of DKD in T1D and mirror epigenetic remodeling in kidney cells, supporting their use in precision nephrology to guide earlier intervention.
Dataset DOI: 10.5061/dryad.69p8cz9h5
Description of the data and file structure
Included files are;
Filename: Sample_Information_Manuscript.tsv
Description: Sample IDs with matching relevant clinical information for all samples (n=121)
-
Dataset overview
The file contains one row per study participant (n = 121) and 16 columns. Participants fall into two groups, given by Disease_Status: "Healthy" (n = 20) non-diabetic controls, and "T1D" (n = 101) individuals with Type 1 Diabetes. Each row records demographic, clinical, and laboratory variables collected for that participant, including a kidney-disease risk classification (KDIGO_class).Variables
- Sample_ID: Unique participant /sample identifier
- Disease_Status: Study group classification. Categorical (text): "Healthy" = non-diabetic control; "T1D" = Type 1 Diabetes Mellitus
- AGE : Participant age at assessment in years
- SEX: Biological sex . Categorical (text) MALE; FEMALE
- HbA1c_baseline: Glycated haemoglobin measured at baseline in % (percent)
- DBP: Diastolic blood pressure in mmHg
- SBP: Systolic blood pressure in mmHg
- logTG: Log10-transformed triglyceride (TG) level / log10 of original TG units in mmol/L); 0 for Healthy controls
- Smoking: Smoking status. 0=No; 1 = Yes (current)
- eGFR: estimated glomerular filtration rate calculated with the CKD-EPI formula in ml/min/1.73m2
- AER: Urinary Albumin Excretion Rate. log-transformed mg/24 hours
Additional sample data is available upon request to authors.
Filename: Steno_T1D_MBDseq_Count_Matrix_Manuscript.tsv
Description: This count matrix contains read-level quantification from methylated DNA binding domain sequencing (MBD-seq) and matched input DNA controls for each biological sample. Rows represent genomic regions (e.g., MACS2 peaks), and columns represent read counts associated with each sample and assay type.
- For every sample, two types of read counts are included:
XXX_MM (MBD-seq Methylated Reads): These columns contain read counts from MBD-enriched libraries, representing methylated CpG–rich regions captured using MBD-based enrichment. Higher counts indicate stronger methylation signal at the corresponding genomic region.
XXX_IN (Input DNA Reads): These columns contain read counts from unenriched genomic DNA (“input”) for the same sample. Input counts are used to normalize MBD-seq signal and control for technical variation such as sequencing depth or local genomic accessibility.
2. Matrix Layout
Rows: Genomic regions or peaks identified from the MBD-seq workflow plus annotation (hg38). Each row typically corresponds "Loci_RegionLengthbp_CpGcount_CpGdens_distanceToTSannotation_EnsemblID_EnsemblGeneName_Genebiotype_RefseqmRNAID_RefseqGeneName"
Row name components:
- Loci: Genomic coordinates of the peak or region (typically formatted as chr:start–end).
- RegionLengthbp: Length of the region in base pairs.
- CpGcount: Total number of CpG dinucleotides within the region.
- CpGdens: CpG density (CpG count normalized to region length).
- distanceToTSS: Distance from the region to the transcription start site of the nearest gene (negative = upstream, positive = downstream).
- annotation: Gene-region annotation (e.g., promoter, intron, intergenic, exon, 5′UTR, 3′UTR).
- EnsemblID: Ensembl gene identifier associated with the nearest or overlapping gene.
- EnsemblGeneName: Standard gene symbol provided by Ensembl.
- Genebiotype: Gene type (e.g., protein_coding, lincRNA, pseudogene, antisense RNA).
- RefseqmRNAID: Corresponding RefSeq mRNA accession ID (if available).
- RefseqGeneName: Gene symbol based on RefSeq annotation.
Variables:
Column names follow the format:Paired MBD-seq (sampleID_MM) and input DNA (sampleID_IN) read counts for each sample.
Access information
Other publicly accessible locations of the data:
- N/A
Data was derived from the following sources:
- N/A
Human subjects data
Human Subjects De-identification Statement
All data included in this submission have been fully de-identified in accordance with applicable legal and ethical guidelines. Explicit informed consent was obtained from all participants, including consent to publish their de-identified data in the public domain through Dryad.
To ensure confidentiality, all direct and indirect personally identifiable information was removed prior to submission. This includes removal of names, contact information, dates of birth, addresses, and any other information that could reasonably be used to re-identify individuals. Free-text fields were reviewed and edited to eliminate accidental disclosure of identifiable details. All dates were generalized or offset, and any potentially sensitive categories were aggregated where necessary to prevent re-identification.
Based on these procedures, the dataset does not contain any information that could be used to identify individual participants, and it complies with Dryad’s human subjects data standards.
Leukocyte DNA was processed to generate genome-wide methylation profiles (MBD-seq). Genomic DNA isolated from blood leukocytes was fragmented and enriched for methylated CpG regions using MBD-capture (MethylMiner™). Sequencing libraries were prepared with the NEBNext® Ultra™ II DNA Library Prep Kit and run on the Illumina NovaSeq 6000 platform (150 bp paired-end). Unenriched genomic DNA was also sequenced to enable normalization.
Raw reads were trimmed with Fastx and aligned to the hg38 reference genome using BWA-MEM. Methylated regions (peaks) were identified with MACS2 (fold-enrichment > 4, P < 0.01), filtered to remove ENCODE blacklisted regions, and annotated with ChIPseeker. Peak quantification used Bedtools multicov, retaining regions with >10 reads.
