Dataset for: Multi-organ acute and delayed effects following exposure to neutron radiation
Data files
Apr 24, 2026 version files 113.95 MB
-
Proteomics_Sample_List.xlsx
11.97 KB
-
Proteomics_tims-diann_result.zip
52.70 MB
-
README.md
6.05 KB
-
Untargeted_Metabolomics_Sample_List.xlsx
34.21 KB
-
Untargeted_ms1_sample_list_neg.xlsx
24.01 KB
-
Untargeted_ms1_sample_list_pos.xlsx
24.01 KB
-
Untargeted_urine_neutron_cohort-1-2_metabolomics_combined_combined_neg_mf.csv
33.54 MB
-
Untargeted_urine_neutron_cohort-1-2_metabolomics_combined_combined_pos_mf.csv
27.60 MB
Abstract
Neutron radiation exposure is a life-threatening consequence of nuclear disasters, a known hazard in space exploration, and a potential side effect in therapeutic proton applications. Prior studies have shown that neutron radiation is more biologically damaging than photon irradiation for acute radiation injury to normal tissue, the rate of tumor induction, and reduced life span. However, the effect of high dose neutron exposures on late responding organs remains unknown. To determine the response of late responding organs including the lung, heart, eye, and kidney, we assessed survival and characterized organ-specific injuries in adult rats exposed to 2 – 10 Gy neutrons with shielding of approximately 5-8% bone marrow at a clinical cyclotron-based neutron therapy facility. Additional cohorts of rats were exposed to 320 kV X-rays at equivalent doses. Survival through the acute and delayed sequelae of organ radiation injury was monitored for up to one year post-exposure. Longitudinal secondary outcomes, including body weight, complete blood cell counts, breathing rate, blood urea nitrogen, urinary proteomics and metabolomics, as well as the incidence of cataracts and intraocular pressure, were monitored. During the acute sequelae of gastrointestinal and hematopoietic injury, neutron exposure was 1.5-1.6 times more damaging than equivalent doses of X-ray. For late-responding organs, the relative biological effect (RBE) of neutron exposure to the kidney was 1.56, compared with 1.3 for the lung or residual bone marrow damage. Consistent with the observed renal sensitivity to neutron exposure, radiation-induced urine proteomic and metabolic profiles were altered in a time- and dose-dependent manner. The urine proteome exhibited dysregulation of proteins associated with oxidative stress and renal injury. In conclusion, the RBE of neutron exposure was found to be organ-dependent during long-term follow-up.
Dataset DOI: 10.5061/dryad.02v6wwqjh
Description of the data and file structure
Untargeted metabolomics and proteomics datasets have been generated from urine samples collected in a rat model of radiation exposure. The metabolomics dataset consists of LC–MS data acquired in both positive and negative electrospray ionization modes using UPLC-QToF-MS. The proteomics dataset comprises diaPASEF mass spectrometry data acquired on a timsTOF HT platform, along with the corresponding spectral library–based protein identification and quantification results.
Files and variables
File: Untargeted_Metabolomics_Sample_List.xlsx
Description: Sample metadata sheet for all urine samples analyzed by untargeted metabolomics. Contains mapping metadata for positive and negative ionization modes for file names, sample identifiers, experimental group information, and analytical batch metadata.
Variables
- POS File Name: Positive mode data acquisition file name
- NEG File Name: Negative mode data acquisition file name
- FILE.TEXT: Sample ID for acquisition
- Cohort: Batch
- Rat ID: Animal Identifier
- Timepoint (D): Timepoint in Days
- Sex: Rat Sex (Male or Female)
- Radiation Dose (Gy): Dose in Gy
File: Untargeted_ms1_sample_list_pos.xlsx
Description: Sample list corresponding to MS1 data acquired in positive electrospray ionization (ESI+) mode. Selected sample data files were converted from raw instrument format for downstream pre-processing and statistical analysis, as indicated by "DBridge" in the PROCESS column, which denotes files processed using the MassLynx DataBridge utility and included in the downstream analysis.
Variables
- FILE.NAME: Data File Name
- FILE.TEXT: Sample Identifier
- SAMPLE.LOCATION: Vial Location for Data Acquisition
- INJ.VOL: Injection Volume
- MS.FILE: Mass Spectrometer Parameter File Name
- INLET.FILE: Chromatography Parameter File Name
- MS.TUNE.FILE: Mass Spectrometer Tuning Parameters
- PROCESS: Parameter For Data File Conversion (Entries labeled "DBridge" indicate that the corresponding raw data files were converted using the DataBridge utility within MassLynx to a unified format suitable for downstream processing)
File: Untargeted_ms1_sample_list_neg.xlsx
Description: Sample list corresponding to MS1 data acquired in negative electrospray ionization (ESI-) mode. Selected sample data files were converted from raw instrument format for downstream pre-processing and statistical analysis, as indicated by "DBridge" in the PROCESS column, which denotes files processed using the MassLynx DataBridge utility and included in the downstream analysis.
Variables
- FILE.NAME: Data File Name
- FILE.TEXT: Sample Identifier
- SAMPLE.LOCATION: Vial Location for Data Acquisition
- INJ.VOL: Injection Volume
- MS.FILE: Mass Spectrometer Parameter File Name
- INLET.FILE: Chromatography Parameter File Name
- MS.TUNE.FILE: Mass Spectrometer Tuning Parameters
- PROCESS: Parameter For Data File Conversion (Entries labeled "DBridge" indicate that the corresponding raw data files were converted using the DataBridge utility within MassLynx to a unified format suitable for downstream processing)
File: Untargeted_urine_neutron_cohort-1-2_metabolomics_combined_combined_pos_mf.csv
Description: Processed feature matrix from untargeted metabolomics acquired in positive ionization mode. Rows contain features detected and columns contain samples and corresponding peak intensities for each feature.
Variables
- mzmed: Median mass-to-charge ratio (m/z) of detected feature
- rtmed: Median chromatographic retention time
- 20220829_cheema_rat_urine_neutron_metabolomics_POS_XXXXX: Positive Mode File Name. Maps to sample list files.
File: Untargeted_urine_neutron_cohort-1-2_metabolomics_combined_combined_neg_mf.csv
Description: Processed feature matrix from untargeted metabolomics acquired in negative ionization mode. Rows contain features detected and columns contain samples and corresponding peak intensities for each feature.
Variables
- mzmed: Median mass-to-charge ratio (m/z) of detected feature
- rtmed: Median chromatographic retention time
- 20220829_cheema_rat_urine_neutron_metabolomics_POS_XXXXX: Positive Mode File Name. Maps to sample list files.
File: Proteomics_Sample_List.xlsx
Description:
Variables
- File Name: Data File Name
- Randomized Order: Acquisition Sequence Number
- Rat ID: Animal Identifier
- Timepoint (D): Timepoint in Days
- Sex: Rat Sex (Male or Female)
- Radiation Dose (Gy): Dose in Gy
File: Proteomics_tims-diann_result.zip
Description: Compressed archive containing processed quantitative proteomics results generated using DIA-NN within the tims-DIA workflow. The archive includes standard DIA-NN output matrices and summary result tables generated from diaPASEF data analysis.
Note on Missing Values: Output tables may contain empty (blank) cells. These blanks are retained as originally generated by DIA-NN and reflect cases where no quantitative value was reported (e.g., signal not detected, below threshold, or not quantified in a given run). Empty cells were not replaced with placeholder values such as "n/a" or "null" to preserve compatibility with downstream analysis pipelines and software that rely on the native DIA-NN output format. Users should interpret blank entries as missing quantitative values rather than zero measurements. The distinction is important, as absence of a value does not imply absence of the analyte, but rather that it was not confidently quantified under the applied analysis criteria.
Code/software
All data files provided can be viewed using any free software that can open tab-delimited text files, such as R, Python, LibreOffice Calc, or Microsoft Excel Online. No specialized software is required to inspect the data.
Untargeted Metabolomics
Sample Preparation
The sample sequence was randomized before sample preparation to avoid bias. Urine samples were prepared by adding 75 µL of an extraction solution containing internal standards made up of 50/50 water/methanol and 10 µL debrisoquine (1 mg/mL in ddH2O), and 50 µL of 4-nitrobenzoic acid (1 mg/mL in Methanol) (per 10 mL) to 25 µL of rodent urine samples. The samples were kept on ice for 20 minutes then 75 µL of acetonitrile was added. The samples were then centrifuged at 15,493 x g for 20 minutes at 4ºC and the supernatant was transferred to a mass spectrometry (MS) vial for LC-MS analysis.
Data Acquisition
A volume of 1 µL of each prepared sample was injected onto a Waters Acquity BEH C18 1.7 μm, 2.1 × 50 mm column using an Acquity UPLC system coupled to a Xevo G2-S quadrupole-time-of-flight mass spectrometer with an electrospray ionization source (UPLC-ESI-QToF-MS) (Waters Corporation, Milford, MA). The mobile phases consisted of 100% water (solvent A) and 100% acetonitrile containing 0.1% formic acid (solvent B). The solvent flow rate for the metabolomics acquisition was set to 0.5 mL/min with the column maintained at 40°C. The LC gradient was as follows: Initial – 95% A, 5% B; 0.5 minutes – 95% A, 5% B; 4.0 minutes – 80% A, 20% B; 8.0 minutes – 5% A, 95% B; 9.0 minutes – 5% A, 95% B; 9.1 minutes – 95% A, 5% B, 11.0 minutes – 95% A, 5% B. The column eluent was introduced into the Xevo G2-S mass spectrometer by electrospray operating in either positive or negative ionization mode. Positive mode had a capillary voltage of 3.00 kV and a sampling cone voltage of 30 V. Negative mode had a capillary voltage of 2.00 kV and had a sampling cone voltage of 30 V. The desolvation gas flow was set to 1,000 L/hour and the desolvation temperature was set to 500°C. The cone gas flow was 25 L/hour and the source temperature was set to 120°C. The data were acquired in the sensitivity MS mode with a scan time of 0.300 seconds and an interscan time of 0.014 seconds. Accurate mass was maintained by infusing Leucine Enkephalin (556.2771 [M+H]+/554.2615 [M-H]-) in 50% aqueous acetonitrile (2.0 ng/mL) at a rate of 10 µL/min via the Lockspray interface every 10 seconds. The data were acquired in centroid mode over a 50.0-1,200.0 m/z mass range for TOF-MS scanning. An aliquot of each sample was pooled and used as a quality control (QC), representing all metabolites present. This QC sample was run at the beginning of the sequence to condition the column and then injected every 10 samples to check mass accuracy, ensure the presence of internal standard, and monitor shifts in retention time and signal intensities.
Data Processing
Data Preprocessing and Normalization
The untargeted raw data was first converted to the NetCDF unified data format using the Databridge tool in MassLynx (Waters Corporation, Milford, MA). An in-house implementation of the XCMS R package (Scripps Institute, La Jolla, CA) was used for peak detection with ordered bijective interpolated warping algorithm utilized for retention time correction. All data processing and statistical analyses were performed in R (version 4.3.1 or later) using packages from the Tidyverse and Bioconductor ecosystems. For untargeted analyses, the intensity of each analyte was normalized to a specific internal standard (IS) intensity value specific to each ionization mode- followed by standard preprocessing using the xcms R package. To create a unified dataset for analysis, the positive and negative ionization mode data were combined. A global feature filter was applied to the combined matrix, removing any feature with more than 20% missing values across all samples. Remaining missing values were imputed using a k-Nearest Neighbors (kNN) algorithm (k=10), as implemented in the bnstruct package. Following imputation, the data were normalized to correct for sample-wise variation. Probabilistic Quotient Normalization (PQN) was applied, using the median spectrum of all pooled quality control (QC) samples as the reference. The data were then log-2 transformed to stabilize variance, better align with the normality assumptions of statistical tests, and to transform multiplicative effects into additive ones.
Batch Correction and Data Scaling
To correct for systematic differences between the two experimental cohorts, the custom ComBat algorithm was applied to the log-2 transformed data. The final batch-corrected, log-2 transformed data matrix was used for calculating log-2-fold changes. A second version of this matrix was generated by applying Pareto scaling (mean-centering and dividing each feature by its standard deviation) for use in statistical modeling.
Proteomics
Sample Preparation
Rat urine samples were thawed on ice and mixed well before the centrifugation (10,000 x g, 10 min, 4°C). Samples were desalted and concentrated as described previously with minor modification 17 . Briefly, 200 μL of the clarified urine samples were loaded onto pre-rinsed 3 kDa Amicon centrifugal filters and centrifuged at 14,000 × g for 45 min at 4°C. Then, centrifugal filters were washed twice with 200 μL of Milli-Q water. Next, 200 μL of Milli-Q water was added to the filter to suspend the protein, and the proteins were extracted by flipping the filters and then centrifuged at 10,000 x g for 2min. The total protein concentration of the samples was determined by the BCA assay (Pierce™ BCA Protein Assay Kit, Thermo). Twenty micrograms of protein samples were processed using the iST-BCT kit in accordance with the manufacturer's instructions from PreOmics. Briefly, 20 µg of protein was added to 50 µL of LYSE-BCT, followed by heating at 95 °C for 10 minutes with shaking at 1,000 rpm to reduce disulfide bridges, alkylate cysteines, and denature proteins. Following a 5-minute room-temperature cooling phase, the mixture was supplemented with Trypsin and LysC, and the proteins were digested for 1 hour at 37 °C. The "Stop" solution was added to halt digestion. Three rounds of washing and elution into the collection plate using the supplied solutions were performed to achieve peptide purification. The samples were centrifuged for 3 minutes at 2,250 x g. Peptides were measured using the Quantitative Fluorometric Peptide Assay, transferred to low-bind tubes, and dried in a vacuum centrifuge, as per manufacturer's recommendations (ThermoFisher, USA). Finally, an estimated 500 ng of peptide per sample was resuspended in water containing 0.1% FA for downstream analysis using mass spectrometry.
Data Acquisition
Peptides from individual samples were separated using a nanoElute 2 (Bruker Daltonik Scientific) coupled online to a timsTOF HT mass spectrometer (Bruker Daltonik). Peptides were analytically separated on a PepSep25 column (75 μm × 25 cm, 1.5 μm, C18) and heated to 50 °C at a flow rate of 300 nL/min. LC mobile phases A and B were water with 0.1% FA (v/v) and ACN with 0.1% FA (v/v), respectively. The nanoLC was coupled to the timsTOF HT via a modified nanoelectrospray ion source (Captive Spray II; Bruker Daltonik). The peptides were separated using a 40-minute gradient from 2% to 95% B at a flow rate of 300 nL/min. The gradient was composed of sequential steps ranging from 2 to 35% B in 30 min, and from 35 to 95% B in 30.5 min. The column was cleaned 7 min step at 95% and equilibrated at initial conditions with four column volumes. The mass spectrometer was fully calibrated before the acquisition. The timsTOF HT was operated in diaPASEF mode with the following parameters: mass range, 100 to 1,700 m/z; 1/K0 start 0.75 V·s/cm 2 and end 1.30 V·s/cm 2 ; ramp time, 100 ms; lock duty cycle to 100%; capillary voltage, 1,600 V; dry gas, 3 L/min; dry temp, 180 °C. The collision energy was ramped linearly as a function of mobility from 59 eV at 1/K0 = 1.6 V·s/cm 2 to 20 eV at 1/K0 = 0.6 V·s/cm 2 . The diaPASEF scan consisted of 60 windows, defined from 350 to 1,250 Da, with a 1 Da overlap. The overall acquisition cycle of 1.18 s comprised one full TIMS-MS scan and 20 parallel accumulation serial fragmentation (PASEF) MS/MS scans.
Data Processing
Proteomics data were analyzed in Realtime ProteoScapeTM (2024, Bruker Scientific LLC, http://www.bruker.com), searched against the rat Swiss-Prot database with the species taxonomy set to Rattus norvegicus (reviewed sequences only; downloaded on May, 2025). TimsTOF acquisition parameters were set in Tims-DiaNN with DIA scan mass tolerance as 20 ppm MS1 mass tolerance as 15 ppm and Protein Q-value filter as 0.01 with classical MBR activation. Raw data from the DIA were processed against the in-house-built rat urine spectral library. Raw diaPASEF data were processed in ProteoScape using the dia-LFQ TIMS-DIA-NN workflow. Spectral identification and quantification were performed with DIA-NN, operating in a project-specific spectral library. The search was configured with an MS1 mass tolerance of 15 ppm and a protein-level false discovery rate (FDR) threshold (Q-value) of 0.001. Protein quantification was performed using the classical match-between-runs (MBR) algorithm, maximizing peptide and protein identification across samples. Differential protein expression analysis was performed using R and the limma package. Protein intensities were log-2-transformed, and missing values were imputed within each experimental group using the median intensity. Linear models were fitted for each protein across sample groups, and empirical Bayes moderation was applied to improve variance estimates. Adjusted p-values were calculated using the Benjamini–Hochberg method to control the false discovery rate (FDR). Proteins with an FDR < 0.05 were considered statistically significant.
