SARS-CoV-2 virus infection of Peromyscus leucopus - differentially expressed genes
Abstract
This dataset comprises RNA sequencing (RNA-seq) data from lung and brain tissues of Peromyscus leucopus infected with SARS-CoV-2, or animals exposed to medium as controls. Tissue samples were collected at day 3 and day 6 post-infection. Total RNA was extracted, and RNA-seq libraries were prepared and sequenced using high-throughput sequencing platforms. Raw sequencing reads were processed and aligned to the reference P. leucopus coding DNA sequence (CDS) DOI: 10.5061/dryad.hdr7sqvwr. Gene expression levels were quantified and normalized, yielding Transcripts Per Million (TPM) DOI: 10.5061/dryad.4j0zpc8sk. TPM data were used for the further study of computational analyses to identify differentially expressed genes (DEGs) across tissues (lung and brain) and time points (day 3 and day 6 post-infection), enabling comparative assessment of transcriptional responses to SARS-CoV-2 infection.
GENERAL INFORMATION
1. Title of Dataset: SARS-CoV-2 virus infection of Peromyscus leucopus
2. Author Information
A. Principal Investigator Contact Information
Name: Alan G. Barbour
Institution: University of California Irvine
Address: 843 Health Sciences Court, Irvine, CA 92697
Email: abarbour@uci.edu
B. Associate or Co-investigator Contact Information
Name: Ana Milovic
Institution: University of California Irvine
Address: 843 Health Sciences Court, Irvine, CA 92697
Email: anamilovic@gmail.com
3. Date of data collection (single date, range, approximate date): 2020-2022
4. Geographic location of data collection: Irvine, California, USA
5. Information about funding sources that supported the collection of the data: National Institute of Allergy and Infectious Diseases grants AI-136523 and AI-157513
SHARING/ACCESS INFORMATION
1. Licenses/restrictions placed on the data: None
2. Links to publications that cite or use the data: https://doi.org/10.64898/2026.03.13.711660
3. Links to other publicly accessible locations of the data: Not applicable
4. Links/relationships to ancillary data sets: See below for BioProject numbers.
5. Was data derived from another source? yes/no
A. If yes, list source(s): No
6. Recommended citation for this dataset: N/A
DATA & FILE OVERVIEW
1. File List: DEGs.xlsx
2. Relationship between files, if important: used DOI: 10.5061/dryad.4j0zpc8sk TPM data for analysis
3. Additional related data collected that was not included in the current data package: None
4. Are there multiple versions of the dataset? No
METHODOLOGICAL INFORMATION
1. Description of methods used for collection/generation of data: This dataset comprises RNA sequencing (RNA-seq) data from lung and brain tissues of Peromyscus leucopus infected with SARS-CoV-2, or animals exposed to medium as controls. Tissue samples were collected at day 3 and day 6 post-infection. Total RNA was extracted, and RNA-seq libraries were prepared and sequenced using high-throughput sequencing platforms. Raw sequencing reads were processed and aligned to the reference P. leucopus coding DNA sequence (CDS). Gene expression levels were quantified and normalized, yielding Transcripts Per Million (TPM) values for each gene. TPM-normalized expression values were used as input for downstream computational analysis. Differential gene expression analyses were conducted in RStudio to identify differentially expressed genes (DEGs) between SARS-CoV-2–infected and control samples at day 3 and day 6 post-infection, separately for lung and brain tissues.
2. Instrument- or software-specific information needed to interpret the data: R studio
3. Standards and calibration information, if appropriate: NA
4. Environmental/experimental conditions: see above
5. Describe any quality-assurance procedures performed on the data: NA
6. People involved with sample collection, processing, analysis and/or submission: Ana Milovic, Alan G. Barbour
DATA-SPECIFIC INFORMATION FOR: Lungs_DEG_V3
1. Number of variables: 5
2. Number of cases/rows: 22599 rows (including header)
3. Variable List: logFC (Log2 fold change in gene expression between the compared conditions. Positive values indicate higher expression in the test condition, while negative values indicate lower expression relative to the reference condition.), logCPM (Log2 counts per million; a normalized measure of gene expression abundance across samples.), PValue (Raw p-value from the statistical test assessing differential gene expression), FDR (False discovery rate–adjusted p-value that accounts for multiple testing.), gene (Gene symbol or short gene name used to identify the transcript or gene in the dataset.),
4. Missing data codes: None
5. Specialized formats or other abbreviations used: None
DATA-SPECIFIC INFORMATION FOR: Lungs_DEG_V6
1. Number of variables: 5
2. Number of cases/rows: 22599 rows (including header)
3. Variable List: logFC (Log2 fold change in gene expression between the compared conditions. Positive values indicate higher expression in the test condition, while negative values indicate lower expression relative to the reference condition.), logCPM (Log2 counts per million; a normalized measure of gene expression abundance across samples.), PValue (Raw p-value from the statistical test assessing differential gene expression), FDR (False discovery rate–adjusted p-value that accounts for multiple testing.), gene (Gene symbol or short gene name used to identify the transcript or gene in the dataset.),
4. Missing data codes: None
5. Specialized formats or other abbreviations used: None
DATA-SPECIFIC INFORMATION FOR: Brain_DEG_V3_E2
1. Number of variables: 5
2. Number of cases/rows: 22599 rows (including header)
3. Variable List: logFC (Log2 fold change in gene expression between the compared conditions. Positive values indicate higher expression in the test condition, while negative values indicate lower expression relative to the reference condition.), logCPM (Log2 counts per million; a normalized measure of gene expression abundance across samples.), PValue (Raw p-value from the statistical test assessing differential gene expression), FDR (False discovery rate–adjusted p-value that accounts for multiple testing.), gene (Gene symbol or short gene name used to identify the transcript or gene in the dataset.),
4. Missing data codes: None
5. Specialized formats or other abbreviations used: None
DATA-SPECIFIC INFORMATION FOR: Brain_DEG_V6_E2
1. Number of variables: 5
2. Number of cases/rows: 22599 rows (including header)
3. Variable List: logFC (Log2 fold change in gene expression between the compared conditions. Positive values indicate higher expression in the test condition, while negative values indicate lower expression relative to the reference condition.), logCPM (Log2 counts per million; a normalized measure of gene expression abundance across samples.), PValue (Raw p-value from the statistical test assessing differential gene expression), FDR (False discovery rate–adjusted p-value that accounts for multiple testing.), gene (Gene symbol or short gene name used to identify the transcript or gene in the dataset.),
4. Missing data codes: None
5. Specialized formats or other abbreviations used: None
