Data from: Insights into longevity and virus-driven adaptation from Myotis bat genomes
Data files
Jul 17, 2026 version files 3.31 GB
-
536Mammals_alignments_noATGPlaceholder_no50pcGap.tgz
1.82 GB
-
536Mammals-ChiropteraOnly-rooted-nodeNames.nwk.tar.gz
1.76 KB
-
ABSREL-Myotis-omegas.csv.tar.gz
13.34 MB
-
ASTRAL4-536Mammals-revision-scoreOnly-detailed.phy.tar.gz
29.29 KB
-
BatsOnly_ABSREL_highlight-bats-nodeNames_Myotis-everything.tsv.bz2
89.68 MB
-
carnivora_genes_human-homolog.txt
576.83 KB
-
cetartio_genes_human-homolog.txt
628.73 KB
-
glires_genes_human-homolog.txt
612.60 KB
-
myotis_genes_human-homolog.txt
659.48 KB
-
README.md
7.73 KB
-
VIP_alignment_bats-no-myotis.tgz
151.27 MB
-
VIP_alignment_carnivores_dogaligned.tgz
228.25 MB
-
VIP_alignment_glires_mousealigned.tgz
452.32 MB
-
VIP_alignment_myotis.tgz
52.29 MB
-
VIP_alignment_primates_homoaligned.tgz
164.31 MB
-
VIP_alignment_ungulates_cowaligned.tgz
328.52 MB
Abstract
The genus Myotis is one of the largest clades of bats and exhibits some of the most extreme variation in lifespans among mammals alongside unique adaptations to viral tolerance and immune defense. To study the evolution of these phenotypes, we generated cell lines and near-complete genome assemblies for 8 closely related Myotis. Using genome-wide screens of positive selection, analyses of structural variation, and functional experiments in primary cells, we identify patterns of adaptation contributing to longevity, cancer resistance, and viral interactions. We demonstrate distinct modes of adaptation to DNA and RNA viruses compared to all other mammals, with bats exhibiting genome-wide overrepresentation of positive selection for DNA virus-interacting proteins and elevated rates of copy number variation for RNA virus-interacting proteins. Characterization of Myotis-specific duplications of the key immune factor protein kinase R (PKR) reveals multiple ancient segregating trans-species copy number polymorphisms. We show that the recurrent evolution of longevity seen in Myotis is associated with positive selection in cancer pathways, and demonstrate a unique response to DNA damage in primary cells of the long-lived M. lucifugus. Together, our results suggest that bats’ remarkable longevity and immunity are linked through pleiotropic adaptations to viruses and aging-related disease.
Dataset DOI: 10.5061/dryad.02v6wwqfh
Description of the data and file structure
This dataset contains the phylogenetic, genomic, and selection scan data used for the analyses described in Vazquez & Lauterbur et al. This includes the phylogeny generated of 536 mammal species from the new alignments presented herein as well as publicly available reference genomes and the ASTRAL scores; alignments and human homolog cross-references for selection and adaptation enrichment analyses; and the results of these selection analyses.
Files and variables
File: 536Mammals-ChiropteraOnly-rooted-nodeNames.nwk.tar.gz
Description: The phylogeny of our Myotis plus other publicly available genomes as listed in Supplemental Table S2.
File: ABSREL-Myotis-omegas.csv.tar.gz
Description: Table of all ABSREL-calculated omega values and statistics from the Myotis genomes generated in this manuscript. Contains columns for species tested, gene identifier, ABSREL output values ("Baseline MG94xREV", "Full adaptive model", "Full adaptive model (non-synonymous subs/site)", "Full adaptive model (synonymous subs/site)", "LRT", "Nucleotide GTR", "Rate classes", "Uncorrected P-value"), "original name" (branch name in the phylogeny, corresponds to "species" column for tips, NA for internal nodes), "pval" (corrected p value), "omega" (ABSREL output), "ENSG" (Ensembl gene ID corresponding to gene annotation ID), "source" (original source for annotation), "seqname" (scaffold where gene is located), "start" (start location of gene), "end" (end location of gene), "strand", "i" (column indicating number of matching lines in corresponding genome annotations: NAs indicate that this gene sequence was manually curated by searching the reference genome using BLAT, and was not present in the original annotation), "human_gene_name" (corresponding to Ensembl ID), "Node" (corresponds to "species" column for tips, MRCA of pairs of species for internal nodes), "species_clean" (human readable version of "Node" column).
File: ASTRAL4-536Mammals-revision-scoreOnly-detailed.phy.tar.gz
Description: ASTRAL scores for the phylogeny generated in this manuscript.
File: BatsOnly_ABSREL_highlight-bats-nodeNames_Myotis-everything.tsv.bz2
Description: A table compilation of all the ABSREL results from the multiple sequence alignments. Contains columns for "species" (species tested), "stat" (label of ABSREL value output, can take values "Baseline MG94xREV", "Baseline MG94xREV omega", "Corrected P-value", "Full adaptive model", "LRT", "Nucleotide GTR", "Rate classes", "Uncorrected P-value", "original name" - corresponds to "species" when species tested is a tip, otherwise not present), "value" (ABSREL value output), "gene".
File: carnivora_genes_human-homolog.txt
Description: Genes to human homolog name correspondence for genes found in "VIP_alignment_carnivores_dogaligned.tgz" archive.
File: cetartio_genes_human-homolog.txt
Description: Genes to human homolog name correspondence for genes found in "VIP_alignment_ungulates_cowaligned.tgz"
File: myotis_genes_human-homolog.txt
Description: Genes to human homolog name correspondence for genes found in "VIP_alignment_myotis.tgz" and "VIP_alignment_bats-no-myotis.tgz"
File: glires_genes_human-homolog.txt
Description: Genes to human homolog name correspondence for genes found in "VIP_alignment_glires_mousealigned.tgz"
All of the above gene to homolog files contain two columns, the first is the Ensembl ID and the second is the annotation gene ID it corresponds to.
File: VIP_alignment_bats-no-myotis.tgz
Description:
Fasta alignments for genes used in VIP enrichment analyses for the full bat clade without Myotis.
File: VIP_alignment_ungulates_cowaligned.tgz
Description:
Fasta alignments for genes used in VIP enrichment analyses for the ungulate clade.
File: VIP_alignment_carnivores_dogaligned.tgz
Description:
Fasta alignments for genes used in VIP enrichment analyses for the carnivora clade.
File: VIP_alignment_glires_mousealigned.tgz
Description:
Fasta alignments for genes used in VIP enrichment analyses for the glires clade.
File: VIP_alignment_primates_homoaligned.tgz
Description:
Fasta alignments for genes used in VIP enrichment analyses for the primate clade.
File: VIP_alignment_myotis.tgz
Description:
This archive contains a fasta alignment (.fa) for each gene used in VIP enrichment analyses for the Myotis clade as well as the corresponding gene tree (.phy).
File: 536Mammals_alignments_noATGPlaceholder_no50pcGap.tgz
Description: Fasta alignments used for ABSREL and 536 mammal phylogeny.
Code/software
The pan-mammal phylogeny was created with IQTREE using all gene alignments with the settings “-B 1000 -m GTR+F3x4+R6.”
The ASTRAL scores for this phylogeny were generated using ASTRAL-IV.
aBSREL results were generated using aBSREL (version 2.5.48) to test for selection at each branch within the Nearctic Myotis clade for 15,734 gene alignments spanning 534 mammals.
Gene annotations for the Myotis genome assemblies completed for this work were generated using a combination of TOGA v1.0.1 using human annotation references, NCBI RefSeq annotations lifted over from Myotis myotis, AUGUSTUS v3.4, genes mapped from the UNIPARC database using miniprot v0.6-r194-dirty, and transcriptome data generated for this publication assembled with TRINITY (v 2.13.2) and mapped to our genomes using minimap (v 2.24). We used EvidenceModeler36 (version 2.0) to generate an initial consensus gene set using only the best lines of evidence (AUGUSTUS, weight 2; high-quality AUGUSTUS, weight 5; TOGA-hg38, weight 12; miniprot-UniParc, weight 5; and LiftOff-mMyoMyo1, weight 5) with hints from protein orthology (miniprot-UniParc, weight 6) and RNA-seq (TRINITY, weight 5) for alternative splicing. We cross-referenced our gene annotations against the SwissProt database using DIAMOND (v. 2.1.4) with settings “--ultra-sensitive --outfmt 6 qseqid bitscore sseqid pident length mismatch gapopen qlen qstart qend slen sstart send ppos evalue --max-target-seqs 1 --evalue 1e-10 ” and kept all genes that matched over 50% of the target sequence, with at least 80% identity, coded for at least 50 amino acids, and contained both a start and stop codon with no internal stop codons.
VIP_alignment_* files contain the alignments created for each gene with one-to-one orthologs (determined using OrthoFinder 2.5.4) in the corresponding taxonomic set indicated by the file name. Alignments were generated in a three-step process: 1. The sequences of each group of orthologs were aligned using MACSE v2 with default settings, using the species indicated in the filename as the reference; 2. Alignments were cleaned with PREQUAL; 3. The remaining parts of orthologous sequences that passed PREQUAL filtering were re-aligned using MACSE v2 with default settings.
Genes in the "536Mammals_alignments_noATGPlaceholder_no50pcGap.tgz" alignments were aligned following the same three-step process as the "VIP_alignment_*" files. Alignments were filtered to remove entries from species whose sequences were comprised of 50% or more gaps.
*genes_human-homolog.txt files contain the correspondence between the gene designations (generated based on the TOGA annotations) and the Ensembl geneID.
Access information
Other publicly accessible locations of the data:
- Genome assemblies and annotations have been submitted to NCBI and are pending review there.
