Data from: Reference genomes and fossils revise bat family phylogeny and biogeography
Data files
Aug 26, 2026 version files 1.49 TB
-
Assemblies_and_species.tsv
14.42 KB
-
README.md
65.02 KB
-
Section_10.tgz
3.28 GB
-
Section_12.tgz
1.86 GB
-
Section_13.zip
4.04 KB
-
Section_14.tgz
507.34 KB
-
Section_15.tgz
6.01 MB
-
Section_16.tgz
21.33 GB
-
Section_17.tgz
375.73 MB
-
Section_18.tgz
451.40 KB
-
Section_5.zip
390.90 MB
-
Section_6.tgz
278.68 GB
-
Section_7.tgz
1.34 MB
-
Section_8_cactus_aligment_projectedTo_hg38.maf.gz
200.90 GB
-
Section_8_cactus_aligment_projectedTo_HLmyoMyo6.maf.gz
237.75 GB
-
Section_8_cactus_aligment_projectedTo_HLrhiFer5.maf.gz
256.70 GB
-
Section_8_final_cactus_alignment.hal
268.85 GB
-
Section_8_multiz_aligment_hg38.maf.gz
76.51 GB
-
Section_8_multiz_aligment_HLmyoMyo6.maf.gz
66.17 GB
-
Section_8_multiz_aligment_HLrhiFer5.maf.gz
72.86 GB
-
Section_9.tgz
291.31 MB
Abstract
Bats are extraordinary amongst mammals, having uniquely evolved powered-flight and laryngeal echolocation, along with disease resistance, extended healthspans and the ability to hibernate. However, bats’ evolutionary history and our understanding of these adaptations remain unresolved. We analysed chromosome-level, long-read genome assemblies from 103 bat species, including 42 new assemblies, representing all 21 bat families. This dataset, unprecedented in taxonomic scope and assembly quality, yielded a new bat phylogeny. We placed Myzopodidae as the earliest branch within Vespertilionoidea, and resolved yangochiropteran relationships, identifying Emballonuroidea and Vespertilionoidea as sister groups. Our analysis revealed a mosaic evolutionary history across bats and explained why previous phylogenetic studies were misled. Chromosomal ancestral state reconstructions supported 26 ancestral bat chromosomes. We integrated a morphological dataset of 699 characters for 65 species, including 44 pre-Quaternary fossils and representatives of most living bat families, with neutrally-evolving genomic sites. Fossilized Birth-Death and Dispersal-Extinction-Cladogenesis analyses showed that bats, and thus powered-flight, likely originated in Europe in the late Palaeocene, refuting African and North American origins. Placement of the fossil †Vielasia in the oldest ‘Eochiroptera’ clade indicates that laryngeal echolocation predates crown-bat diversification. Total evidence dating including the fossil taxa significantly reduced Unrepresented-Basal-Branch-Lengths compared to molecular-only divergence estimates. By integrating comprehensive genomic and morphological data sets, analysed using innovative methods, we resolve long-standing controversies in bat biology and provide new insights into bats’ evolutionary history and trait diversification.
This repository represents the supplementary data for sections 5-10 and 12-18 of the above referenced manuscript.
Sections 1, 2, 3, 4, and 11 are not included and correspond to supplemental documents associated with the manuscript.
Data are divided by Supplementary Material section number:
- Section 5. Annotation of coding genes
- Section 6. Annotation of TEs
- Section 7. Annotation and Evolution of microRNAs
- Section 8. Multiple Genome Alignments
- Section 9. Phylogenomic Inference (Part I)
- Section 10. Phylogenomic Inference (Part II)
- Section 12. Phylogenomic Inference (Part IV)
- Section 13. Phylogenomic Inference from an X-linked Recombination Desert region (XLRD)
- Section 14. Phylogenetic inference: Mitochondrial Genome
- Section 15. Reconstructing the ancestral bat karyotype
- Section 16. Fossils, Morphological Data, and Dates
- Section 17. Tip dating: Fossilized Birth Death Models
- Section 18. Biogeography
All sections are tarred and/or zipped to save space and keep section identifiers. Exceptionally large data (8, for example) are divided into multiple files.
In some analyses, particularly section 8, two alternate naming systems were used, ToLIDs (e.g. mAntPal2.1.pri) vs. Senckenberg IDs (e.g. HLantPal2). Replacing those names in the very, very large whole genome alignment files is prohibitively time consuming, thus we are providing a key for interconversion
File types and access:
Archive files
.tgz - Tarred and gzipped archives. Accessible by using tar -xvf archive on linux operating systems or using 7zip.
.zip - Zipped archives. Accessible using unzip on linux operating systems or any of several compression packages including 7zip.
Tabular data
- .bed - Text file with data in bed format. Viewable in any text editor.
- .tsv - Tab-separated values file. Plain text file that uses tab characters to separate columns and new lines to separate row. View in any text editor.
- .align.gz - Compressed text file. Decompress using gzip or 7zip and view in any text editor.
- .out.gz - Compressed text file. Decompress using gzip or 7zip and view in any text editor.
- .summary.gz - Compressed text file. Decompress using gzip or 7zip and view in any text editor.
- .divsum.gz - Compressed text file. Decompress using gzip or 7zip and view in any text editor.
- .xlsx/.xls - Excel file.
Image files
- .png - Image file. View in any image software.
Sequence data
- .fa.gz - Compressed fasta formatted sequence file. Decompress using gzip or 7zip and view in any text editor of fasta viewer, e.g. BioEdit, Aliview, etc.
- .fasta - Fasta formatted sequence file. View in any text editor of fasta viewer, e.g. BioEdit, Aliview, etc.
Alignments
- .hal - Hierarchical alignment file. A compressed, graph-based data format used in comparative genomics to store whole-genome multiple sequence alignments and ancestral genome reconstructions. View and analyze using HalTools.
- .maf.gz - Mutation Annotation Format file. A tab-delimited text file that stores processed, somatic mutation data (such as single nucleotide variants and indels) discovered in cancer genomics projects. View with any text editor or view and analyze with maftools.
Tree files
- mod.zip, .neutralWindows.zip, trees.wholeGenome.zip, trees_slidingWindows.zip - Compressed text file. Compressed text files. Decompress using unzip (on linux operating systems) or 7zip and view in any text editor.
- .tre - Text file. Stores phylogenetic tree data (evolutionary relationships between species, genes, or languages). View and analyze using any text editor or specializes software. e.g. FigTree, iTOL, etc.
- .newick - Newick formatted tree file. View and analyze using any text editor or specializes software. e.g. FigTree, iTOL, etc.
- .treefile - Newick formatted tree file. View and analyze using any text editor or specializes software. e.g. FigTree, iTOL, etc.
- .nex - Nexus file for phylogenetic character data. View and analyze with any text editor or phylogenetic analysis packages, e.g. MrBayes, PAUP*, FigTree, etc.
- .log - Log file from the associated analysis. Text format. View in any text editor.
- .iqtree - Text file report from IQ-TREE software. View and analyze using any text editor, FigTree, iTOL, etc.
- .mldist - Text Maximum Likelihood distance matrix generated by IQ-TREE. View and analyze using any text editor and cluster/heatmap analysis tools, etc.
- .contree - Text consensus tree output by IQ-TREE. View and analyze using any text editor, FigTree, iTOL, etc.
- .bionj - Text tree output by IQ-TREE. View and analyze using any text editor, FigTree, iTOL, etc.
- .ckp.gz - Compressed binary checkpoint file with intermediate states during a maximum likelihood analysis. Not human readable.
Analysis scripts
- .pl - perl script.
- .R - R script.
Description of the data and file structure
File: Assemblies_and_species.tsv
Description: In some analyses, particularly section 8, alternative naming systems were used, ToLIDs (e.g. mAntPal2.1.pri) vs. Senckenberg IDs (e.g. HLantPal2). Replacing those names in the very, very large whole genome alignment files is prohibitively time consuming, thus we are providing this file as a key for interconversion.
File: Section_5.zip
Description: Single-file archive. 02.TOGAv1_bed.zip
Protein-coding gene annotations
This dataset contains protein-coding gene annotations generated with TOGA (commit v.c4bce48) for 103 bat genome assemblies representing all 21 extant bat families and seven non-chiropteran mammalian outgroups. Gene annotations were projected from the human reference genome (hg38; GRCh38.p12) using pairwise whole-genome alignments between human and each query genome.
TOGA jointly infers orthologous loci and projects protein-coding gene structures from an annotated reference genome to query genomes. The analyses used the human GENCODE v38 annotation (Ensembl 104), comprising 39,664 transcripts from 19,456 protein-coding genes. These annotations were used throughout the study for analyses requiring orthologous protein-coding genes, including assessments of genome completeness and downstream phylogenomic analyses.
The dataset is provided as the gzip-compressed archive:
The archive contains TOGA BED-format protein-coding gene annotations for 110 query genome assemblies with the appropriate genome assembly ID in place of <assembly_ID>:
- 103 bat genome assemblies
vs_<assembly_ID>.annotationTOGA_v1.bed
- 7 non-chiropteran mammalian outgroup assemblies
vs_HLeleMax2.annotationTOGA_v1.bed
vs_mm39.annotationTOGA_v1.bed
vs_HLsorFum1.annotationTOGA_v1.bed
vs_HLbosTau10.annotationTOGA_v1.bed
vs_HLequAsi3.annotationTOGA_v1.bed
vs_HLneoVis2.annotationTOGA_v1.bed
vs_HLmanPen4.annotationTOGA_v1.bed
Human (hg38) is not included as a query annotation because it served as the reference genome from which gene annotations were projected.
Each BED file contains the genomic coordinates and exon structure of protein-coding transcripts projected by TOGA onto the corresponding query genome assembly. Coordinates refer to the assembly for the species indicated by the corresponding assembly identifier/file name. Transcript and gene identities are derived from the human GENCODE v38 reference annotation.
Users should use the genome assembly corresponding to each annotation file when interpreting coordinates or extracting genomic sequences. Genome assembly identifiers and associated species information are provided in the accompanying manuscript and Supplementary Tables.
File: Section_6.tgz
Description: Multi-file archive. Repeatmasker output files for each genome assembly, the following files are provided.
<assembly_ID>.fa.align.gz - standard .align file generated by RepeatMasker. Contains the pairwise alignments between each genomic sequence match and the repeat consensus sequence that was used to annotate it. It is much more detailed than the .out file and is primarily used for manual inspection, curation, and evaluating the quality of repeat annotations.
The header reports:
Smith-Waterman score
Percent divergence
Percent deletions
Percent insertions
Query sequence
Genomic coordinates
Strand
Repeat name
Repeat class/family
Coordinates on the consensus sequence
Alignment block contains the genomic sequence (9query)compared to thee TE match.
<assembly_ID>.fa.raw.out.gz - Native .out file for each assembly from RepeatMasker. Primary annotation file produced by RepeatMasker. It contains one line per repeat annotation, summarizing where each transposable element (TE) or other repetitive sequence occurs in the genome and how well it aligns to a repeat consensus sequence.
The header reports:
Smith-Waterman score
Percent divergence
Percent deletions
Percent insertions
Query sequence
Begin
End
(Left)
Strand
Repeat name
Repeat class/family
Repeat begin
Repeat end
(Left)
ID
Overlap indicator (*)
<assembly_ID>.resolved.out.gz - Modifed RepeatMasker output in which overlapping hits are resolved in favor of the hit with lower divergence. Headers are identical to the raw.out.gz file.
<assembly_ID>.removed.out.gz - Overlap lines removed from raw.out.gz to generate resolved.out.gz. Headers are identical to raw.out.gz file.
<assembly_ID>.summary.gz - Native RepeatMasker summary file. Overall statistical summary of the RepeatMasker run data in <assembly_ID>.fa.raw.out.gz. Unlike the .out file, which contains one record per repeat annotation, the .summary file aggregates results across the entire input sequence or genome.
Top section:
Total sequence length analyzed
GC content
Total bases masked
Percentage of the genome masked
Breakdown of repeat classes and families
Repeat Class Summary section:
Column : Description
Number of elements : Total repeat annotations assigned to the class or family.
Length occupied : Total number of genomic bases covered by that class.
Percent of sequence : Fraction of the analyzed genome occupied by that class.
<assembly_ID>.divsum.gz - RepeatMasker divsum file. Reports the genomic abundance of repeats as a function of sequence divergence from their consensus sequences in <assembly_ID>.fa.raw.out.gz. It is primarily used to generate repeat landscapes, which visualize the historical accumulation of transposable elements (TEs) in a genome.
Column : Description
Div : Divergence bin (Kimura divergence in percent).
Repeat class columns : Total genomic bases assigned to that repeat class within the divergence bin.
<assembly_ID>.divsum.landscape.tsv - Landscape data from <assembly_ID>.divsum.gz in .tsv format and converted to genomic proportions.
Column : Description
Divergence : Divergence bin (Kimura divergence in percent).
TE Class/Family : Proportion of genome occupied by sequence identified by the TE Class/Family (total bp/genome length).
<assembly_ID>_stacked_bar_data.tsv - non-RepeatMasker stacked bar data. Calculated from genome proportions of each major TE Class in <assembly_ID>.resolved.out.gz.
Column : Description
Div : Divergence bin (Kimura divergence in percent).
Repeat class columns : Total genomic proportions assigned to that repeat class within the divergence bin (total bp/genome length).
<assembly_ID>_line.png - Non-RepeatMasker landscape plot. Line plot of TE occupancy calculated from genome proportions of each major TE Class. Derived from <assembly_ID>_stacked_bar_data.tsv.
<assembly_ID>_stackedbar.png - Non-RepeatMasker landscape plot. Stacked bar plot of TE occupancy calculated from genome proportions of each major TE Class. Derived from <assembly_ID>_stacked_bar_data.tsv.
<assembly_ID>_pie_data.tsv - Non-RepeatMasker pie data. Total genome proportions of each major TE Class calculated from <assembly_ID>.resolved.out.gz.
<assembly_ID>_pie.png - Non-RepeatMasker pie plot of overall TE occupancy. Derived from <assembly_ID>_pie_data.tsv.
<assembly_ID>.wUnknown_stacked_bar_data.tsv - Same as <assembly_ID>_stacked_bar_data.tsv but including annotations of repeats identified as Unknown.
<assembly_ID>.wUnknown_line.png - Same as <assembly_ID>_line.png but including annotations of repeats identified as Unknown.
<assembly_ID>.wUnknown_stackedbar.png - Same as <assembly_ID>_stacked_bar.png but including annotations of repeats identified as Unknown.
<assembly_ID>.wUnknown_pie_data.tsv - Same as <assembly_ID>_pie_data.tsv but including annotations of repeats identified as Unknown.
<assembly_ID>.wUnknown_pie.png - Same as <assembly_ID>_pie.png but including annotations of repeats identified as Unknown.
bat1k_named_20250905.fa.gz - Fasta formatted RepeatMasker library generated for this project. Standard RepeatMasker formatted headers, TE_ID#Class/Family.
For details on the methods used to generate the files as described in the associated manuscript, go to https://github.com/davidaray/Bat1k_TE_curation.
File: Section_7.tgz
Description: Single-file archive. Dataset 7.1.xlsx
This file provides microRNA gene annotations for 103 bat genomes, generated by the Bat1K consortium. Data are organised into 103 individual sheets—one per species—each formatted as a standard six-column BED table (chromosome, start, end, miRNA family name, alignment score, and strand). The dataset enables comparative genomics and evolutionary studies of miRNA repertoires across Chiroptera.
File: Section_8_final_cactus_alignment.hal
Description: Whole genome alignment generated using Progressive cactus.
File: Section_8_cactus_aligment_projectedTo_hg38.maf.gz
Description: Cactus alignment projected to human hg38.
File: Section_8_cactus_aligment_projectedTo_HLmyoMyo6.maf.gz
Description: Cactus alignment projected to Myotis myotis HLmyoMyo6.
File: Section_8_cactus_aligment_projectedTo_HLrhiFer5.maf.gz
Description: Cactus alignment projected to Rhinolophus ferrumequinum HLrhiFer5.
File: Section_8_multiz_aligment_hg38.maf.gz
Description: MultiZ alignment projected to human hg38.
File: Section_8_multiz_aligment_HLmyoMyo6.maf.gz
Description: MultiZ alignment projected to Myotis myotis HLmyoMyo6.
File: Section_8_multiz_aligment_HLrhiFer5.maf.gz
Description: MultiZ alignment projected to Rhinolophus ferrumequinum HLrhiFer5.
File: Section_9.tgz
Description: Multi-file archive. phyloFit_4Dsites.mod.zip, tree_v2RLv3.rooted_17436trees.tre, trees.neutralWindows.zip, trees.wholeGenome.zip, trees_slidingWindows.zip
Phylogenomic inference (Part I)
This dataset contains phylogenetic trees and neutral substitution models generated for the phylogenomic analyses described in Section 9 of the Supplementary Materials. The analyses integrate protein-coding genes, neutral intergenic regions, whole-genome multiple alignments, and non-overlapping genomic sliding windows across 103 bat species and 8 mammalian outgroups.
Phylogenomic analyses were performed using multiple genome alignments generated with two alignment approaches, MULTIZ and CACTUS, and three reference or projected-reference genomes: Homo sapiens (hg38), Myotis myotis (HLmyoMyo6), and Rhinolophus ferrumequinum (HLrhiFer5).
phyloFit_4Dsites.mod.zip`
This archive contains four PHAST .mod files:
├── 21FamT1.multiZHLmyoMyo6_chrASM.mod
├── 21FamT1.multiZHLrhiFer5_chrASM.mod
├── 21FamT1.multiZhg38_chrASM.mod
└── 21FamT2.multiZhg38_chrASM.mod
These files contain neutral substitution models estimated from fourfold degenerate (4D) sites using phyloFit. Fourfold degenerate sites were extracted from the whole-genome MULTIZ alignments using protein-coding gene annotations generated with TOGA. Phylogenetic models were estimated using the REV substitution model, EM optimization, and medium precision.
The files correspond to models estimated using two alternative reference tree topologies:
- T1: the phylogenetic topology inferred from protein-coding genes.
- T2: the phylogenetic topology inferred from intergenic neutral windows, and independently recovered by analyses of the X chromosomes.
For T1, models are provided for the three MULTIZ reference alignments: human (hg38), Myotis myotis (HLmyoMyo6), and Rhinolophus ferrumequinum (HLrhiFer5). For T2, the model provided here corresponds to the MULTIZ-human reference alignment (hg38).
These phyloFit models were used as neutral expectations in downstream phastCons and phyloP analyses.
tree_v2RLv3.rooted_17436trees.tre
This file contains the rooted species tree (T1) in Newick format, inferred from protein-coding gene trees.
Protein-coding phylogenetic analyses were based on 1:1 orthologous coding genes identified with TOGA. For each gene, individual orthologous exons were aligned with MACSE v.2, concatenated into complete codon alignments, and filtered with HmmCleaner. Maximum-likelihood gene trees were reconstructed with RAxML under a GTR+GAMMA substitution model.
trees.neutralWindows.zip
This archive contains phylogenetic trees in Newick format inferred from intergenic neutral windows based on MULTIZ alignments based on human (hg38), Myotis myotis (HLmyoMyo6), and Rhinolophus ferrumequinum (HLrhiFer5).
├── multiZHLmyoMyo6.concatNeutralWin_500_COV80.scored.tre
├── multiZHLrhiFer5.concatNeutralWin_500_COV80.scored.tre
└── multiZhg38.concatNeutralWin_500_COV80.scored_U2.tre
Neutral regions were defined by excluding annotated and constrained genomic elements and retaining regions with alignment coverage across at least 80 species. Windows ≥500 bp were used for phylogenomic analyses, and trees were inferred with CASTER.
trees.wholeGenome.zip
This archive contains phylogenetic trees inferred from whole-genome multiple alignments (MULTIZ and CACTUS) based on human (hg38), Myotis myotis (HLmyoMyo6), and Rhinolophus ferrumequinum (HLrhiFer5)
├── cactusHLmyoMyo6.tree_CASTERsite_Genome__CACTUS.scored.tre
├── cactusHLriFer5.tree_CASTERsite_Genome__CACTUS.scored.tre
├── cactushg38.tree_CASTERsite_Genome__CACTUS.scored.tre
├── multiZhg38.tree_CASTERsite_Genome.scored.tre
├── multizHLmyoMyo6.tree_CASTERsite_Genome.scored.tre
└── multizHLrhiFer5.tree_CASTERsite_Genome.scored.tre
Whole-genome phylogenies were reconstructed with CASTER using genome alignments generated with MULTIZ and CACTUS and the three reference or projected-reference genomes: human (hg38), Myotis myotis (HLmyoMyo6), and Rhinolophus ferrumequinum (HLrhiFer5).
trees_slidingWindows.zip
This archive contains phylogenetic trees inferred from non-overlapping genomic sliding windows across the whole-genome alignments (MULTIZ and CACTUS) based on human (hg38), Myotis myotis (HLmyoMyo6), and Rhinolophus ferrumequinum (HLrhiFer5).
HLmyoMyo6
├── CACTUSnD/keep_less80Miss
└── multiZ/keep_less80Miss
HLrhiFer5
├── CACTUSnD/keep_less80Miss
└── multiZ/keep_less80Miss
hg38
├── CACTUSnD/keep_less80Miss
└── multiZ/keep_less80Miss
Whole-genome alignments were subdivided into non-overlapping windows whose boundaries were guided by genic elements and alignment structure rather than by uniform fixed lengths. Human-based alignments primarily comprised windows of approximately 8–9 kb, whereas windows in the bat-based alignments extended up to approximately 2 Mb.
Windows containing more than 20% missing data were excluded. Phylogenetic trees for individual windows were inferred with CASTER.
File: Section_10.tgz
Description: Multi-file archive. Supermatrix_Phylogeny.tar.gz, aLRT_Trees_Astral_run.zip, AlternativeTopologies.zip
Supermatrix_Phylogeny.tar.gz:
aln_files - all gene alignment files used in the formation of the supermatrix
gene_trees - contains the gene tree inferred for each alignment in the ‘aln_files’ directory. Each gene file has an equivalent “X.treefile”, “X.log”, “X.iqtree” and equivalent intermediate files as output by IQTREE2.
supermatrix - supermatrix - contains the concatenated alignment and information (FcC_info.xls, FcC_supermatrix), partition file (NucPartitions.nex), and all equivalent IQTREE2 output files (NucPartitions.nex.
aLRT_Trees_Astral_run.zip:
all_alrt_collapsed.tre - file containing all gene trees where nodes below 70 support have been collapsed
astral_all_alrt_collapsed.tre - the inferred astral species tree using "all_alrt_collapsed.tre" as input
weighted_astral_tree.tre - the inferred weighted astral species tree
AlternativeTopologies.zip:
All_alternative_topologies_identified.txt - list of trees found to contain specific monophyletic subtree combinations
check_support.pl - perl script to count all nodes above a certain score threshold
make_subtrees.R - R script to break a tree into all possible subtrees
run.pl - master perl stcript for running other scripts on all input trees
count_monophyletic_topolgies_allbats.pl - perl script to count all monophyletic trees when all bats are present
count_monophyletic_topolgies_NOT_allbats.pl - perl script to count all monophyletic trees when not all bats are present
AllBats_Subtrees.zip - all subtree files for each gene tree where all bats are present
NotAllBats_Subtrees.zip - all subtree files for each gene tree where not all bats are present
File: Section_12.tgz
Description: Multi-file archive. Dataset_12_1_phyloP.zip
FASTA alignments of phyloP-based SNP subsets from human, Myotis myotis and Rhinolophus ferrumequinum CACTUS and MultiZ alignments.
This archive contains SNP alignment files extracted from CACTUS and MultiZ whole-genome alignments. SNPs are grouped by phyloP category: accelerated, conserved and neutral. These alignments were used as input for downstream phylogenetic tree inference in Dataset_12_2, Dataset_12_3 and Dataset_12_4.
Human_Cactus_subset_accelerated_SNPs.fasta - FASTA alignment of accelerated SNPs from the human CACTUS alignment.
Human_Cactus_subset_conserved_SNPs.fasta - FASTA alignment of conserved SNPs from the human CACTUS alignment.
Human_Cactus_subset_neutral_SNPs.fasta - FASTA alignment of neutral SNPs from the human CACTUS alignment.
Human_MultiZ_subset_accelerated_SNPs.fasta - FASTA alignment of accelerated SNPs from the human MULTIZ alignment.
Human_MultiZ_subset_conserved_SNPs.fasta - FASTA alignment of conserved SNPs from the human MULTIZ alignment.
Human_MultiZ_subset_neutral_SNPs.fasta - FASTA alignment of neutral SNPs from the human MULTIZ alignment.
Myotis_Cactus_subset_accelerated_SNPs.fasta - FASTA alignment of accelerated SNPs from the Myotis myotis CACTUS alignment.
Myotis_Cactus_subset_conserved_SNPs.fasta - FASTA alignment of conserved SNPs from the Myotis myotis CACTUS alignment.
Myotis_Cactus_subset_neutral_SNPs.fasta - FASTA alignment of neutral SNPs from the Myotis myotis CACTUS alignment.
Myotis_MultiZ_subset_accelerated_SNPs.fasta - FASTA alignment of accelerated SNPs from the Myotis myotis MULTIZ alignment.
Myotis_MultiZ_subset_conserved_SNPs.fasta - FASTA alignment of conserved SNPs from the Myotis myotis MULTIZ alignment.
Myotis_MultiZ_subset_neutral_SNPs.fasta - FASTA alignment of neutral SNPs from the Myotis myotis MULTIZ alignment.
Rhinolophus_Cactus_subset_accelerated_SNPs.fasta - FASTA alignment of accelerated SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhinolophus_Cactus_subset_conserved_SNPs.fasta - FASTA alignment of conserved SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhinolophus_Cactus_subset_neutral_SNPs.fasta - FASTA alignment of neutral SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhinolophus_MultiZ_subset_accelerated_SNPs.fasta - FASTA alignment of accelerated SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhinolophus_MultiZ_subset_conserved_SNPs.fasta - FASTA alignment of conserved SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhinolophus_MultiZ_subset_neutral_SNPs.fasta - FASTA alignment of neutral SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Description: Multi-file archive. Dataset_12_2_MLtrees.zip
Maximum likelihood trees inferred from phyloP-based SNP subsets from human, Myotis myotis and Rhinolophus ferrumequinum CACTUS and MultiZ alignments.
This archive contains maximum likelihood phylogenetic trees in Newick format. Trees were inferred from the accelerated, conserved and neutral SNP alignments in Dataset_12_1_phyloP.zip.
Human_cactus_accelerated_ML.newick - Maximum likelihood tree inferred from accelerated SNPs from the human CACTUS alignment.
Human_cactus_conserved_ML.newick - Maximum likelihood tree inferred from conserved SNPs from the human CACTUS alignment.
Human_cactus_neutral_ML.newick - Maximum likelihood tree inferred from neutral SNPs from the human CACTUS alignment.
Human_multiz_accelerated_ML.newick - Maximum likelihood tree inferred from accelerated SNPs from the human MULTIZ alignment.
Human_multiz_conserved_ML.newick - Maximum likelihood tree inferred from conserved SNPs from the human MULTIZ alignment.
Human_multiz_neutral_ML.newick - Maximum likelihood tree inferred from neutral SNPs from the human MULTIZ alignment.
Myotis_cactus_accelerated_ML.newick - Maximum likelihood tree inferred from accelerated SNPs from the Myotis myotis CACTUS alignment.
Myotis_cactus_conserved_ML.newick - Maximum likelihood tree inferred from conserved SNPs from the Myotis myotis CACTUS alignment.
Myotis_cactus_neutral_ML.newick - Maximum likelihood tree inferred from neutral SNPs from the Myotis myotis CACTUS alignment.
Myotis_multiz_accelerated_ML.newick - Maximum likelihood tree inferred from accelerated SNPs from the Myotis myotis MULTIZ alignment.
Myotis_multiz_conserved_ML.newick - Maximum likelihood tree inferred from conserved SNPs from the Myotis myotis MULTIZ alignment.
Myotis_multiz_neutral_ML.newick - Maximum likelihood tree inferred from neutral SNPs from the Myotis myotis MULTIZ alignment.
Rhino_cactus_accelerated_ML.newick - Maximum likelihood tree inferred from accelerated SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_cactus_conserved_ML.newick - Maximum likelihood tree inferred from conserved SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_cactus_neutral_ML.newick - Maximum likelihood tree inferred from neutral SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_multiz_accelerated_ML.newick - Maximum likelihood tree inferred from accelerated SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhino_multiz_conserved_ML.newick - Maximum likelihood tree inferred from conserved SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhino_multiz_neutral_ML.newick - Maximum likelihood tree inferred from neutral SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Description: Multi-file archive. Dataset_12_3_SVDQuartetstrees.zip
SVDQuartets trees inferred from phyloP-based SNP subsets from human, Myotis myotis and Rhinolophus ferrumequinum CACTUS and MultiZ alignments.
This archive contains SVDQuartets phylogenetic trees in Newick format. Trees were inferred from the accelerated, conserved and neutral SNP alignments in Dataset_12_1_phyloP.zip.
Human_cactus_accelerated_SVDQ.newick - SVDQuartets tree inferred from accelerated SNPs from the human CACTUS alignment.
Human_cactus_conserved_SVDQ.newick - SVDQuartets tree inferred from conserved SNPs from the human CACTUS alignment.
Human_cactus_neutral_SVDQ.newick - SVDQuartets tree inferred from neutral SNPs from the human CACTUS alignment.
Human_multiz_accelerated_SVDQ.newick - SVDQuartets tree inferred from accelerated SNPs from the human MULTIZ alignment.
Human_multiz_conserved_SVDQ.newick - SVDQuartets tree inferred from conserved SNPs from the human MULTIZ alignment.
Human_multiz_neutral_SVDQ.newick - SVDQuartets tree inferred from neutral SNPs from the human MULTIZ alignment.
Myotis_cactus_accelerated_SVDQ.newick - SVDQuartets tree inferred from accelerated SNPs from the Myotis myotis CACTUS alignment.
Myotis_cactus_conserved_SVDQ.newick - SVDQuartets tree inferred from conserved SNPs from the Myotis myotis CACTUS alignment.
Myotis_cactus_neutral_SVDQ.newick - SVDQuartets tree inferred from neutral SNPs from the Myotis myotis CACTUS alignment.
Myotis_multiz_accelerated_SVDQ.newick - SVDQuartets tree inferred from accelerated SNPs from the Myotis myotis MULTIZ alignment.
Myotis_multiz_conserved_SVDQ.newick - SVDQuartets tree inferred from conserved SNPs from the Myotis myotis MULTIZ alignment.
Myotis_multiz_neutral_SVDQ.newick - SVDQuartets tree inferred from neutral SNPs from the Myotis myotis MULTIZ alignment.
Rhino_cactus_accelerated_SVDQ.newick - SVDQuartets tree inferred from accelerated SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_cactus_conserved_SVDQ.newick - SVDQuartets tree inferred from conserved SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_cactus_neutral_SVDQ.newick - SVDQuartets tree inferred from neutral SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_multiz_accelerated_SVDQ.newick - SVDQuartets tree inferred from accelerated SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhino_multiz_conserved_SVDQ.newick - SVDQuartets tree inferred from conserved SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhino_multiz_neutral_SVDQ.newick - SVDQuartets tree inferred from neutral SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Description: Multi-file archive. Dataset_12_4_CASTERtrees.zip
CASTER trees inferred from phyloP-based SNP subsets from human, Myotis myotis and Rhinolophus ferrumequinum CACTUS and MultiZ alignments.
This archive contains CASTER phylogenetic trees in Newick format. Trees were inferred from the accelerated, conserved and neutral SNP alignments in Dataset_12_1_phyloP.zip.
Human_cactus_accelerated_CASTER.newick - CASTER tree inferred from accelerated SNPs from the human CACTUS alignment.
Human_cactus_conserved_CASTER.newick - CASTER tree inferred from conserved SNPs from the human CACTUS alignment.
Human_cactus_neutral_CASTER.newick - CASTER tree inferred from neutral SNPs from the human CACTUS alignment.
Human_multiz_accelerated_CASTER.newick - CASTER tree inferred from accelerated SNPs from the human MULTIZ alignment.
Human_multiz_conserved_CASTER.newick - CASTER tree inferred from conserved SNPs from the human MULTIZ alignment.
Human_multiz_neutral_CASTER.newick - CASTER tree inferred from neutral SNPs from the human MULTIZ alignment.
Myotis_cactus_accelerated_CASTER.newick - CASTER tree inferred from accelerated SNPs from the Myotis myotis CACTUS alignment.
Myotis_cactus_conserved_CASTER.newick - CASTER tree inferred from conserved SNPs from the Myotis myotis CACTUS alignment.
Myotis_cactus_neutral_CASTER.newick - CASTER tree inferred from neutral SNPs from the Myotis myotis CACTUS alignment.
Myotis_multiz_accelerated_CASTER.newick - CASTER tree inferred from accelerated SNPs from the Myotis myotis MULTIZ alignment.
Myotis_multiz_conserved_CASTER.newick - CASTER tree inferred from conserved SNPs from the Myotis myotis MULTIZ alignment.
Myotis_multiz_neutral_CASTER.newick - CASTER tree inferred from neutral SNPs from the Myotis myotis MULTIZ alignment.
Rhino_cactus_accelerated_CASTER.newick - CASTER tree inferred from accelerated SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_cactus_conserved_CASTER.newick - CASTER tree inferred from conserved SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_cactus_neutral_CASTER.newick - CASTER tree inferred from neutral SNPs from the Rhinolophus ferrumequinum CACTUS alignment.
Rhino_multiz_accelerated_CASTER.newick - CASTER tree inferred from accelerated SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhino_multiz_conserved_CASTER.newick - CASTER tree inferred from conserved SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Rhino_multiz_neutral_CASTER.newick - CASTER tree inferred from neutral SNPs from the Rhinolophus ferrumequinum MULTIZ alignment.
Description: Multi-file archive. Dataset_12_5_Annotated_alignments.zip
Per-chromosome and merged FASTA alignments of phyloP-based neutral SNPs from annotated regions across the whole human MultiZ alignment.
This archive contains neutral SNP alignments extracted from annotated genomic regions in the whole human MultiZ alignment. Files are provided separately for each chromosome, together with merged alignments for autosomes, autosomes plus chromosome X, and chromosome X.
human_multiz_annotation_chr1.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 1.
human_multiz_annotation_chr2.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 2.
human_multiz_annotation_chr3.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 3.
human_multiz_annotation_chr4.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 4.
human_multiz_annotation_chr5.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 5.
human_multiz_annotation_chr6.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 6.
human_multiz_annotation_chr7.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 7.
human_multiz_annotation_chr8.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 8.
human_multiz_annotation_chr9.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 9.
human_multiz_annotation_chr10.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 10.
human_multiz_annotation_chr11.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 11.
human_multiz_annotation_chr12.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 12.
human_multiz_annotation_chr13.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 13.
human_multiz_annotation_chr14.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 14.
human_multiz_annotation_chr15.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 15.
human_multiz_annotation_chr16.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 16.
human_multiz_annotation_chr17.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 17.
human_multiz_annotation_chr18.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 18.
human_multiz_annotation_chr19.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 19.
human_multiz_annotation_chr20.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 20.
human_multiz_annotation_chr21.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 21.
human_multiz_annotation_chr22.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome 22.
human_multiz_annotation_chrX.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome X.
human_multiz_annotation_SNPs_autosomes_merged.fasta - Merged FASTA alignment of neutral SNPs from annotated regions across all autosomes.
human_multiz_annotation_SNPs_autosomes_chrX_merged.fasta - Merged FASTA alignment of neutral SNPs from annotated regions across all autosomes and chromosome X.
human_multiz_annotation_SNPs_chrX.fasta - FASTA alignment of neutral SNPs from annotated regions on chromosome X.
Description: Multi-file archive. Dataset_12_6_Dark_alignments.zip
Per-chromosome and merged FASTA alignments of phyloP-based neutral SNPs from non-annotated regions across the whole human MultiZ alignment.
This archive contains neutral SNP alignments extracted from non-annotated regions in the whole human MultiZ alignment. These non-annotated regions correspond to the "dark" genomic regions analysed in the manuscript. Files are provided separately for each chromosome, together with merged alignments for autosomes, autosomes plus chromosome X, and chromosome X.
human_multiz_none_annotation_chr1.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 1.
human_multiz_none_annotation_chr2.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 2.
human_multiz_none_annotation_chr3.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 3.
human_multiz_none_annotation_chr4.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 4.
human_multiz_none_annotation_chr5.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 5.
human_multiz_none_annotation_chr6.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 6.
human_multiz_none_annotation_chr7.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 7.
human_multiz_none_annotation_chr8.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 8.
human_multiz_none_annotation_chr9.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 9.
human_multiz_none_annotation_chr10.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 10.
human_multiz_none_annotation_chr11.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 11.
human_multiz_none_annotation_chr12.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 12.
human_multiz_none_annotation_chr13.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 13.
human_multiz_none_annotation_chr14.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 14.
human_multiz_none_annotation_chr15.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 15.
human_multiz_none_annotation_chr16.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 16.
human_multiz_none_annotation_chr17.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 17.
human_multiz_none_annotation_chr18.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 18.
human_multiz_none_annotation_chr19.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 19.
human_multiz_none_annotation_chr20.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 20.
human_multiz_none_annotation_chr21.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 21.
human_multiz_none_annotation_chr22.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome 22.
human_multiz_none_annotation_chrX.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome X.
human_multiz_dark_SNPs_autosomes_merged.fasta - Merged FASTA alignment of neutral SNPs from non-annotated regions across all autosomes.
human_multiz_dark_SNPs_autosomes_chrX_merged.fasta - Merged FASTA alignment of neutral SNPs from non-annotated regions across all autosomes and chromosome X.
human_multiz_dark_SNPs_chrX.fasta - FASTA alignment of neutral SNPs from non-annotated regions on chromosome X.
Description: Multi-file archive. Dataset_12_7_MLtrees2.zip
Maximum likelihood trees inferred from phyloP-based neutral SNPs from annotated and non-annotated regions across the whole human MultiZ alignment.
This archive contains maximum likelihood phylogenetic trees in Newick format. Trees were inferred from the neutral SNP alignments from annotated regions in Dataset_12_5_Annotated_alignments.zip and from non-annotated, or dark, regions in Dataset_12_6_Dark_alignments.zip. Trees are provided separately for each chromosome and for the merged autosomal alignments.
human_multiz_annotation_SNPs_ML_tree_chr1.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 1.
human_multiz_annotation_SNPs_ML_tree_chr2.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 2.
human_multiz_annotation_SNPs_ML_tree_chr3.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 3.
human_multiz_annotation_SNPs_ML_tree_chr4.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 4.
human_multiz_annotation_SNPs_ML_tree_chr5.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 5.
human_multiz_annotation_SNPs_ML_tree_chr6.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 6.
human_multiz_annotation_SNPs_ML_tree_chr7.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 7.
human_multiz_annotation_SNPs_ML_tree_chr8.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 8.
human_multiz_annotation_SNPs_ML_tree_chr9.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 9.
human_multiz_annotation_SNPs_ML_tree_chr10.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 10.
human_multiz_annotation_SNPs_ML_tree_chr11.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 11.
human_multiz_annotation_SNPs_ML_tree_chr12.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 12.
human_multiz_annotation_SNPs_ML_tree_chr13.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 13.
human_multiz_annotation_SNPs_ML_tree_chr14.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 14.
human_multiz_annotation_SNPs_ML_tree_chr15.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 15.
human_multiz_annotation_SNPs_ML_tree_chr16.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 16.
human_multiz_annotation_SNPs_ML_tree_chr17.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 17.
human_multiz_annotation_SNPs_ML_tree_chr18.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 18.
human_multiz_annotation_SNPs_ML_tree_chr19.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 19.
human_multiz_annotation_SNPs_ML_tree_chr20.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 20.
human_multiz_annotation_SNPs_ML_tree_chr21.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 21.
human_multiz_annotation_SNPs_ML_tree_chr22.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome 22.
human_multiz_annotation_SNPs_ML_tree_chrX.newick - Maximum likelihood tree inferred from neutral SNPs from annotated regions on chromosome X.
human_multiz_annotation_SNPs_ML_tree_autosomes_merged.newick - Maximum likelihood tree inferred from the merged neutral SNP alignment from annotated regions across all autosomes.
human_multiz_dark_SNPs_ML_tree_chr1.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 1.
human_multiz_dark_SNPs_ML_tree_chr2.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 2.
human_multiz_dark_SNPs_ML_tree_chr3.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 3.
human_multiz_dark_SNPs_ML_tree_chr4.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 4.
human_multiz_dark_SNPs_ML_tree_chr5.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 5.
human_multiz_dark_SNPs_ML_tree_chr6.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 6.
human_multiz_dark_SNPs_ML_tree_chr7.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 7.
human_multiz_dark_SNPs_ML_tree_chr8.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 8.
human_multiz_dark_SNPs_ML_tree_chr9.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 9.
human_multiz_dark_SNPs_ML_tree_chr10.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 10.
human_multiz_dark_SNPs_ML_tree_chr11.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 11.
human_multiz_dark_SNPs_ML_tree_chr12.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 12.
human_multiz_dark_SNPs_ML_tree_chr13.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 13.
human_multiz_dark_SNPs_ML_tree_chr14.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 14.
human_multiz_dark_SNPs_ML_tree_chr15.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 15.
human_multiz_dark_SNPs_ML_tree_chr16.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 16.
human_multiz_dark_SNPs_ML_tree_chr17.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 17.
human_multiz_dark_SNPs_ML_tree_chr18.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 18.
human_multiz_dark_SNPs_ML_tree_chr19.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 19.
human_multiz_dark_SNPs_ML_tree_chr20.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 20.
human_multiz_dark_SNPs_ML_tree_chr21.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 21.
human_multiz_dark_SNPs_ML_tree_chr22.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome 22.
human_multiz_dark_SNPs_ML_tree_chrX.newick - Maximum likelihood tree inferred from neutral SNPs from non-annotated regions on chromosome X.
human_multiz_dark_SNPs_ML_tree_autosomes_merged.newick - Maximum likelihood tree inferred from the merged neutral SNP alignment from non-annotated regions across all autosomes.
Description: Multi-file archive. Dataset_12_8_CASTERtrees2.zip
CASTER trees inferred from phyloP-based neutral SNPs from annotated and non-annotated regions across the whole human MultiZ alignment.
This archive contains CASTER phylogenetic trees in Newick format. Trees were inferred from the neutral SNP alignments from annotated regions in Dataset_12_5_Annotated_alignments.zip and from non-annotated, or dark, regions in Dataset_12_6_Dark_alignments.zip. Trees are provided separately for each chromosome and for the merged autosomal alignments.
human_multiz_annotation_SNPs_chr1_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 1.
human_multiz_annotation_SNPs_chr2_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 2.
human_multiz_annotation_SNPs_chr3_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 3.
human_multiz_annotation_SNPs_chr4_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 4.
human_multiz_annotation_SNPs_chr5_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 5.
human_multiz_annotation_SNPs_chr6_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 6.
human_multiz_annotation_SNPs_chr7_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 7.
human_multiz_annotation_SNPs_chr8_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 8.
human_multiz_annotation_SNPs_chr9_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 9.
human_multiz_annotation_SNPs_chr10_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 10.
human_multiz_annotation_SNPs_chr11_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 11.
human_multiz_annotation_SNPs_chr12_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 12.
human_multiz_annotation_SNPs_chr13_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 13.
human_multiz_annotation_SNPs_chr14_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 14.
human_multiz_annotation_SNPs_chr15_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 15.
human_multiz_annotation_SNPs_chr16_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 16.
human_multiz_annotation_SNPs_chr17_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 17.
human_multiz_annotation_SNPs_chr18_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 18.
human_multiz_annotation_SNPs_chr19_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 19.
human_multiz_annotation_SNPs_chr20_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 20.
human_multiz_annotation_SNPs_chr21_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 21.
human_multiz_annotation_SNPs_chr22_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome 22.
human_multiz_annotation_SNPs_chrX_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from annotated regions on chromosome X.
human_multiz_annotation_SNPs_autosomes_CASTER_tree.newick - CASTER tree inferred from the merged neutral SNP alignment from annotated regions across all autosomes.
human_multiz_dark_SNPs_chr1_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 1.
human_multiz_dark_SNPs_chr2_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 2.
human_multiz_dark_SNPs_chr3_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 3.
human_multiz_dark_SNPs_chr4_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 4.
human_multiz_dark_SNPs_chr5_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 5.
human_multiz_dark_SNPs_chr6_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 6.
human_multiz_dark_SNPs_chr7_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 7.
human_multiz_dark_SNPs_chr8_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 8.
human_multiz_dark_SNPs_chr9_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 9.
human_multiz_dark_SNPs_chr10_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 10.
human_multiz_dark_SNPs_chr11_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 11.
human_multiz_dark_SNPs_chr12_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 12.
human_multiz_dark_SNPs_chr13_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 13
human_multiz_dark_SNPs_chr14_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 14.
human_multiz_dark_SNPs_chr15_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 15.
human_multiz_dark_SNPs_chr16_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 16.
human_multiz_dark_SNPs_chr17_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 17.
human_multiz_dark_SNPs_chr18_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 18
human_multiz_dark_SNPs_chr19_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 19
human_multiz_dark_SNPs_chr20_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 20.
human_multiz_dark_SNPs_chr21_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 21.
human_multiz_dark_SNPs_chr22_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome 22.
human_multiz_dark_SNPs_chrX_CASTER_tree.newick - CASTER tree inferred from neutral SNPs from non-annotated regions on chromosome X
human_multiz_dark_SNPs_autosomes_CASTER_tree.newick - CASTER tree inferred from the merged neutral SNP alignment from non-annotated regions across all autosomes.
File: Section_13.zip
Description: Multi-file archive. CASTER trees inferred from phyloP-based neutral SNPs from annotated and non-annotated regions within the X-linked recombination desert region of the human MultiZ alignment.
This archive contains CASTER phylogenetic trees in Newick format. Trees were inferred from neutral SNPs within the X-linked recombination desert (XLRD) region. Separate trees are provided for SNPs from annotated regions and from non-annotated, or dark, regions.
caster_XLRD_anno_SNPs_human_multiz.newick - CASTER tree inferred from neutral SNPs from annotated regions within the X-linked recombination desert region of the human MultiZ alignment.
caster_XLRD_dark_SNPs_human_multiz.newick - CASTER tree inferred from neutral SNPs from non-annotated, or dark, regions within the X-linked recombination desert region of the human MultiZ alignment.
File: Section_14.tgz
Description: bats_mito.tar.gz
Mitochondrial genomes identified from 103 bat genome assemblies.
This archive contains mitochondrial genome sequences identified from 103 bat genome assemblies. These mitochondrial genome sequences were used for downstream mitochondrial genome alignment and phylogenetic analyses.
bats_mito - mitochondrial genome sequences identified from bat genome assemblies.
Description: mito_alignment.tar.gz
Alignment of mitochondrial genomes identified from bat genome assemblies.
This archive contains the mitochondrial genome alignment generated from the mitochondrial genome sequences in bats_mito.tar.gz. The alignment was used for downstream mitochondrial phylogenetic analyses.
mitochondrial_genomes_alignment.fasta - FASTA alignment of mitochondrial genome sequences identified from bat genome assemblies.
File: Section_15.tgz
Description: Multi-file archive. Dataset_S1_DM.xlsx, Dataset_S2_DM.xlsx
Description: Dataset_S1_DM.xlsx
Each column represents:
- original scaffolds generated by Bat1K
- renamed chromosomes (random naming, except Myotis myotis species) based on size distribution
- scaffold size in bp
Description: Dataset_S2_DM.xlsx
Each tab except "Chiroptera" represents alignments of scaffolds of 103 bat species with reconstructed ancestral chromosomes:
- bat species scaffold ID
- start coordinates of each synteny block on bat scaffolds aligned with RACFs named after Chiropteran ancestral chromosomes
- end coordinates of each synteny block on bat scaffolds aligned with RACFs named after Chiropteran ancestral chromosomes
- Chiropteran ancestral chromosome IDs
- Distribution shows how scaffolds align to ancestral chromosomes (0 - whole scaffold aligns to a single ancestral Chiropteran chromoosome, 1,2, etc - scaffold alignes to multiple reconstructed chromosomes)
- start coordinates of each synteny block on a RACF aligned to species scaffold
- end coordinates of each synteny block on a RACF aligned to species scaffold
- block orientation
Chiroptera tab shows RACFs, corresponding ancestral chromosomes and their sizes in bp.
File: Section_16.tgz
Description: Multi-file archive. Morphobank file, Divergence_time_estimates.tar.gz, Morph.zip
From Morph.zip:
File: 21fam.bat1k.MorphMatrix.nex
Description: original Nexus file with all characters and taxa initially assessed before pruning taxa and identifying correlated characters.
File: matrix480.nex
Description: processed Nexus file with only uncorrelated characters and taxa subsequently analyzed using the total evidence fossilized birth-death (FBD) approach.
From Divergence_time_estimates.tar.gz:
Astral_tree:
coding_sequences:
bin_000-bin077 - all bins of 100 coding genes and their concatenation used to infer divergence times based on astral tree (T1)
mean_divtree.tre - tree with mean divergence estimates per node from across all bins
neutral_sites:
perChrom:
Chr<num>_bin<num> - all bins containing randomly assigned genes, split by chromosome
mean_divtree.tre - tree with mean divergence estimates per node from across all bins
random:
bin_000-bin104 - all bins of 100 coding genes and their concatenation used to infer divergence times based on astral tree (T1)
mean_divtree.tre - tree with mean divergence estimates per node from across all bins
Caster_tree:
coding_sequences:
bin_000-bin077 - all bins of 100 coding genes and their concatenation used to infer divergence times based on caster tree (T2)
mean_divtree.tre - tree with mean divergence estimates per node from across all bins
neutral_sites:
perChrom:
Chr<num>_bin<num> - all bins containing randomly assigned genes, split by chromosome
mean_divtree.tre - tree with mean divergence estimates per node from across all bins
random:
bin_000-bin104 - all bins of 100 coding genes and their concatenation used to infer divergence times based on caster tree (T2)
mean_divtree.tre - tree with mean divergence estimates per node from across all bins
File: Section_17.tgz
Description Multi-file archive. FBD.zip
Description: Multi-file archive. T1_21fam_neutral.nex,T1_21fam_neutral.nex, FBD.model.log, FBD.trees, FBD.mcc.tre, FBD.TOGA.model.log, FBD.TOGA.trees, FBD.mcc.tre
From FBD.zip:
File: T1_21fam_neutral.nex
Description: Nexus file for conducting FBD analyses in the RevBayes programming language. This file includes both the filtered morphological characters and neutrally evolving nucleotide sites.
File: FBD.model.log
Description: log file of results of unconstrained FBD analyses including estimated posterior parameters.
File: FBD.trees
Description: file with phylogenies, posterior, likelihood, and prior resulting from unconstrained FBD analyses.
File: FBD.mcc.tre
Description: maximum clade credibility Nexus tree from unconstrained FBD analyses.
File: FBD.TOGA.model.log
Description: log file of results of FBD analyses constrained to the phylogeny resulting from TOGA or T1, including estimated posterior parameters.
File: FBD.TOGA.trees
Description: file with phylogenies, posterior, likelihood, and prior resulting from from FBD analyses constrained to the phylogeny resulting from TOGA or T1.
File: FBD.TOGA.mcc.tre
Description: maximum clade credibility Nexus tree from FBD analyses constrained to the phylogeny resulting from TOGA or T1.
File: Section_18.tgz
Description Multi-file archive. DEC.zip
Description: Multi-file archive. DEC.ranges.nex,DEC_FBDtopology_bats.ase.tre, DEC_TOGAtopology_bats.ase.tre, DEC_TOGAtopology_bats.model.log, DEC_FBDtopology_bats.model.log, DEC_TOGAtopology_bats.stoch.log, DEC_FBDtopology_bats.states.log
From DEC.zip:
File: DEC.ranges.nex
Description: Nexus file for conducting Dispersal-Extinction-Cladogenesis (DEC) analyses in the RevBayes programming language. This file includes the distribution of each tip in the phylogeny.
File: DEC_FBDtopology_bats.ase.tre
Description: Nexus tree that together with the geographical ranges comprise the input for DEC analyses. This is the FBD tree.
File: DEC_TOGAtopology_bats.ase.tre
Description: Nexus tree that together with the geographical ranges comprise the input for DEC analyses. This is the TOGA or T1 tree.
File: DEC_TOGAtopology_bats.model.log
Description: DEC model results, parameters of the DEC model for T1.
File: DEC_FBDtopology_bats.model.log
Description: DEC model results, parameters of the DEC model for the FBD phylogeny.
File: DEC_TOGAtopology_bats.stoch.log
Description: stochastic character mapping of DEC model states, to T1.
File: DEC_FBDtopology_bats.states.log
Description: beginning and end state for each node based on DEC model results for the FBD phylogeny.
Sharing/Access information
Links to other publicly accessible resources relevant to these data:
Section 5
- TOGA source code and documentation: https://github.com/hillerlab/TOGA (commit v.c4bce48)
Data were derived from the following sources:
- Human reference genome hg38 (GRCh38.p12), used as the reference for pairwise genome alignments and gene projection.
- Human GENCODE v38 annotation (Ensembl 104), containing 39,664 input transcripts representing 19,456 protein-coding genes.
- Genome assemblies for the 103 bat species and seven non-chiropteran mammalian outgroups described in the associated manuscript and Supplementary Tables.
Section 9
Data were derived from the following sources:
- Protein-coding gene annotations and orthology relationships generated with TOGA as described in Section 5 of the Supplementary Materials.
- MULTIZ and CACTUS whole-genome multiple alignments described in Section 8 of the Supplementary Materials.
- Conservation and acceleration analyses generated with phastCons and phyloP as described in Section 9 of the Supplementary Materials.
Code/Software
General
https://github.com/Bat1K-21families/21families-analyses
This central repository provides links for repositories associated with:
- Genome assembly
- Annotations including proteins (see below also), transposable elements, and miRNAs
- Alignements including coding genes, multiz, and cactus
- Phylogenomic inference including coalescent-based, whole-genome, chromosome, sliding window, neutral intergenic, and monophyly tests. Also includes methods associated with concatenation, neutral sites, mitocondria, and concordance factors.
- Molecular dating with fossil calibrations
- Fossilized birth-death models (molecular + morphology)
- Reconstructing the ancestral bat karyotype
- Biogeography
Section 5: Protein-coding gene annotations were generated using:
- TOGA v1.0, commit c4bce48: https://github.com/hillerlab/TOGA
- LASTZ v1.04.15, used to generate local pairwise genome alignments between human hg38 and each query assembly.
- axtChain, used to chain local alignments.
- RepeatFiller, used to recover repeat-overlapping alignments missed during the initial alignment procedure.
- chainCleaner, used to improve alignment-chain specificity.
- RepeatModeler, used to generate de novo repeat libraries for individual genome assemblies.
- RepeatMasker v4.0.9, used to soft-mask repetitive sequences before genome alignment.
TOGA used the resulting human-to-query whole-genome alignment chains together with the human GENCODE v38 annotation to infer orthologous protein-coding loci and generate the BED annotations provided in this dataset.
For complete software parameters and methodological details, see Section 5, “Annotation of Protein-Coding Genes,” of the Supplementary Materials accompanying the associated publication.
