Data and code from: The Y chromosome is a reliable marker for deep-time phylogenetic inference
Data files
Jul 21, 2026 version files 693.87 MB
-
README_alignments.txt
992 B
-
README_alternativeTopology.txt
899 B
-
README_cactus.txt
661 B
-
README_coalescence.txt
483 B
-
README_concatenation.txt
932 B
-
README_gene_conversion.txt
768 B
-
README_identif-Y-er.txt
709 B
-
README_phyloP.txt
778 B
-
README.md
3.50 KB
-
S1(cactus).zip
421.71 MB
-
S2(alignments).zip
24.73 MB
-
S3(phylop).zip
4.86 MB
-
S4(concatenation).zip
84.75 KB
-
S5(coalescence).zip
15.66 KB
-
S6(alternative_topology).zip
6.68 KB
-
S7(gene_conversion).zip
235.21 KB
-
S8(identif-Y-er).zip
242.22 MB
Abstract
The Y chromosome has been omitted from almost all phylogenomic analyses to date, with its complex structure making it difficult to assemble accurately. Yet, the Y chromosome is, theoretically, an optimal phylogenetic marker. To overcome difficulties associated with using the Y chromosome, we developed a novel approach to identify Y-linked sequences in placental mammals, allowing us to generate a 63-species Y-chromosome alignment. This alignment allowed us to test the robustness and accuracy of the Y-chromosome applied in a comparative phylogenetic context. Our results demonstrate that not only can the Y chromosome be aligned across species divergences spanning more than 100 million years, but that it performs robustly in a comparative phylogenetic context. Here, we have included all of the data necessary to replicate our study.
Dataset DOI: 10.5061/dryad.fttdz095p
Description of the data and file structure
This repository contains the necessary data required to replicate analyses in Alexander, Foley, and Murphy (in review). This manuscript evaluated whether the mammalian Y-chromosome can be used for the inference of deep-time phylogenetic relationships.
Files and variables
All repositories contain their own README file, which describes the files that are available in each compressed folder.
S1(cactus).zip - input genomes and alignments generated using the software Progressive Cactus
S2(alignments).zip - all filtered and subsetted alignments used to generate trees
S3(phylop).zip - accelerated, neutral, and conserved quintile ranges and phyloP score files
S4(concatenation).zip - all of the maximum-likelihood phylogenies generated under concatenation using IQ-TREE
S5(coalescence).zip - trees generated under multi-species coalescence using CASTER
S6(alternative_topology).zip - hypotheses used to perform alternative tree tests using IQ-TREE
S7(gene_conversion).zip - data related to the gene conversion analysis
S8(identif-Y-er).zip - contains the data generated using our novel pipeline, Identi-Y-er, which identifies Y-linked scaffolds in scaffolded genome assemblies.
Code & Software
All code used for this study can be accessed either through GitHub or Zenodo.
Each folder contains a README that states the purpose of the script(s). Folders are named after their respective section in the manuscript's methods:
- data collection
- sex-linked sequence identification
- alignments
- measuring evolutionary conservation of sites
- phylogenetic reconstruction
- alternative tree topology test
- identification of gene conversion
Dependencies
Most scripts use basic Python packages (argparse, os) or R packages; therefore, the scripts do not require special dependencies or installation instructions. I do not provide conda environment files for most scripts, but dependencies are clearly listed at the start of the code.
All code was ran on a personal MacOS M1 Pro or the Texas A&M HPRC cluster. Hence, some scripts are written to be executed on a SLURM cluster.
Citation
If you use any of this code, please cite both the following paper and Zenodo repository:
Alexander, E.P., Foley, N.M., and Murphy, W.J. Accepted. "Using the mammalian Y-chromosome for deep-time phylogenetic inference." Systematic Biology.
Support
If you need assistance with navigating this repository, running any of these scripts, or just general questions, please don't hesitate to reach out via email or create an issue on GitHub.
