Data from: Malus (Rosaceae) systematics and diversification: Insights from a RAD-seq phylogeny
Data files
Jul 15, 2026 version files 2.69 GB
-
m15_Malus_20211215x6387_refgen.loci
611.91 MB
-
m15_Malus_20211215x6387_refgen.phy
903.80 MB
-
m15_Malusx6387_20241010.loci
400.48 MB
-
m15_Malusx6387_20241010.phy
556.26 MB
-
m15_Malusx6387_20241010.snps.hdf5
212.90 MB
-
Malus-TreeMix-JupyterNotebookscript-LindseyMix-Final.ipynb
4.08 MB
-
RAxML_trMalusall20211215x6387_v2.1_v1.tre
6.56 KB
-
RAxML_trMalusx6387_20241010.1_v1.tre
7.23 KB
-
README.md
3.48 KB
Abstract
Phylogenies were estimated using maximum likelihood on the concatenated RAD-seq loci in RAxML version 8.2.4 (Stamatakis 2014), employing rapid bootstrap analysis and search for the best-scoring ML tree in one program run (-f a) with 20 random-start trees, using the GTRCAT approximation for the GTR+gamma model. Non-parametric bootstrapping was used to estimate clade support. Two phylogenies were generated: 1) for the 94 individuals that passed quality checks, and 2) a phylogeny of 64 individuals that excluded hybrids and cultivars. To investigate the possible effects of reticulate evolution in Malus, we used TreeMix (Pickrell and Pritchard 2012) as implemented in ipyrad v 0.9.105 (Eaton and Overcast 2020). TreeMix uses SNP frequencies to reconstruct population trees under maximum likelihood. The method ignores mutations, which may introduce imprecision in phylogenetic contexts, but it is computationally efficient and has been effectively used for macroevolutionary questions in previous studies (Mason et al. 2019; Cui et al., in review). We inspected nine trees from each TreeMix analysis, inferred from nine random subsamples of one SNP per locus.
lworcester@mortonarb.org (2026-05-26)
Roalson, Eric H., Lindsey Worcester, and Andrew L. Hipp, in rev. Malus (Rosaceae) systematics and diversification: insights from a RAD-seq phylogeny. Systematic Botany.
Description of data and files
The scope of this project was to understand relationships between species, hybrids, and cultivars as sampled for this project using standard phylogenomic methods (Maximum Likelihood) to generate the trees (full dataset and reduced dataset - without hybrids and cultivars). Analyses were used to evaluate the phylogenetic effects of reticulation using TreeMix (script listed below) and F-Branch in Dsuite (PDF only for the F-branch analysis; see Supplemental Fig. 2).
Data Files:
m15_Malus_20211215x6387_refgen.loci — data matrix of nucleotide data assembled in ipyrad. All genomic data analyzed in this paper are assembled into the loci file before processing into alignments (.phy). This is the full dataset (95 samples).
m15_Malus_20211215x6387_refgen.phy — concatenated DNA matrix generated from the loci file for analysis in RAxML (full dataset).
m15_Malusx6387_20241010.loci — data matrix subset from m15_Malus_20211215x6387_refgen.loci, excluding cultivars, hybrids, and rogue taxa (reduced data set).
m15_Malusx6387_20241010.phy — concatenated DNA matrix generated from the loci files for analysis in RAxML (reduced dataset).
m15_Malusx6387_20241010.snps.hdf5 — database file that includes genotype calls and linkage information for subsampling unlinked SNPs and bootstrap resampling (per the ipyrad TreeMix tutorial).
Tree Files:
RAxML_trMalusall20211215x6387_v2.1_v1.tre — tree file generated from RAxML and labeled in R with species names. This file corresponds to Fig. 1 in the Malus manuscript.
RAxML_trMalusx6387_20241010.1_v1.tre — tree file from the reduced dataset generated from RAxML, and labeled in R with species names. This file corresponds to Fig. 2 in the Malus manuscript.
The scripts for using TreeMix were written in and Python in a Jupyter Notebook in the file "Malus-TreeMix-JupyterNotebookscript-LindseyMix-Final.ipynb". All of the scripts came from Deren Eaton's "ipyrad-analysis toolkit: treemix" (https://ipyrad.readthedocs.io/en/master/API-analysis/cookbook-treemix.html) tutorials. This script is modified for the best fit for the Malus dataset for the various analyses.
Data:
The data file used was the SNPs.hdf5 file, generated in RAxML as the base dataset for Malus: /home/lindsey/Documents/Analyses/m15_Malus/Malus_ms_files/m15_Malusx6387_20241010_outfiles/m15_Malusx6387_20241010.snps.hdf5
Analyses:
The best fit for the data was found for each analysis by using the "1. Finding the best value for m" instructions with the Malus dataset. Analyses were iterated over different subsamples of SNPs to see how the initial resulting tree was supported, or if there were closely supported alternative results.
TreeMix was used to test for reticulation between different sections for the sectional analysis.
Other analyses on all the samples in the dataset were condensed by looking at where the samples of the same species formed a clade, and using one individual to represent the species for that clade. Analyses are shown with one outgroup per analysis.
