Tree of Archaeplastida as a showcase for systematic artifacts in nuclear and plastid phylogenomics
Data files
Aug 03, 2026 version files 67.43 MB
-
README.md
9.91 KB
-
Suppl_Data_ver260724.tar.gz
67.42 MB
Abstract
Archaeplastida is defined as a taxonomic assemblage comprising three sub-clades, namely Chloroplastida, Glaucophyta, and Rhodophyta plus two non-photosynthetic lineages sister to Rhodophyta (collectively termed “Rhodozoa” here). Recent phylogenomic analyses stably recovered the monophyly of Archaeplastida, but uncertainty remains in the relationship among the three sub-clades in this assemblage. The phylogenomic analyses of nucleus-encoded proteins (nuc-proteins) grouped Chloroplastida and Glaucophyta together, excluding Rhodozoa in the Archaeplastida clade, albeit the union of Chloroplastida and Rhodophyta was often inferred from the phylogenomic analyses of plastid-encoded proteins (pld-proteins). Our detailed analyses of nuc-protein and pld-protein supermatrices revealed that taxon sampling can invoke different types of phylogenetic artifacts into the inferences from both supermatrices examined here. In the end, we propose a working hypothesis for the ToA and provide future perspectives toward phylogenomics aiming to resolve the global tree of eukaryotes.
Dataset DOI: 10.5061/dryad.b2rbnzssh
Description of the data and file structure
After extracting the tar.gz file "Suppl_Data_ver260724.tar.gz", you will find a directory named "Supp_Data_v4". In this directory, there are two subdirectories named "Nuclear_analyses" and "Plastid_analyses". The structure and contents of the two subdirectories are given below.
NOTE:
- Here, we write “Δ” as “D” and “*” as “s” to use only ASCII in the filename.
- *.log files are the IQ-TREE analysis log files, each corresponding to the treefile with the same base filename.
Nuclear_analyses/
All the files in this directory are related to the phylogenomic analyses based on 317 or 276 nuclear proteins.
ML-Bayes_analyses/
- *.fasta: The supermatrices for ML and Bayesian tree search analyses. Please refer to Table 1 in the paper for details.
- *_ML.treefile: The ML trees in the Newick format, each corresponding to the supermatrix with the same base filename. If multiple support values are provided as node labels, the order is as follows: the value from the ultrafast bootstrap approximation (UFBP), the value from the maximum likelihood bootstrap analysis (MLBP), and the Bayesian posterior probabilities (BPP) from left to right.
- *_Bayes.treefile: The Bayesian consensus tree with BPPs, based on two chains. Each treefile corresponds to the supermatrix with the same base filename.
- *_Bayes_Chain1only.treefile: The Bayesian consensus tree with BPPs, but based only on chain1.
- *_Bayes_Chain2only.treefile: The Bayesian consensus tree with BPPs, but based only on chain2.
- *_single_fasta/: The directory contains single-protein fasta files for the phylogenomic analyses conducted in the study. There are two files for each single-protein, with and without trimming the positions inappropriate for phylogeny (e.g., ambiguously aligned positions), namely the files named as “_trimed.fasta” and those named as “_mafft.fasta”. By concatenating the trimmed files, you can generate two nuclear protein supermatrices, the nuc317 or nuc276.
AUtest_trees/
- *_3trees.treefile: The set of test tree topologies that were used in the AU tests. Each file was prepared from the ML tree of the corresponding supermatrix, as indicated by the filename.
FPR_analyses/
- nuc276_site-rate.tsv: The site rates calculated from nuc276 using IQ-TREE's
--rateoption. The meaning of each column is as follows:- "Site": The site number in the nuc276 supermatrix.
- "Rate": Site-specific evolutionary rate. These values are used in this study for fast-evolving position removal (FPR) analyses.
- "Cat": The category of the site in the G4 model (4-category discrete gamma model).
- "C_Rate": The evolutionary rate of each category.
- nuc276-ALLinONE_FPR-alignments.nxs: nuc276 in the NEXUS format. The top 20%, 40%, 60%, and 80% fastest-evolving positions are stored as
CHARSET. By combining the exclusion of the character set of interest and that of Rhodelphidia and/or Picozoa, all the supermatrices analyzed in this study can be restored. - *.treefile: Each treefile in the Newick format contains the four ML trees inferred from supermatrices after excluding the top 20%, 40%, 60%, and 80% fastest-evolving positions, which correspond to tree #1, tree #2, tree #3, and tree #4, respectively.
- LG+C60+F+G_summary_node-and-clade.ods: The raw UFBP values plotted in the paper. Each sheet corresponds to a supermatrix, and the meaning of each column is as follows:
- "delete_per": The percentage of fastest-evolving positions removed.
- "Rhodo+Outg": The number of bootstrap trees that support the Rhodo+Outg topology, satisfying the following conditions: rhodophytes formed a clade; other archaeplastid species (chloroplastids and glaucophytes) formed a clade; and non-archaeplastid species formed a clade (i.e., outgroup).
- "Glauco+Outg": The number of bootstrap trees that support the Glauco+Outg topology.
- "Chloro+Outg": The number of bootstrap trees that support the Chloro+Outg topology.
- "Archae-mono": The number of bootstrap trees that support the monophyly of the Archaeplastida clade.
- "CAM": The number of bootstrap trees that support the monophyly of the CAM clade (Archaeplastida + Pancryptista).
RPR_analyses/
- nuc276-ALLinONE_RPR-alignments.nxs: This NEXUS file is the same as nuc276-ALLinONE_FPR-alignments.nxs, except for
CHARSET, which stores the sets of the randomly selected positions excluded in the random position removal (RPR) analyses. - *.treefile: Each treefile in the Newick format contains the 20 ML trees inferred after excluding 20%, 40%, 60%, or 80% of random positions.
- LG+C60+F+G_RSR_summary_node-and-clade.ods: The raw UFBP values plotted in the paper. Each sheet corresponds to a supermatrix, and the meaning of each column is as follows:
- "delete_per": The percentage of positions removed.
- "trial_num": The trial number (1–20) at each deletion percentage, corresponding to the order of the
CHARSETblocks in the nexus file and the 20 ML trees in the tree files. - "Rhodo+Outg": The number of bootstrap trees that support the Rhodo+Outg topology.
- "Glauco+Outg": The number of bootstrap trees that support the Glauco+Outg topology.
- "Chloro+Outg": The number of bootstrap trees that support the Chloro+Outg topology.
- "Archae-mono": The number of bootstrap trees that support the monophyly of the Archaeplastida clade.
- "CAM": The number of bootstrap trees that support the monophyly of the CAM clade.
- RPR-UFBP_Wilcoxon-signed-rank-test.ods: The p-values of the Wilcoxon signed-rank test on the difference between the UFBPs based on the RPR-processed nuc276ΔP and those based on the RPR-processed nuc276ΔR, nuc276ΔP*, or nuc276ΔP**. Each sheet corresponds to a supermatrix, and the meaning of each row is as follows:
- "delete_per": The percentage of positions removed.
- "Archae-mono": The p-value from comparing the distribution of the number of bootstrap trees supporting the Archaeplastida monophyly between nuc276ΔP and the supermatrix indicated by the sheet name.
- "Rhodo+Outg": The p-value from comparing the distribution of the number of bootstrap trees supporting the Rhodo+Outg topology between nuc276ΔP and the supermatrix indicated by the sheet name.
- "Glauco+Outg": The p-value from comparing the distribution of the number of bootstrap trees supporting the Glauco+Outg topology between nuc276ΔP and the supermatrix indicated by the sheet name.
Plastid_analyses/
All the files in this directory are related to the phylogenomic analyses based on 54 plastid proteins.
ML-Bayes_analyses/
The naming schemes for files in this directory are the same as those for files in Nuclear_analyses/ML-Bayes_analyses/.
AUtest_trees/
The naming schemes for files in this directory are the same as those for files in Nuclear_analyses/AUtest_trees/.
FPR_analyses/
The naming schemes for files in this directory are the same as those for files in Nuclear_analyses/FPR_analyses/.
RPR_analyses/
The naming schemes for files in this directory are the same as those for files in Nuclear_analyses/RPR_analyses/.
RateRatioRemoval_analyses/
- comp_rate_ratio_pld54-pld54Dgsrs.tsv: The list of site rates calculated from pld54, those calculated from pld54Δgsrs, and the ratios of the two corresponding site rates. The meaning of each column is as follows:
- "site": The site number in the supermatrix.
- "i1_rate(pld54)": Site-specific evolutionary rate. These values are the same as values in
../FPR_analyses/pld54_site-rate.tsv. - "i2_rate(pld54Dgsrs)": Site-specific evolutionary rate. These values are the same as values in
../FPR_analyses/pld54Dgsrs_site-rate.tsv. - "rate ratio(i1 / i2)": The values are calculated as "i1_rate(pld54)" / "i2_rate(pld54Dgsrs)".
- pld54_RateRatioRemoval-alignments.nxs: pld54 in the NEXUS format with sets of the positions associated with the 5% and 10-90% largest site-rate ratios stored as
CHARSET. By excluding the positions in one of the character sets, the supermatrices used in the analyses to produce Figs. 6a-c. - pld54_RateRatioRemoval-BothTop5per_ML_FigS15.fasta: The fasta-formatted pld54 after removing the alignment positions with the 5% largest and 5% smallest site-rate ratios.
- *.treefile: The ML trees in the Newick format inferred from the modified pld54 after the exclusion of the positions according to the site-rate ratios.
- cpREV+C60+F+I+G_RateRatioRemoval_summary_node-and-clade.ods: The raw UFBP values plotted in Figs. 6b and c. The meaning of each column is the same as in the FPR and RPR analyses.
WitChi_analysis/
- pld54_wasserstein_s1_pruned.fasta: The supermatrices that pld54 after removing 3,152 positions bearing potential bias in amino acid composition.
- pld54_wasserstein_s1_pruned.fasta.treefile: The ML trees in the Newick format inferred from the above supermatrix.
- Other files: Log files by WitChi. You can trace which positions were removed as bearing potential bias in amino acid composition.
GHOST_model_analyses/
- FigS16_pld54_optlen-BFGS_H4_teTree2.treefile: The trees corresponding to each category of the H4 model (the GHOST model with four categories). These trees were fitted while calculating the log-likelihood of the Glauco+Outg topology in order to evaluate the model's performance.
- Figs17_pld54_optlen-BFGS_H8_teTree1.treefile: The trees corresponding to each category of the H8 model (the GHOST model with eight categories). These trees were fitted while calculating the log-likelihood of the Rhodo+Outg topology in order to evaluate the model's performance.
