Data and code from: Association with dynamic habitats disrupts speciation and reduces the efficacy of purifying selection in flightless beetles
Data files
Jul 20, 2026 version files 3.52 GB
-
1_orthofinder.zip
537.25 MB
-
2_pipeline.zip
39.49 KB
-
3_final_dataset.zip
32.91 MB
-
4_analyses.zip
2.95 GB
-
README.md
98.47 KB
Abstract
Association with temporally dynamic habitats has been proposed as a strong predictor of effective population size (Ne) and speciation rates, and thus also potentially of molecular evolution rates, but empirical insights remain limited. We use a genomic dataset of over 2,000 single-copy orthologs to evaluate how habitat association shapes molecular evolution in closely related flightless beetle lineages (Coleoptera: Tenebrionidae: Eutagenia), co-distributed across Eastern Mediterranean islands. The focal taxa occupy either dynamic coastal sand dunes or stable compact-soil habitats but share uniform life-history traits and morphology. Species delimitation analyses identified a single widespread coastal dune species, whereas the stable-habitat clade has diversified into nine allopatric species across the same geographic space and timeframe. Despite its wide distribution, the dynamic-habitat lineage exhibits the lowest long-term Ne and accordingly the highest mean nonsynonymous-to-synonymous substitution rate ratio (dN/dS), yet its total substitution rate is not significantly higher. We propose that recurrent local population extinction in dynamic habitats hinders the completion of speciation and reduces the efficacy of purifying selection, thus elevating the dN/dS ratio. More broadly, these findings indicate that diversification and molecular evolution rates may become decoupled under recurrent population turnover, a hypothesis that requires further empirical and theoretical investigation.
REPOSITORY INFORMATION
Dataset overview
A detailed description of the general framework and specific methodology that were followed in order to generate and process all included files can be found in the relevant publication listed below.
This dataset was generated using a whole-genome sequencing (WGS) approach. Briefly, 59 WGS libraries were constructed, one for each specimen with adequate DNA yield (at least 200 ng) and fragmentation level (at least 5.8 DIN value), from 46 localities across the Eastern Mediterranean. Libraries were prepared using the Illumina TruSeq DNA Nano prep kit (350 bp insert size) and sequenced on the Illumina NovaSeq6000 platform (2 × 150 bp paired-end reads), with target coverage of 20X. Additionally, a previously sequenced specimen (WGS; Illumina MiSeq platform; 2 × 250 bp paired-end reads; 550 bp insert size; ~40X coverage) was provided by the NHM (London, UK). Raw Illumina reads and relevant metadata have been deposited in the NCBI Sequence Read Archive (SRA), under BioProject PRJNA1399320 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA1399320). Assembly of single-copy protein-coding genes was performed using a custom local assembly pipeline approach, as detailed in the relevant publication.
Corresponding author information
Name: Emmanouil Meramveliotakis
ORCID: https://orcid.org/0000-0002-6399-575X
Affiliation: Department of Biological Sciences, Faculty of Pure and Applied Sciences, University of Cyprus, Nicosia, Cyprus
email_1: emeram01@ucy.ac.cy
email_2: m.meramveliotakis@gmail.com
Alternative contact information
Name: Anna Papadopoulou
ORCID: https://orcid.org/0000-0002-4656-4894
Affiliation: Department of Biological Sciences, Faculty of Pure and Applied Sciences, University of Cyprus, Nicosia, Cyprus
email: papadopoulou.g.anna@ucy.ac.cy
Related publication
Meramveliotakis, E., Ntatsopoulos, K., Vogler, A. P., Papadopoulou, A. (ACCEPTED; Jul-2026) Association with dynamic habitats disrupts speciation and reduces the efficacy of purifying selection in flightless beetles. Evolution
Funding information & facility acknowledgements
This work was funded by the project EXCELLENCE/0421/0419 - EVOLHAB: A genomic approach to evaluate the effects of habitat instability on the evolutionary dynamics of insular lineages. The project EXCELLENCE/0421/0419 was co-financed by the European Regional Development Fund and the Republic of Cyprus through the Research and Innovation Foundation (RIF).
Computational resources were provided by the High Performance Computing facility of the University of Cyprus (UCY HPC).
GENERAL NOTES
Scripts and paths
The majority of scripts, especially the SLURM batch scripts, contain absolute file or directory paths. These paths have been replaced with the /ABSOLUTE_PATH/ string and should be modified before running any analyses on another system. The scripts are intended to document the analytical workflow and may require software installation and/or modification for another HPC environment or directory structure.
Most computational scripts are batch/shell scripts (.slurm, .sh), Python scripts (.py), or R scripts (.R, .r). SLURM batch scripts were designed for an HPC system using the SLURM workload manager.
Repeated analyses
Some analyses were performed both on the complete dataset and on the Metazoa-level BUSCO subset (mzlbusco). For the case of molecular evolution analyses, when applicable, commands relevant to the mzlbusco subset are not included if they are identical to the commands for the complete dataset.
File formats and general notations
In the file description tables below, entries ending with / indicate directories rather than individual files. These directory entries are used to group large sets of repeated files and subdirectories, or to represent typical software output directories. Entries containing * are wildcard patterns that represent multiple files with the same extension or naming structure. Terms in square brackets, such as [assembly], [taxon], [batch], [dataset], [run], or [locus], are placeholders describing repeated naming components. For directory entries, the File type column is listed as Multiple file types when the directory contains more than one type of file; otherwise, it lists the single file type present in that directory. The Open with column lists representative ways to open or work with the corresponding file(s) or grouped directory contents.
| Extension / format | Description | Software required to open or use |
|---|---|---|
.fasta, .faa, .fna |
FASTA sequence files or multiple sequence alignments; nucleotide or amino-acid plain sequences or multiple sequence alignments | Text editor, sequence editor, Unix shell, R, Python, alignment software and/or viewers |
.treefile, .tre, .nw, .nex, .trees |
Phylogenetic tree files, usually Newick or NEXUS format | Text editor, tree viewers such as FigTree, Unix shell, R, Python |
.tsv, .txt, .csv |
Plain text, tabular or log files | Text editor, Unix shell, R, Python, spreadsheet software |
.phy, .nexus, .xml, .ctl, .ctr |
Analysis input / control files | Relevant analysis software such as BPP, BEAST2, PAUP*, PAML/MCMCTree, or text editor, Unix shell, R, Python |
.json |
Structured result files | Text editor, Unix shell, R, Python, or .json dedicated tools like jq. |
.pdf |
Figures or reports | PDF viewer |
.R, .r, .py, .sh, .slurm |
Analysis scripts | Text editor, Unix shell, R, Python, and/or SLURM to execute |
Frequently used abbreviations and codes
| Code | Interpretation |
|---|---|
WGS |
Whole-genome sequencing |
USCO |
Universal single-copy ortholog |
MSA |
Multiple sequence alignment |
ML |
Maximum Likelihood |
BI |
Bayesian Inference |
mzlbusco |
Metazoa-level BUSCO subset |
all_sites |
Analyses using complete alignments |
4fold_sites |
Analyses using only four-fold degenerate sites |
PSA_EM |
Psammophilous Eastern Mediterranean Eutagenia lineage |
GEN_CY |
Generalist/eurytopic Cyprus Eutagenia lineage |
GEO_* |
Geophilous Eutagenia lineage(s). Specific suffixes correspond to different lineages based on the species delimitation analyses as explained in the relevant manuscript |
STENOSIS |
Stenosis sp. outgroup taxon used in phylogenomic analyses |
ARCHIVE OVERVIEW
| Archive | Main contents | Relevance |
|---|---|---|
1_orthofinder.zip |
OrthoFinder outputs, USCO reference files, bait sequences, BUSCO-related files, and scripts relevant to the in silico bait-selection step | Identification of Tenebrionidae single-copy orthologs used as in silico baits |
2_pipeline.zip |
SLURM batch scripts and helper scripts for the custom local assembly pipeline | Processing WGS reads, local assembly of bait loci, consensus calling, alignment, filtering, gene tree inference, and final dataset generation |
3_final_dataset.zip |
Final curated MSAs, gene trees, locus statistics, and mzlbusco subset |
Core dataset(s) used as input for downstream analyses |
4_analyses.zip |
Input / output files, scripts, and summarized results for manuscript analyses | Species delimitation, species tree inference, effective population size estimation, functional annotation of assembled protein-coding genes, dN/dS ratio estimation, and substitution rate analyses |
1_orthofinder.zip
The compressed master directory orthofinder contains a total of 77 directories and 160,939 files. It includes the input/output files and relevant scripts for the bait-selection step, as implemented using OrthoFinder and custom scripts. The target clade for bait extraction was Tenebrionidae, corresponding to clade N7 in the OrthoFinder species tree.
Sub-directories
-
0_data
Contains the reference taxon input files and the main OrthoFinder output directory used for the bait-selection step. The included files document the reference taxa used in the analysis, the primary transcripts used as OrthoFinder input, and the final OrthoFinder output directories retained for downstream USCO identification. The OrthoFinder
WorkingDirectory/, which contains large intermediate files used for restarting or modifying the run, was excluded from the repository.Naming conventions:
Primary transcript input files are named using the pattern
[assembly]_[taxon].fafor nucleotide coding sequences and[assembly]_[taxon].faafor amino-acid sequences, where[assembly]is the genome assembly accession and[taxon]is the reference taxon name. Most files insideorthofinder_run/follow the native OrthoFinder naming scheme. Orthogroup-specific files use the prefixOG[orthogroup_number], where eachOGidentifier corresponds to one orthogroup inferred by OrthoFinder. Gene tree files follow the patternOG[orthogroup_number]_tree.txt, orthogroup sequence and alignment files follow the patternOG[orthogroup_number].fa, pairwise orthologue tables follow the pattern[taxon_A]__v__[taxon_B].tsv, and node-specific hierarchical orthogroup files follow the patternN[node_number].tsv, whereNrefers to a node in the OrthoFinder species tree.File(s) / Directory File type Open with Role Description list.ref.taxa.tenebrionidae.txtPlain text file Text editor, Unix shell, R, Python Input helper file for get.uscos.N7.teneb.shList of reference Tenebrionidae taxa used for the OrthoFinder analysis and subsequent USCO retrieval. primary_transcripts_cds/[assembly]_[taxon].famulti-FASTA nucleotide files Text editor, sequence viewer, Unix shell, R, Python Potential OrthoFinder input Primary transcript coding sequences for the reference Coleoptera assemblies that can be used as OrthoFinder input. primary_transcripts_prot/[assembly]_[taxon].faamulti-FASTA amino-acid files Text editor, sequence viewer, Unix shell, R, Python OrthoFinder input Primary transcript protein sequences for the reference Coleoptera assemblies. These files were used as the main sequence input for OrthoFinder in this study. orthofinder_run/Multiple file types N/AMain OrthoFinder output directory Final OrthoFinder output directory retained for the bait-selection workflow. Refer to the OrthoFinder documentation for detailed description of output files and directories. The large WorkingDirectory/intermediate directory is not included. -
1_info
Contains helper lists and summary tables used to identify, inspect, and filter Tenebrionidae universal single-copy orthologs (USCOs) from the OrthoFinder results. These files document (a) which orthogroups were retained as candidate Tenebrionidae-level USCOs, (b) which were excluded because of duplication, repeated assignment, or xenology-related filters, and (c) which USCOs were recovered for each selected reference taxon.
Naming conventions:
Files containing
N7.tenebrionidaerefer to the OrthoFinder species tree node used here to represent the Tenebrionidae-level comparison. Files beginning withlist.usco.[assembly]_[taxon]contain taxon-specific USCO lists for individual reference assemblies.File(s) / Directory File type Open with Role Description blacklist.N7.tenebrionidae.txtPlain text file Text editor, Unix shell, R, Python Helper file for filtering List of orthogroups excluded from the Tenebrionidae-level USCO set during bait selection. list.duplications.N7.tenebrionidae.txtPlain text file Text editor, Unix shell, R, Python Helper file for filtering List of orthogroups associated with duplication events at the Tenebrionidae-level node. Used to exclude duplicated candidates from the single-copy ortholog set. list.repeated.OGs.N7.tenebrionidae.txtPlain text file Text editor, Unix shell, R, Python Helper file for filtering List of repeated orthogroup identifiers detected during the USCO selection procedure, to be excluded from the single-copy ortholog set. list.usco.[assembly]_[taxon].txtPlain text files Text editor, Unix shell, R, Python Taxon-specific USCO lists Lists of USCO identifiers recovered for each selected reference taxon. File names include the genome assembly accession and taxon name. list.usco.N7.tenebrionidae.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Summary table Tabular list of the Tenebrionidae-level USCOs retained from each Tenebrionidae representative. list.xenologs.txtPlain text file Text editor, Unix shell, R, Python Helper file for filtering List of orthogroups flagged as xenolog candidates to be excluded from the single-copy ortholog set. table.usco.N7.tenebrionidae.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Summary table Main summary table describing the final Tenebrionidae-level USCO set used for downstream bait retrieval, including the relevant OrthoFinder columns / information. -
2_usco_references
Contains the selected Tenebrionidae-level USCO reference sequence files for the
N7node. For each selected reference taxon, both amino-acid and nucleotide sequence files are included. These FASTA files were generated byget.uscos.N7.teneb.shand were retained as potential reference USCO datasets for the assembly of single-copy protein-coding genes.Naming conventions:
Files follow the pattern
[assembly]_[taxon]_usco.[extension], where[assembly]is the genome assembly accession,[taxon]is the reference taxon name, anduscoindicates that the sequences correspond to Tenebrionidae universal single-copy orthologs. Files ending in.faacontain amino-acid sequences, and files ending in.fnacontain nucleotide sequences.File(s) / Directory File type Open with Role Description N7/[assembly]_[taxon]_usco.faamulti-FASTA amino-acid files Text editor, sequence viewer, Unix shell, R, Python Reference Tenebrionidae USCO sequence files Amino-acid sequences of the selected USCOs for each reference taxon. N7/[assembly]_[taxon]_usco.fnamulti-FASTA nucleotide files Text editor, sequence viewer, Unix shell, R, Python Reference Tenebrionidae USCO sequence files Nucleotide sequences corresponding to the selected USCOs for each reference taxon. -
3_baits
Contains the final bait sequence files generated from the selected Tenebrionidae-level USCOs. Files are organized by reference taxon. For each reference taxon, the directory contains coding sequence and protein multi-FASTA files, together with per-locus coding sequence and protein FASTA files. The
GCA_004193795.1_Asbolus_verrucosus/busco/directory contains the BUSCO input, logs, summaries, and BUSCO-classified sequence files used to identify the Metazoa-level BUSCO subset of baits. Thebusco/directory only exists for theGCA_004193795.1_Asbolus_verrucosusassembly, as this was the selected reference in this study.Naming conventions:
The top-level reference directories are named using the pattern
[assembly]_[taxon], where[assembly]is the genome assembly accession and[taxon]is the reference taxon name. Taxon-level multi-FASTA bait files follow the patterns[assembly]_[taxon]_usco_cds.fasta,[assembly]_[taxon]_usco_prot.fasta, and[assembly]_[taxon]_usco_prot.stop.fasta. Per-locus FASTA bait files insidecds/andprotein/use the patternOG[orthogroup_number]_[protein_or_transcript_accession].fasta, where theOGidentifier refers to the OrthoFinder orthogroup and the accession identifies the reference sequence.File(s) / Directory File type Open with Role Description [assembly]_[taxon]/[assembly]_[taxon]_usco_cds.fastamulti-FASTA nucleotide file Text editor, sequence viewer, Unix shell, R, Python Combined bait file Coding sequence multi-FASTA file containing the selected USCO bait sequences for the corresponding reference taxon. [assembly]_[taxon]/[assembly]_[taxon]_usco_prot.fastamulti-FASTA amino-acid file Text editor, sequence viewer, Unix shell, R, Python Combined bait file Protein multi-FASTA file containing the selected USCO bait sequences for the corresponding reference taxon. [assembly]_[taxon]/[assembly]_[taxon]_usco_prot.stop.fastamulti-FASTA amino-acid file Text editor, sequence viewer, Unix shell, R, Python Combined bait file Protein multi-FASTA file with added stop codons at the end of the protein sequences, as implemented in get.uscos.N7.teneb.sh[assembly]_[taxon]/cds/FASTA nucleotide files Text editor, sequence viewer, Unix shell, R, Python Per-locus bait directory Directory containing individual coding sequence FASTA files for each selected USCO bait locus. [assembly]_[taxon]/protein/FASTA amino-acid files Text editor, sequence viewer, Unix shell, R, Python Per-locus bait directory Directory containing individual protein FASTA files for each selected USCO bait locus, including versions with and without terminal stop codon. GCA_004193795.1_Asbolus_verrucosus/stats.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Summary table Summary statistics for the generated bait files in the GCA_004193795.1_Asbolus_verrucosusreference directory.GCA_004193795.1_Asbolus_verrucosus/busco/Multiple file types N/ABUSCO input/output directory Directory containing the analysis to identify the Metazoa-level BUSCO subset. -
scripts
Contains the custom scripts used during OrthoFinder input processing and USCO retrieval. These scripts support the extraction of primary transcripts and the selection of Tenebrionidae-level USCOs from the OrthoFinder results.
Naming conventions:
Script names describe their main function. The shell script
get.uscos.N7.teneb.shrefers to retrieval of USCOs from theN7Tenebrionidae OrthoFinder node. The Python scriptprimary_transcript.py(a modified version of the original script provided by the OrthoFinder authors in the relevant repository; https://github.com/davidemms/OrthoFinder) was used to retain only primary transcripts prior to the OrthoFinder run, as recommended.File(s) / Directory File type Open with Role Description get.uscos.N7.teneb.shShell script Text editor, Unix shell Custom analysis script Script used to retrieve the Tenebrionidae-level USCO set from the OrthoFinder results and to generate taxon-specific USCO FASTA files and bait sequence files. primary_transcript.pyPython script Text editor, Python Modified input processing script for OrthoFinder Script used to process OrthoFinder input files, in order to retain only primary transcripts for the OrthoFinder analyses.
2_pipeline.zip
The compressed master directory pipeline contains a total of 2 directories and 18 files. It includes the SLURM batch scripts and helper scripts used to run the custom local assembly pipeline. The numbered SLURM batch scripts represent the main assembly workflow, beginning with read processing and database construction, followed by local sequence assembly using SRAssembler, exon identification using Exonerate, exon evaluation and stitching using the custom ExAMiner tool, read self-mapping, consensus calling, alignment, alignment cleaning/filtering, partitioning, gene tree inference, and final dataset generation. The helper_scripts/ directory contains custom R and Python scripts called by the SLURM batch workflow. Note that the scripts start from 2- as step 1- is considered the bait selection step implemented using OrthoFinder, as described above.
Naming conventions:
The numbered prefixes indicate the order of the pipeline steps. Scripts beginning with 2-, 3-, 4-, etc. correspond to successive major steps in the workflow. Scripts with decimal numbering, such as 5.1, 5.2, and 5.3, represent consecutive substeps within the same broader stage. File names including sra correspond to scripts adapted for SRAssembler-related steps. File names including DP4B_iupac_N25S90 indicate the consensus generation and filtering settings used in the final pipeline: DP4 refers to the minimum read depth setting, B denotes the default Bayesian consensus mode as implemented in Samtools, iupac indicates that ambiguous bases were represented using IUPAC ambiguity codes, and N25S90 indicates filtering thresholds, where N is the allowed percentage of missing data per assembled sequence and S is the minimum specimen coverage percentage. Files ending in .slurm are SLURM batch scripts, .R files are R helper scripts, and .py files are Python helper scripts.
| File(s) / Directory | File type | Open with | Role | Description |
|---|---|---|---|---|
2-process_reads.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / read processing and filtering | SLURM batch script used to process raw Illumina reads from WGS. |
3-make.database.sra.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / SRAssembler database construction | SLURM batch script used to build per-specimen SRAssembler database chunks from the processed read data. |
4-mining.sra.local.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / SRAssembler local assembly | SLURM batch script used for local mining of target loci from the SRAssembler databases using the selected in silico bait protein sequences. The script uses temporary directories with common names for all specimens / array jobs, because it is assumed that each array runs in a separate HPC node. If this is not the case the temporary directory names should be modified accordingly. |
5.1-reformer.sra.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / Exonerate input preparation | SLURM batch script used to restructure SRAssembler output. |
5.2-exonerate.sra.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / exon identification | SLURM batch script used to run Exonerate with a protein-to-genome model for each assembled bait/specimen combination. |
5.3-ExAMiner_stitcher.sra.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / exon stitching and mapping-reference preparation | SLURM batch script used to run the custom ExAMiner Python script. This step evaluates and stitches Exonerate exons, generates stitched locus FASTA files, and writes exon/intron coordinate information used during the read self-mapping step and consensus extraction. |
6.1-mapper.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / read self-mapping | SLURM batch script used to map reads back to assembled loci with BWA, filter mapped reads with Samtools, clip read overlaps, index final BAM files, and write mapping statistics. |
6.2-consensus.DP4B_iupac.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / consensus calling | SLURM batch script used to generate per-locus consensus FASTA files from mapped reads using Samtools. |
7-aligner.DP4B_iupac_N25S90.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / sequence filtering and multiple sequence alignment | SLURM batch script used to filter consensus sequences by missing-data content, keep loci represented in enough samples, and align retained loci with MACSE. |
8.1-cleaning.DP4B_iupac_N25S90.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / alignment cleaning and ML gene tree inference | SLURM batch script used to filter alignments for internal stop codons and frameshifts, trim alignments, infer Maximum Likelihood gene trees with IQ-TREE, and generate alignment files for downstream filtering. |
8.2-filtering.DP4B_iupac_N25S90.corr.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / final alignment filtering | SLURM batch script used to compute pre-filtering and post-filtering AMAS statistics, run the custom filtering script physsu.R, and prepare the post-filtered dataset. |
9.1.1-partfinder.DP4B_iupac_N25S90.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / PartitionFinder2 analyses | SLURM batch script used to convert filtered alignments to PHYLIP format, prepare PartitionFinder2 configuration files from a template, and run PartitionFinder2 for each retained locus. |
9.1.2-mrbayes.DP4B_iupac_N25S90.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / Bayesian gene tree inference | SLURM batch script used to prepare MrBayes NEXUS files using the PartitionFinder2 best scheme, and run MrBayes. |
9.2-datasets.DP4B_iupac_N25S90.slurm |
SLURM batch script | Text editor, Unix shell, SLURM scheduler | Pipeline script / dataset preparation | SLURM batch script used to run the custom script datagene.R to select the final dataset, root retained trees, and extract the Metazoa-level BUSCO subset. |
helper_scripts/datagene.R |
R script | Text editor, R | Helper script / final dataset selection | Custom R helper script used to calculate locus statistics from AMAS tables and ML/BI gene trees, filter loci based on a series of thresholds, write statistics, and produce a whitelist of retained loci. Note that the helper script uses code from the https://github.com/marekborowiec/good_genes GitHub repository, which is linked to Borowiec et al. (2015, BMC Genomics 16:987) and other related publications. Please, make sure to check the linked repository and cite the relevant publications when applicable. |
helper_scripts/physsu.R |
R script | Text editor, R | Helper script / alignment and gene tree filtering | Custom R helper script used to summarize gene tree statistics, calculate clock-likeness and related filtering statistics, and write whitelists for downstream dataset preparation. Note that the helper script uses code from the https://github.com/marekborowiec/good_genes GitHub repository, which is linked to Borowiec et al. (2015, BMC Genomics 16:987) and other related publications. Please, make sure to check the linked repository and cite the relevant publications when applicable. |
helper_scripts/ExAMiner/ExAMiner.py |
Python script | Text editor, Python | Main exon stitching script | Custom Python tool used to parse Exonerate output, infer exon/intron coordinates, evaluate and stitch exon sequences. |
helper_scripts/ExAMiner/utils.py |
Python script | Text editor, Python | Helper module | Necessary auxiliary code to run ExAMiner.py. |
3_final_dataset.zip
The compressed master directory final_dataset contains a total of 6 directories and 16,863 files. It includes the final curated phylogenomic dataset used for downstream analyses: multiple sequence alignments, gene trees, alignment summary statistics, and the corresponding subset of loci identified as Metazoa-level BUSCOs (mzlbusco). The main alignments/ and trees/ directories contain the complete post-filtering dataset, whereas mzlbusco/ contains the reduced BUSCO-based subset used for comparison analyses.
Naming conventions:
Alignment and tree files follow the pattern [locus].[dataset_or_method].[status].[extension]. The [locus] prefix is divided into the OG[orthogroup] component that indicates the OrthoFinder orthogroup, and the [reference_accession] component that identifies the initial reference bait. Alignment files ending in .all.fasta include all sampled specimens, including outgroups, whereas files ending in .ingroup.fasta include only ingroup specimens. Tree files ending in .BI.treefile are Bayesian Inference gene trees, and files ending in .ML.treefile are Maximum Likelihood gene trees. Files containing .rooted. are rooted versions of the corresponding gene trees. The mzlbusco/ directory follows the same naming conventions as the complete dataset but contains only the Metazoa-level BUSCO subset.
| File(s) / Directory | File type | Open with | Role | Description |
|---|---|---|---|---|
alignments/[locus].all.fasta |
FASTA alignment files | Text editor, alignment viewer, Unix shell, R, Python | Final alignment dataset | Curated MSAs for the complete post-filtering dataset, including all sampled specimens and outgroups. |
alignments/[locus].ingroup.fasta |
FASTA alignment files | Text editor, alignment viewer, Unix shell, R, Python | Final alignment dataset | Curated MSAs for the complete post-filtering dataset, including only ingroup specimens. |
trees/[locus].BI.treefile |
Tree files in text/Newick format | Text editor, tree viewers such as FigTree, Unix shell, R, Python | Final gene tree dataset | Unrooted Bayesian Inference gene trees inferred for each locus in the complete post-filtering dataset. |
trees/[locus].BI.rooted.treefile |
Tree files in text/Newick format | Text editor, tree viewers such as FigTree, Unix shell, R, Python | Final gene tree dataset | Rooted Bayesian Inference gene trees inferred for each locus in the complete post-filtering dataset. |
trees/[locus].ML.treefile |
Tree files in text/Newick format | Text editor, tree viewers such as FigTree, Unix shell, R, Python | Final gene tree dataset | Unrooted Maximum Likelihood gene trees inferred for each locus in the complete post-filtering dataset. |
trees/[locus].ML.rooted.treefile |
Tree files in text/Newick format | Text editor, tree viewers such as FigTree, Unix shell, R, Python | Final gene tree dataset | Rooted Maximum Likelihood gene trees inferred for each locus in the complete post-filtering dataset. |
stats/stats_prefilter.tsv |
Tab-delimited table | Text editor, Unix shell, R, Python, spreadsheet software | Summary statistics | Per-locus alignment and gene tree statistics calculated before the final filtering step. |
stats/stats_postfilter.tsv |
Tab-delimited table | Text editor, Unix shell, R, Python, spreadsheet software | Summary statistics | Per-locus alignment and gene tree statistics calculated after the final filtering step. |
stats/stats_postfilter.pdf |
PDF file | PDF viewer | Summary figure | Visual summary of post-filtering statistics for the final curated dataset. |
mzlbusco/ |
Multiple file types | N/A |
BUSCO subset | Directory containing the Metazoa-level BUSCO (considering the metazoa_odb10 lineage dataset) subset of the final dataset, including alignments and gene trees that follow the same naming conventions as the complete dataset. |
-
Column descriptions for statistics included in the
stats/directoryColumn Interpretation Units / values LocusLocus identifier. Categorical with OG[orthogroup]_[reference_accession]pattern (e.g.,OG0008521_RZC42532.1)No_of_taxaNumber of specimens represented in the MSA for the corresponding locus. Integer Alignment_lengthLength of the MSA after filtering. Number of nucleotide sites Missing_percentPercentage of missing data in the MSA. Percentage (range 0 to 100) Proportion_parsimony_informativeProportion of MSA sites that are parsimony informative. Proportion (range 0 to 1) GC_contentGC content of the MSA. Proportion (range 0 to 1) CIDClustering Information Distance between the ML and BI gene trees for the corresponding locus, used as a measure of topological dissimilarity. Distance value (lower values indicate more similar topologies) Average_bootstrap_MLMean bootstrap support across internal nodes of the ML gene tree. Bootstrap support value (range 0 to 100) Average_prob_BIMean posterior probability across internal nodes of the BI gene tree. Posterior probability value (range 0 to 1) Average_branch_length_MLMean branch length of the maximum likelihood gene tree. Substitutions per site Average_branch_length_BIMean branch length of the Bayesian inference gene tree. Substitutions per site Clocklikeness_MLClocklikeness score for the ML gene tree, as defined in https://github.com/marekborowiec/good_genes (Borowiec et al.; 2015, BMC Genomics 16:987). Lower values indicate a more clocklike tree Clocklikeness_BIClocklikeness score for the BI gene tree, as defined in https://github.com/marekborowiec/good_genes (Borowiec et al.; 2015, BMC Genomics 16:987). Lower values indicate a more clocklike tree
4_analyses.zip
The compressed master directory analyses contains a total of 2,793 directories and 18,335 files. It includes the files and scripts used in the main downstream analyses reported in the relevant manuscript. It is organized into three major analysis sections: species delimitation, species tree inference, and molecular evolution analyses. Each section contains input files, helper files, scripts, and software outputs. The directory includes both the complete dataset analyses and, where applicable, analyses of the Metazoa-level BUSCO subset (mzlbusco).
-
1_species_delimitation
Contains the input files, helper files, scripts, and outputs for the species delimitation analyses. It includes two main analyses:
speedemon/, which contains Speedemon input XML files and output summaries across locus batches, replicate runs, and ε settings, andtr2/, which contains TR2 input trees and resulting species delimitation summaries based on Bayesian Inference and Maximum Likelihood gene tree sets. Thespeedemon/1_info/svdq_qage/subdirectory contains auxiliary SVDQuartets/QAGE files used to select the range of explored ε values.Naming conventions:
Speedemon files follow the pattern
speedemon_[batch]_[replicate]_e[epsilon]_inv_g1_strict, where[batch]identifies the 100-locus batch,[replicate]is replicateaorb,e[epsilon]encodes the ε value used for the run, andstrictrefers to the strict clock setting. Speedemon output directories and files keep the same base name as the corresponding XML input. In the table below,[run]refers to one such Speedemon run directory and[batch]refers to the corresponding batch directory.TR2 files beginning with
crossval.BIuse Bayesian Inference gene trees, whereas files beginning withcrossval.MLuse Maximum Likelihood gene trees. Files containingguidedrefer to analyses using an external weighted-ASTRAL guide tree, while files containinginternalrefer to delimitation results using an internally inferred guide tree. Directory entries ending with/group multiple related input, output, or script files.File(s) / Directory File type Open with Role Description speedemon/0_data/*.xmlXML files Text editor, BEAST2-compatible tools BEAST2/Speedemon input Input XML files for Speedemon species delimitation analyses across 100-locus batches, replicate runs, and ε settings. speedemon/1_info/lists/list.combinations.100loci.strict.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Helper file List of batch, replicate, and ε combinations used by the Speedemon run and result-collection SLURM batch scripts. speedemon/1_info/svdq_qage/Multiple file types N/ASupplementary analysis directory for Speedemon Directory containing input files, scripts, logs, and output trees for the supplementary SVDQuartets/QAGE analysis used to help define ε value ranges for the Speedemon runs. speedemon/2_output/[batch]/[run]/*.con.treTree files Text editor, tree viewers such as FigTree, Unix shell, R, Python Speedemon output Consensus tree files from individual Speedemon runs, produced using BEAST2 treeannotator.speedemon/2_output/[batch]/[run]/*.ess.values.txtPlain text summary files Text editor, Unix shell, R, Python Speedemon diagnostics Effective sample size summaries from individual Speedemon log files, produced using BEAST2 loganalyser.speedemon/2_output/[batch]/[run]/*.species.txtPlain text summary files Text editor, Unix shell, R, Python Speedemon output Species delimitation summaries from individual Speedemon tree posterior files, produced using ClusterTreeSetAnalyser.speedemon/2_output/[batch]/[run]/*.treesTree posterior files Text editor, tree viewers such as FigTree, Unix shell, R, Python Speedemon output Posterior tree files generated by individual Speedemon runs. speedemon/scripts/SLURM batch scripts Text editor, Unix shell, SLURM scheduler Analysis and result-collection scripts Scripts used to run the Speedemon analyses across 100-locus batches and to summarize their results. tr2/0_data/crossval.BI.treesTree set file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 input generation Bayesian Inference gene tree set used to generate TR2 input and weighted-ASTRAL guide trees. tr2/0_data/crossval.BI.rooted.treesTree set file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 input Rooted Bayesian Inference gene tree set used as TR2 input. tr2/0_data/crossval.BI.wastral.guide.treeTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 guide tree input generation Unrooted weighted-ASTRAL guide tree based on Bayesian Inference gene trees. tr2/0_data/crossval.BI.wastral.guide.rooted.treeTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 guide tree input Rooted weighted-ASTRAL guide tree based on Bayesian Inference gene trees. tr2/0_data/crossval.ML.b20.treesTree set file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 input generation Maximum likelihood gene tree set used to generate TR2 input and weighted-ASTRAL guide trees. Gene tree branches with bootstrap support <20 have been collapsed. tr2/0_data/crossval.ML.rooted.b20.treesTree set file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 input Rooted Maximum Likelihood gene tree set used as TR2 input. Gene tree branches with bootstrap support <20 have been collapsed. tr2/0_data/crossval.ML.wastral.guide.treeTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 guide tree input generation Weighted-ASTRAL guide tree based on Maximum Likelihood gene trees. tr2/0_data/crossval.ML.wastral.guide.rooted.treeTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 guide tree input Rooted weighted-ASTRAL guide tree based on Maximum Likelihood gene trees. tr2/1_info/list.outgroups.eut_data.txtPlain text file Text editor, Unix shell, R, Python Helper file List of outgroup specimens used for TR2 input preparation and tree rooting. tr2/2_output/crossval.BI.guided.table.txtPlain text table Text editor, Unix shell, R, Python, spreadsheet software TR2 output Tabular TR2 species delimitation result based on Bayesian Inference trees and external guide tree information. tr2/2_output/crossval.BI.guided.treTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 output TR2 species delimitation tree based on Bayesian Inference trees and external guide tree information. tr2/2_output/crossval.BI.internal.table.txtPlain text table Text editor, Unix shell, R, Python, spreadsheet software TR2 output Tabular TR2 species delimitation result based on Bayesian Inference trees and an internally inferred guide tree. tr2/2_output/crossval.BI.internal.treTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 output TR2 species delimitation tree based on Bayesian Inference trees and an internally inferred guide tree. tr2/2_output/crossval.ML.guided.table.txtPlain text table Text editor, Unix shell, R, Python, spreadsheet software TR2 output Tabular TR2 species delimitation result based on Maximum Likelihood trees and external guide tree information. tr2/2_output/crossval.ML.guided.treTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 output TR2 species delimitation tree based on Maximum Likelihood trees and external guide tree information. tr2/2_output/crossval.ML.internal.table.txtPlain text table Text editor, Unix shell, R, Python, spreadsheet software TR2 output Tabular TR2 species delimitation result based on Maximum Likelihood trees and an internally inferred guide tree. tr2/2_output/crossval.ML.internal.treTree file Text editor, tree viewers such as FigTree, Unix shell, R, Python TR2 output TR2 species delimitation tree based on Maximum Likelihood trees and an internally inferred guide tree. tr2/analyses.tr2.slurmSLURM batch script Text editor, Unix shell, SLURM scheduler Analysis script SLURM batch script used to generate weighted-ASTRAL guide tree, reroot guide tree, and run TR2 species delimitation analyses with and without external guide trees. -
2_species_tree
Contains the input files, scripts, and outputs for species tree inference using weighted-ASTRAL, BPP, and SVDQuartets/QAGE. The
astral/subdirectory contains gene tree inputs and weighted-ASTRAL species tree outputs from the complete dataset. Thebpp/subdirectory contains BPP control files, sequence alignments, imap files, constraints, scripts, and majority-rule consensus trees for batches of 100 loci and themzlbuscosubset. Thesvdquartets/subdirectory contains NEXUS input files, scripts, logs, and species tree outputs from the SVDQuartets/QAGE analysis.Naming conventions:
Files beginning with
crossval_c05refer to the complete dataset used for species tree inference.BIindicates Bayesian Inference gene trees, whereasMLindicates Maximum Likelihood gene trees.b20indicates that branches in ML gene trees with bootstrap support <20 have been collapsed. Files beginning withwastralare weighted-ASTRAL species tree outputs. BPP files use the prefixa01, indicating the BPP A01 analysis. In BPP paths,[dataset]refers to either100loci_batches_crossvalormzlbusco, depending on the subdirectory. Batch directories follow the patternbatch_[01-24]to indicate the specific batch number out of 24 batches, and control files includestarttree_1orstarttree_2to indicate the starting tree used for that run. Directory entries ending with/, such asbpp/scripts/andsvdquartets/scripts/, group (one or more) analysis scripts.File(s) / Directory File type Open with Role Description astral/0_data/crossval_c05.BI.treesTree set file Text editor, tree viewers such as FigTree, Unix shell, R, Python, ASTRAL-compatible tools ASTRAL input Bayesian Inference gene tree set used as input for weighted-ASTRAL species tree inference. astral/0_data/crossval_c05.ML.b20.treesTree set file Text editor, tree viewers such as FigTree, Unix shell, R, Python, ASTRAL-compatible tools ASTRAL input Maximum likelihood gene tree set used as input for weighted-ASTRAL species tree inference. Gene tree branches with bootstrap support <20 have been collapsed. astral/0_data/imap.outgroup.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software ASTRAL input / Helper file Taxon-mapping file used for weighted-ASTRAL species tree inference. astral/1_output/wastral.2408loci.BI.sptree.nwNewick tree file Text editor, tree viewers such as FigTree, Unix shell, R, Python ASTRAL output Weighted-ASTRAL species tree inferred from the 2,408-locus Bayesian Inference gene tree set. astral/1_output/wastral.2408loci.ML.b20.sptree.nwNewick tree file Text editor, tree viewers such as FigTree, Unix shell, R, Python ASTRAL output Weighted-ASTRAL species tree inferred from the 2,408-locus Maximum Likelihood gene tree set. astral/analyses.astral.slurmSLURM batch script Text editor, Unix shell, SLURM scheduler Analysis script SLURM batch script used to run weighted-ASTRAL species tree inference from BI and ML gene tree sets. bpp/0_data/100loci_batches_crossval/[batch]/*.ctrBPP control files Text editor, Unix shell, R, Python, BPP BPP input / control files BPP A01 control files for each 100-locus batch, including alternative starting tree conditions. bpp/0_data/100loci_batches_crossval/[batch]/*.phyPHYLIP alignment files Text editor, Unix shell, R, Python, BPP, sequence-alignment tools BPP input Multilocus PHYLIP alignment files used by BPP for the corresponding batch. bpp/0_data/100loci_batches_crossval/[batch]/constraints.txtPlain text file Text editor, Unix shell, R, Python, BPP BPP input / helper file Constraint file used by BPP for the corresponding batch. bpp/0_data/100loci_batches_crossval/[batch]/imap.outgroup.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software, BPP BPP input / helper file Taxon-mapping file used by BPP for the corresponding batch. bpp/0_data/mzlbusco/Multiple file types N/ABPP input directory / BUSCO subset Directory containing BPP A01 control files, constraints, imap file, and multilocus PHYLIP input file for the Metazoa-level BUSCO subset. bpp/1_output/*.treTree files Text editor, tree viewers such as FigTree, Unix shell, R, Python BPP output Majority-rule consensus trees from BPP A01 analyses of the complete dataset batches and the mzlbuscosubset.bpp/scripts/SLURM batch scripts Text editor, Unix shell, SLURM scheduler Analysis scripts Scripts used to run BPP A01 analyses for the 100-locus complete dataset batches and the mzlbuscosubset.svdquartets/0_data/concat_alinments_mod.nexusNEXUS alignment file Text editor, Unix shell, R, Python, PAUP*, sequence-alignment viewer SVDQuartets input Concatenated alignment file used for SVDQuartets/QAGE species tree inference. svdquartets/0_data/svdq_qage_run.nexusNEXUS command/input file Text editor, Unix shell, R, Python, PAUP* SVDQuartets input / analysis script NEXUS run file for the SVDQuartets/QAGE species tree analysis. svdquartets/1_output/qage_sp_tree_crossval.cons.nexNEXUS tree file Text editor, tree viewers such as FigTree, PAUP*, Unix shell, R, Python QAGE output Consensus species tree from the QAGE analysis. svdquartets/1_output/svdq_sp_tree_crossval.cons.nexNEXUS tree file Text editor, tree viewers such as FigTree, PAUP*, Unix shell, R, Python SVDQuartets output Consensus species tree from the SVDQuartets analysis. svdquartets/1_output/svdq_sp_tree_crossval.nexNEXUS tree file Text editor, tree viewers such as FigTree, PAUP*, Unix shell, R, Python SVDQuartets output SVDQuartets species tree output file. svdquartets/1_output/svdq_qage.log.txtPlain text log file Text editor, Unix shell, R, Python Analysis log Log file from the SVDQuartets/QAGE species tree analysis. svdquartets/scripts/SLURM batch scripts Text editor, Unix shell, SLURM scheduler Analysis scripts Script directory used to run the SVDQuartets/QAGE species tree analysis. -
3_molecular_evolution
Contains input / output files, and scripts for the molecular evolution analyses. The
0_mol_evol_data/directory contains input files and outputs from BPP, HyPhy, MCMCTree, and eggNOG-mapper functional annotation. The1_mol_evol_analyses/directory contains summarized per-locus or per-batch result tables used by the master R script for downstream analyses and plotting. Each batch of loci for the molecular evolution analyses includes 195 loci.Naming conventions:
Files and directories containing
all_sitescorrespond to analyses using all alignment sites, whereas files and directories containing4fold_sitescorrespond to analyses restricted to four-fold degenerate sites. In the table below,[dataset]therefore refers to eitherall_sitesor4fold_sites, unless otherwise specified. Files containingmzlbuscocorrespond to the Metazoa-level BUSCO subset. Batch directories follow the patternbatch_[01-11]to indicate specific batch number out of 11 batches, and MCMCTree replicate directories follow the patternrep[1-4]to indicate the specific replicate number out of 4 replicates. Locus-specific directories and files use the patternOG[orthogroup]_[reference_accession], abbreviated as[locus]in the table. Files beginning withcrossval.BI.wastralrefer to weighted-ASTRAL guide trees based on Bayesian Inference gene trees. Files ending in.renamed.tsvcontain versions of MCMCTree output tables with loci renamed to their originalOG[orthogroup]_[reference_accession]code, and taxa or internal nodes renamed according to the relevant study code. Directory entries ending with/group multiple related input, output, or script files.-
0_mol_evol_dataFile(s) / Directory File type Open with Role Description bpp/[dataset]/bpp_runs/[batch]/*.ctrBPP control files Text editor, BPP, Unix shell, R, Python BPP input / control files BPP A00 control files for estimating population-genetic parameters from the specified dataset ( all_sitesor4fold_sites) in each 195-locus batch.bpp/[dataset]/bpp_runs/[batch]/imap.bpp.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software, BPP BPP input / helper file Taxon-mapping file used by BPP for the specified dataset in each 195-locus batch. bpp/[dataset]/bpp_runs/[batch]/*.phyPHYLIP alignment files Text editor, Unix shell, R, Python, BPP, sequence-alignment tools BPP input Multilocus PHYLIP alignment file for the corresponding BPP A00 batch and dataset. bpp/[dataset]/bpp_runs/groups.txtPlain text file Text editor, Unix shell, R, Python Helper file Grouping file to define BPP batches bpp/scripts/SLURM batch scripts Text editor, Unix shell, SLURM scheduler Analysis scripts Scripts used to run BPP A00 analyses for both all_sitesand4fold_sites195-locus batch datasets.functional_annotation/0_data/multi-FASTA file Text editor, Unix shell, R, Python, sequence viewer, eggNOG-mapper Functional annotation input directory Directory containing the query FASTA file used for eggNOG-mapper functional annotation. functional_annotation/1_results/Multiple file types N/AFunctional annotation output directory Directory containing eggNOG-mapper annotation and associated files. hyphy/0_data/[locus]/[locus].all.fastaFASTA alignment files Text editor, alignment viewer, HyPhy, Unix shell, R, Python HyPhy input Locus-specific codon alignment files used for HyPhy analyses. hyphy/0_data/[locus]/*.treeTree files Text editor, tree viewers such as FigTree, HyPhy, Unix shell, R, Python HyPhy input Marked rooted guide tree files, including pruned versions matching the taxa present in the corresponding locus alignment. hyphy/0_data/[locus]/list.tips.txtPlain text file Text editor, Unix shell, R, Python HyPhy helper file List of tips used for the corresponding locus-specific HyPhy analysis. hyphy/1_info/Multiple file types N/AHyPhy helper directory Directory containing the marked rooted reference species tree and list of loci used to prepare locus-specific HyPhy inputs. hyphy/2_output/*.jsonJSON files Text editor, JSON viewer, Unix shell, R, Python HyPhy output Locus-specific HyPhy JSON output files, used to extract branch-specific dN/dS ratios. hyphy/2_output/*.logLog files Text editor, Unix shell, R, Python HyPhy output Locus-specific HyPhy log files. hyphy/scripts/SLURM batch scripts Text editor, Unix shell, SLURM scheduler Analysis scripts Script directory used to run the locus-specific HyPhy analyses. mcmctree/[dataset]/mcmctree_runs/[batch]/step1/Multiple file types N/AMCMCTree step 1 input Directory containing step 1 MCMCTree input and control files for the specified dataset and 195-locus batch, including calibrated tree, multilocus alignment, generated control file. mcmctree/[dataset]/mcmctree_runs/[batch]/step2/Multiple file types N/AMCMCTree step 2 input Directory containing input files for multiple replicate runs of MCMCTree step 2, for the specified dataset and 195-locus batch, including calibrated tree, multilocus alignment, generated control files, and the in.BVfile created in step 1.mcmctree/[dataset]/mcmctree_runs/mzlbusco/Multiple file types N/AMCMCTree BUSCO-subset input directory Directory containing MCMCTree step 1 and step 2 files for the Metazoa-level BUSCO subset, following the same structure as the batch directories. mcmctree/[dataset]/mcmctree_runs/groups.txtPlain text file Text editor, Unix shell, R, Python Helper file Grouping file used to map SLURM array task IDs to MCMCTree batch or mzlbuscodirectories.mcmctree/[dataset]/mcmctree_output/*.tsvTab-delimited tables Text editor, Unix shell, R, Python, spreadsheet software MCMCTree summarized output MCMCTree output tables for the specified dataset, including batch and replicate summaries and renamed versions used for downstream analyses. mcmctree/mcmctree_locus_map.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Helper file Mapping table linking MCMCTree locus identifiers to the corresponding loci. mcmctree/mcmctree_node_map.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Helper file Mapping table linking MCMCTree node identifiers to the corresponding renamed tips or internal tree nodes. mcmctree/scripts/SLURM batch scripts Text editor, Unix shell, SLURM scheduler Analysis scripts Scripts used to run MCMCTree step 1 and step 2 analyses for both all_sitesand4fold_sitesdatasets, including batch analyses and themzlbuscosubset where applicable. -
1_mol_evol_analysesFile(s) / Directory File type Open with Role Description 0_data/dnds/*.tsvTab-delimited tables Text editor, Unix shell, R, Python, spreadsheet software Summarized analysis input Per-locus dN/dS summary tables extracted from HyPhy outputs. 0_data/theta/[dataset]/*.tsvTab-delimited tables Text editor, Unix shell, R, Python, spreadsheet software Summarized analysis input Per-batch θsummary tables extracted from BPP A00 analyses of theall_sitesand4fold_sitesdatasets.0_data/rates/mcmctree_master_rates_terminals_only.tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Summarized analysis input Master table of substitution rate estimates summarized from MCMCTree output. 0_data/functional_annotation/[...].tsvTab-delimited table Text editor, Unix shell, R, Python, spreadsheet software Summarized analysis input Functional annotation summary table. mol_evol_analyses.master.rR script Text editor, R, RStudio Master analysis script R script used run the molecular evolution analyses and comparisons. - Column descriptions for summarized tables included in the
1_mol_evol_analyses/0_data/directory-
0_data/dnds/*.tsvColumn Interpretation Units / values BranchTerminal branch or internal node identifier. Categorical (e.g., NODE_01)dNEstimated nonsynonymous substitution rate for the corresponding branch. Nonsynonymous substitutions per nonsynonymous site dSEstimated synonymous substitution rate for the corresponding branch. Synonymous substitutions per synonymous site dNdSRatio of nonsynonymous to synonymous substitution rates, calculated as dN / dS.Ratio lbLower confidence bound for the dNdSestimate.Ratio ubUpper confidence bound for the dNdSestimate.Ratio pRaw p-value associated with the branch-specific dNdSestimate.Probability (range 0 to 1) corrected_pMultiple-testing corrected p-value. Probability (range 0 to 1) combined_branch_lengthTotal branch length estimate Substitutions per site synonymous_branch_lengthSynonymous branch length estimate. Synonymous substitutions per site nonsynonymous_branch_lengthNonsynonymous branch length estimate. Nonsynonymous substitutions per site -
0_data/rates/mcmctree_master_rates_terminals_only.tsvColumn Interpretation Units / values speciesTerminal branch identifier. Categorical (e.g., GEN_CY,GEO_WAEC,PSA_EM)locusLocus identifier. Categorical with OG[orthogroup]_[reference_accession]patternrateTotal substitution rate estimate from MCMCTree in the original analysis time scale. Numeric rate estimate rate_per_MYTotal substitution rate converted to a per-million-year scale. Substitutions per site per million years rate_per_genTotal substitution rate converted to a per-generation scale. Substitutions per site per generation (one generation is assumed to correspond to a year) 4fold_rateFour-fold degenerate site substitution rate estimate from MCMCTree in the original analysis time scale. Numeric rate estimate 4fold_rate_per_MYFour-fold degenerate site substitution rate converted to a per-million-year scale. Substitutions per four-fold degenerate site per million years. 4fold_rate_per_genFour-fold degenerate-site substitution rate converted to a per-generation scale. Substitutions per four-fold degenerate site per generation (one generation is assumed to correspond to a year). batch_age_MYTerminal branch or internal node age estimate from the corresponding MCMCTree batch analysis. Million years batch_age_HPDlo_MYLower bound of the 95% HPD interval for batch_age_MY.Million years batch_age_HPDup_MYUpper bound of the 95% HPD interval for batch_age_MY.Million years mzlbusco_age_MYCorresponding age estimate from the Metazoa-level BUSCO MCMCTree analysis. Million years mzlbusco_age_HPDlo_MYLower bound of the 95% HPD interval for mzlbusco_age_MY.Million years mzlbusco_age_HPDup_MYUpper bound of the 95% HPD interval for mzlbusco_age_MY.Million years batchMCMCTree batch from which the estimate was obtained. Categorical (e.g., batch_01)repMCMCTree batch replicate from which the estimate was obtained. Categorical (e.g., rep1,rep2,rep3,rep4)batch_typeDataset type used for the MCMCTree run. Categorical (e.g., totalif it corresponds to a batch run from the complete dataset ormzlbuscoif it corresponds to the Metazoa-level BUSCO subset) -
0_data/theta/[dataset]/*.tsvColumn / pattern Interpretation Units / values sampleRow index identifier. Integer filenamePath to the original BPP MCMC output file from which the summary was extracted. File path theta:[number]:[lineage].meanPosterior mean of the BPP population size parameter thetafor the specified terminal or internal branch.Numeric BPP thetaestimatetheta:[number]:[lineage].stderrStandard error of the posterior mean for the corresponding thetaestimate.Numeric theta:[number]:[lineage].stddevPosterior standard deviation of the corresponding thetaestimateNumeric theta:[number]:[lineage].medianPosterior median of the corresponding thetaestimate.Numeric theta:[number]:[lineage].95%HPDloLower bound of the 95% HPD interval for the corresponding thetaestimate.Numeric theta:[number]:[lineage].95%HPDupUpper bound of the 95% HPD interval for the corresponding thetaestimate.Numeric theta:[number]:[lineage].ACTAutocorrelation time for the MCMC chain of the corresponding thetaparameter.Numeric theta:[number]:[lineage].ESSEffective sample size for the MCMC chain of the corresponding thetaparameter.Numeric theta:[number]:[lineage].geometric-meanGeometric mean of the posterior samples for the corresponding thetaestimate.Numeric tau:[number]:[node_label].meanPosterior mean of the BPP divergence time parameter taufor the specified internal node.Numeric BPP tauestimatetau:[number]:[node_label].stderrStandard error of the posterior mean for the corresponding tauestimate.Numeric tau:[number]:[node_label].stddevPosterior standard deviation of the corresponding tauestimate.Numeric tau:[number]:[node_label].medianPosterior median of the corresponding tauestimate.Numeric tau:[number]:[node_label].95%HPDloLower bound of the 95% HPD interval for the corresponding tauestimate.Numeric tau:[number]:[node_label].95%HPDupUpper bound of the 95% HPD interval for the corresponding tauestimate.Numeric tau:[number]:[node_label].ACTAutocorrelation time for the MCMC chain of the corresponding tauparameter.Numeric tau:[number]:[node_label].ESSEffective sample size for the MCMC chain of the corresponding tauparameter.Numeric tau:[number]:[node_label].geometric-meanGeometric mean of the posterior samples for the corresponding tauestimate.Numeric nu_bar.meanPosterior mean of the BPP variance of branch rates parameter. Numeric nu_bar.stderrStandard error of the posterior mean for nu_bar.Numeric nu_bar.stddevPosterior standard deviation of nu_bar.Numeric nu_bar.medianPosterior median of nu_bar.Numeric nu_bar.95%HPDloLower bound of the 95% HPD interval for nu_bar.Numeric nu_bar.95%HPDupUpper bound of the 95% HPD interval for nu_bar.Numeric nu_bar.ACTAutocorrelation time for the MCMC chain of nu_bar.Numeric nu_bar.ESSEffective sample size for the MCMC chain of nu_bar.Numeric nu_bar.geometric-meanGeometric mean of the posterior samples for nu_bar.Numeric lnL.meanPosterior mean of the log-likelihood. Log-likelihood value lnL.stderrStandard error of the posterior mean for lnL.Numeric lnL.stddevPosterior standard deviation of lnL.Numeric lnL.medianPosterior median of lnL.Log-likelihood value lnL.95%HPDloLower bound of the 95% HPD interval for lnL.Log-likelihood value lnL.95%HPDupUpper bound of the 95% HPD interval for lnL.Log-likelihood value lnL.ACTAutocorrelation time for the MCMC chain of lnL.Numeric lnL.ESSEffective sample size for the MCMC chain of lnL.Numeric lnL.geometric-meanGeometric mean of the posterior samples for lnL.Numeric -
0_data/functional_annotation/functional_annotation_summary.clean.tsvColumn Interpretation Units / values locusLocus identifier. Categorical with OG[orthogroup]_[reference_accession]patternCOG_categoryFunctional category assigned from eggNOG-mapper (COG terms can be found in https://www.ncbi.nlm.nih.gov/research/cog#). Single-letter COG category code DescriptionFunctional description associated with the best annotation or inferred functional category. Plain text Preferred_namePreferred gene or protein name assigned by eggNOG-mapper, where available. Plain text GOsGene Ontology terms associated with the annotation. Comma-separated GO identifiers
-
- Column descriptions for summarized tables included in the
-
MAIN SOFTWARE
| Software / tool | Version | Main usage |
|---|---|---|
fastp |
v.0.23.4 |
Read processing, trimming, adapter detection, duplicate filtering, and quality filtering |
OrthoFinder |
v.2.5.4 |
Orthogroup inference and Tenebrionidae-level USCO identification |
SRAssembler |
v.1.2.0 |
Local assembly of genomic regions targeted by the in silico baits |
Exonerate |
v.2.4.0 |
Identification of putative exons from locally assembled contigs |
ExAMiner.py |
custom script | Evaluation of Exonerate output and production of stitched exonic sequences |
BWA |
v.0.7.17 |
Mapping reads back to assembled loci |
SAMtools |
v.1.17 |
Consensus calling and low-coverage masking |
MACSE |
v.2.07 |
Codon-aware alignment of consensus coding sequences |
IQ-TREE |
v.2.3.6 |
Maximum likelihood gene-tree inference |
MrBayes |
v.3.2.6 |
Bayesian inference gene-tree estimation |
PartitionFinder2 |
v.2.1.1 |
Partitioning scheme and substitution model selection for MrBayes |
AMAS |
not specified | Alignment summary statistics and alignment manipulation |
BUSCO |
v.5.5.0 |
Identification of the Metazoa-level BUSCO subset |
eggNOG-mapper |
v.2.1.13 |
Functional annotation of the 2,408 single-copy orthologs |
TR2 |
not specified | Species delimitation from rooted gene trees |
weighted-ASTRAL |
v.1.16.3.4 |
Species tree inference from gene trees |
GOTREE |
v.0.4.5 |
Tree rooting and manipulation |
Speedemon |
v.1.1.0 |
Species delimitation implemented in BEAST2 |
BEAST2 |
v.2.7.6 |
Speedemon species delimitation analyses and associated output summarization, including use of BEAST2 tools such as loganalyser, treeannotator, and ClusterTreeSetAnalyser. |
BEAUti |
v.2.7.7 |
Generation of BEAST2/Speedemon XML input files |
SVDQuartets in PAUP* |
v.4.0a168 |
Quartet-based species tree inference |
QAGE in PAUP* |
v.4.0a168 |
Node-height estimation for SVDQuartets-derived trees |
BPP |
v.4.8.2 |
A01 species tree inference and A00 θ parameter estimation analyses |
MCMCTree in PAML |
v.4.10.10 |
Substitution rate and node age estimation |
CDSKIT |
v.0.14.5 |
Extraction of four-fold degenerate sites |
HyPhy |
v.2.5.64 |
Branch-specific dN/dS estimation using FitMG94 |
