Supplementary files from: The evolutionary history of praying mantises (Dictyoptera, Mantodea)
Data files
Jun 29, 2026 version files 256.80 KB
-
README.md
7.14 KB
-
Supplement.pdf
53.88 KB
-
Supplementary_file_1_-_dating_analyses_1.tre
9.06 KB
-
Supplementary_file_2_-_dating_analyses_2.tre
4.67 KB
-
Supplementary_file_3_-_dating_analyses_3.tre
4.67 KB
-
Supplementary_file_4_-_dating_analyses_4.tre
4.67 KB
-
Supplementary_file_5_-_FcLM_result_summary.pdf
150.80 KB
-
Supplementary_file_6_-_ML_tree
5.13 KB
-
Supplementary_file_7_-_wAstral_tree
5.81 KB
-
Supplementary_table_1_-_Taxon_sampling.csv
10.97 KB
Abstract
This dataset supports a phylogenomic study reconstructing the evolutionary history of praying mantises (Dictyoptera, Mantodea). It comprises transcriptome-derived sequence data and phylogenetic results for 51 mantodean species sampled across 14 of the 16 currently recognised superfamilies (22 of 29 families), together with a representative set of polyneopteran outgroup taxa. The transcriptomes were generated as part of the 1KITE initiative from specimens acquired between 2012 and 2013, sequenced on an Illumina HiSeq 2000 platform (2×150 bp paired-end), and assembled with SOAPdenovo-Trans. Single-copy orthologs were identified using Orthograph, aligned at the amino acid level, and filtered for alignment quality, randomness, and information content. The core data product is a concatenated amino acid supermatrix of 1,160 single-copy ortholog groups totalling 584,496 amino acid positions, along with the corresponding partitioning scheme and substitution models. The dataset also includes the 616 locus alignments and individual maximum likelihood gene trees used for coalescent-based species-tree estimation, as well as the resulting phylogenetic trees: the maximum likelihood tree (IQ-TREE), the coalescent species tree (wASTRAL), and four time-calibrated trees produced with mcmctree (PAML) under six fossil calibration points. Files include node ages, confidence intervals, and branch support values (non-parametric bootstrap, SH-aLRT, local posterior probabilities, gene- and site-concordance factors), four-cluster likelihood mapping (FcLM) results, and the phylogenetic justification for each calibration fossil. Specimen accession numbers and taxonomic sampling details accompany the alignments. These data enable reanalysis of mantodean relationships, exploration of alternative tree-inference and dating approaches, integration with morphological or biogeographic datasets, and reuse as outgroup or backbone data in broader polyneopteran phylogenomic studies.
Dataset DOI: 10.5061/dryad.hqbzkh1zd
Description of the data and file structure
Dataset Overview
Raw sequence reads and assembled transcriptomes are not stored in Dryad; they are archived at NCBI (see Sharing/Access information below). Accession numbers linking each specimen to its NCBI records are provided in Supplementary Table 1.
Files and variables
Supplementary_table_1_-_Taxon_sampling.csv
Specimen and sampling metadata for all sequenced mantodean taxa, one row per specimen. Columns:
- Species: Genus and species (or morphospecies designation) of the sampled specimen.
- Author: Taxonomic authority and year for the species name ("-" where not applicable).
- NCBI Taxonomy ID: NCBI Taxonomy database identifier for the species.
- Family: Family-level classification (sensu Schwarz and Roy, 2019).
- Superfamily: Superfamily-level classification (sensu Schwarz and Roy, 2019).
- Lifestage: Developmental stage and sex of the specimen where known ("?" where unknown).
- LibID: Internal 1KITE library identifier for the sequenced sample.
- #contigs: Number of contigs in the assembled transcriptome.
- #orthologs: Number of single-copy orthologs recovered for the specimen.
- BioProject Accession or Source (downloaded): NCBI BioProject accession. An asterisk (*) marks records originating from earlier projects/sources.
- BioSample Accession: NCBI BioSample accession.
- Run Accession: NCBI Sequence Read Archive (SRA) run accession.
- TSA Accession: NCBI Transcriptome Shotgun Assembly (TSA) accession.
- TSA Version: Versioned TSA accession.
- Origin: Geographic origin or source of the specimen (e.g., lab culture, field collection locality).
Missing values appear as "-" or "?" as noted above.
Supplement.pdf
Supplementary document providing the phylogenetic justification for the six fossil calibration points used in the divergence-time analyses, plus an index of the separately provided supplementary files. For each fossil, the document records the calibrated node, the fossil taxon and its original description, additional descriptive accounts, the type locality and age (in millions of years ago, MYA), repository and specimen voucher numbers, and a structured justification (fields CR1-CR5) following the best-practice fossil-calibration reporting scheme. The six calibration fossils are: Archeorhinotermes rossi, Valditermes brenanae, "Gyna" obesa, Qilianiblatta namurensis, Juramantophasma sinica, and Alexarasnia rossica. Node labels (A-F) correspond to the calibrated nodes marked in Figure 3 of the article.
Supplementary file 1-4 (time-calibrated trees; .tre/.tree)
Four time-calibrated phylogenetic trees, each representing one of four independent dating analyses run in mcmctree (PAML v4.9) under an independent-rates clock. All four runs produced very similar node ages and confidence intervals; analysis 1 was selected for interpretation in the article (Figure 3). Each tree file includes estimated node ages and associated confidence intervals as well as branch support values.
- Supplementary_file_1_-_dating_analyses_1.tre: dating analysis 1 of 4 (the tree used in the main text).
- Supplementary_file_2_-_dating_analyses_2.tre: dating analysis 2 of 4.
- Supplementary_file_3_-_dating_analyses_3.tre: dating analysis 3 of 4.
- Supplementary_file_4_-_dating_analyses_4.tre: dating analysis 4 of 4.
Node ages are expressed in millions of years (Ma). These files are standard phylogenetic tree files (Newick/Nexus) and can be opened in tree-viewing software such as FigTree or read programmatically (e.g., with ape in R, or DendroPy/Bio.Phylo in Python).
Supplementary_file_5_-_FcLM_result_summary.pdf
Summary of the four-cluster likelihood mapping (FcLM) analyses used to assess support for the backbone and other low-support internodes (13 tested configurations). The workbook contains three sheets:
- Cluster definitions: For each of the 13 analyses, the four taxon clusters (Clust 1-4) and the three alternative resolutions tested (labelled Top, Left, and Right), with a plain-language statement of the sister-group hypothesis each resolution represents.
- Results: For each analysis, the percentage of quartets supporting the Top/Left/Right resolutions under three versions of the supermatrix - (i) unmodified, (ii) amino acids randomised at their original frequency with missing-data patterns retained (rand_equal), and (iii) amino acids randomised at equal frequency with missing-data patterns retained (rand_freq) - together with derived "Apparent signal," "Noise," and "Remaining signal" values and a conclusion as to whether the result is biased by missing data or compositional heterogeneity.
- Support in concatenation tree: For each tested resolution, the corresponding bootstrap support (BS) value recovered in the concatenation analysis ("NA" where the resolution was not recovered).
Percentages are on a 0-100 scale; derived signal/noise values are proportions on a 0-1 scale. Empty cells and "-" denote configurations not analysed (e.g., where a single resolution received full support and randomised tests were not required).
Supplementary_file_6_-_ML_tree (maximum likelihood tree; .tre/.tree)
The best-scoring maximum likelihood tree inferred with IQ-TREE from the concatenated amino acid supermatrix, with branch support values (non-parametric bootstrap; SH-aLRT and concordance factors as described in the article).
Supplementary_file_7_-_wAstral_tree (coalescent tree; .tre/.tree)
The coalescent species tree was estimated with wASTRAL from 616 locus-specific maximum likelihood gene trees, with local posterior probabilities as support values.
Code/software
No custom code is included in this deposit. The trees and alignments were produced with the following published software: SOAPdenovo-Trans (assembly), Orthograph v0.4.5 (orthology), MAFFT v7.221 (alignment), Aliscore v1.2 and AliCUT v2.3 (alignment masking), FASconCAT (concatenation), MARE v0.1.2-rc (information content), PartitionFinder 2 (partitioning), ModelFinder/IQ-TREE v1.5.1 and v3.0.1 (model selection and maximum likelihood inference), RAxML-ng with pargenes (gene trees), wASTRAL (coalescent species tree), AnomalyFinder (anomaly-zone detection), modeltest-ng (per-locus models), and mcmctree/codeml from PAML v4.9 (divergence-time estimation). Tree files use standard Newick/Nexus formats and can be viewed with FigTree or parsed with R or Python.
Sharing/Access information
Raw sequencing reads, assembled transcriptomes, and per-specimen accessions are archived at the NCBI Sequence Read Archive (SRA) and Transcriptome Shotgun Assembly (TSA) databases. The relevant BioProject, BioSample, SRA run, and TSA accession numbers are listed for each specimen in Supplementary Table 1. Specimens generated within the 1KITE project are associated with the BioProject accessions given in that table.
