Data from: Mitogenomic phylogeny of Eurytoma Illiger (Hymenoptera: Eurytomidae): genus delimitation, species-group assessment, and recurrent host-use transitions
Data files
Apr 10, 2026 version files 2.79 GB
-
ASR_value.xlsx
141.55 MB
-
Dataset_1_IQ_partition.txt
782 B
-
Dataset_1_IQ.phy
1.78 MB
-
Dataset_1_IQ.treefile
9.26 KB
-
Dataset_2_BI.nex
27.64 MB
-
Dataset_2_BI.nex.mcmc
43.09 MB
-
Dataset_2_BI.nex.run1.p
1.29 GB
-
Dataset_2_BI.nex.run2.p
1.29 GB
-
Dataset_2_BI.tre
64.12 KB
-
Dataset_3_IQ_partition.txt
777 B
-
Dataset_3_IQ.phy
2.11 MB
-
Dataset_3_IQ.treefile
10.69 KB
-
README.md
4.26 KB
Abstract
This Dryad dataset contains mitochondrial genome–based data and derived analysis files for a phylogenetic framework of the Palaearctic genus Eurytoma (Hymenoptera: Eurytomidae). The dataset includes 166 terminals (159 ingroup specimens and seven outgroups) spanning 13 of the 18 historically recognized Eurytoma species-groups. Contents comprise: (i) assembled mitochondrial genomes and annotations (sequence files and feature tables for protein-coding genes, rRNAs, and tRNAs); (ii) locus-wise and concatenated nucleotide alignments for 13 mitochondrial protein-coding genes plus rrnL/rrnS, including trimming/curation outputs and partition information; (iii) phylogenetic tree files from maximum-likelihood and Bayesian inference (e.g., Newick format with node-support values, where applicable); (iv) specimen-level metadata (taxon names, species-group assignments, voucher/locality fields, and—when available—links/accessions for underlying public reads); and (v) trophic-strategy character matrices used in ancestral-state reconstruction, provided under both binary (entomophagy vs. phytophagy) and multistate codings, with unknown states consistently treated as missing data. Sequence data are standard nucleotide characters (A/C/G/T with IUPAC ambiguity codes where needed); alignment gaps are coded as “-”. These files can be reused for taxonomic placement of newly sampled taxa, comparative mitogenomics, testing species-group diagnoses, and broader trait-evolution analyses in Chalcidoidea.
Dataset DOI: 10.5061/dryad.vx0k6dk6g
Description of the data and file structure
-
“Dataset 1” comprised 13 mitochondrial protein-coding genes (PCGs) from 140 specimens (IQ-TREE analysis). (Dataset_1_IQ.phy)
-
“Dataset 2” comprised 13 mitochondrial PCGs from 140 specimens (MrBayes analysis).
-
“Dataset 3” comprised 13 mitochondrial PCGs from 140 specimens plus COI sequences from 22 additional specimens (IQ-TREE analysis). (Dataset_3_IQ.phy)
-
“Dataset_1_IQ_partition.txt” and “Dataset_3_IQ_partition.txt” are partition files used for IQ-TREE analyses.
-
“Dataset_1_IQ.treefile” and “Dataset_3_IQ.treefile” are the resulting maximum likelihood trees from IQ-TREE.
-
“Dataset_2_BI.nex” and associated files (Dataset_2_BI.nex.mcmc, Dataset_2_BI.nex.run1.p, Dataset_2_BI.nex.run2.p, Dataset_2_BI.tre) are the Bayesian inference process files from MrBayes.
-
"ASR_value.xlsx" comprised posterior probabilities from ancestral state reconstruction (ASR) analyses of trophic strategy/host association within Eurytomidae, summarized for each node under binary and multistate models.
-
ASR_value.xlsx contains two worksheets:
1) Sheet: “Table 1” (binary ASR; state symbols: a, b)
2) Sheet: “Table 2” (multistate ASR; state symbols: a–j)
State definitions
- Binary coding (Table 1):
- a = phytophagy / plant association (phytophagous)
- b = parasitism on animal hosts (entomophagous)
- Multistate coding (Table 2):
- a = stem
- b = seed
- c = flower bud
- d = Ficus
- e = gall (gall-associated)
- f = Hymenoptera (parasitism on hymenopteran hosts)
- g = Lepidoptera (parasitism on lepidopteran hosts)
- h = Coleoptera (parasitism on coleopteran hosts)
- i = hyperparasitoid (secondary parasitism)
- j = egg (egg parasitism)
Column definitions (applies to both sheets unless noted)
Core MCMC (Markov Chain Monte Carlo) / model columns
- Iteration: MCMC iteration number recorded in the log (sampled every “Sample Period” iterations).
- Lh: Log-likelihood (log L) at the recorded iteration.
- Tree No: Tree index used for the run (e.g., 1 when a single tree is analysed).
- No Off Parmeters: Number of free (active) parameters in the current reversible-jump model at that iteration.
- No Off Zero: Number of rate parameters constrained to zero in the current model at that iteration.
- Model string: Compact reversible-jump model descriptor output by BayesTraits. ‘Z’ indicates a rate parameter estimated (non-zero) in that iteration; ‘0’ indicates a rate fixed to zero.
Transition-rate columns
- qXY (e.g., qab, qba, qac …): Instantaneous transition-rate parameter from state X to state Y, sampled by the RJ-MCMC.
• Table 1 includes qab and qba (two-state model: a ↔ b).
• Table 2 includes all directed rates among a–j (10-state model).
Estimated tip-state columns
- Est <taxon_label> - 1 (e.g., “Est Eurytoma_nr_artemisiae_PDYEURY0015 - 1”):
Sampled state assignment for the specified terminal taxon at each recorded iteration. Values are the state symbols (a–b in Table 1; a–j in Table 2).
Marginal posterior probability columns (node reconstructions)
- Root P(x): Marginal posterior probability that the root is in state x at the recorded iteration (x = a/b for Table 1; x = a–j for Table 2).
- nodeN P(x): Marginal posterior probability that MRCA node N (node1, node2, …) is in state x at the recorded iteration.
• Table 1 includes node1–node18.
• Table 2 includes node1–node23.
Node numbers correspond to the MRCA node labels defined in the BayesTraits run (see the node-tag definitions and the manuscript/Supporting Information for the mapping of node numbers to clades).
Reversible-jump hyperparameter
- RJRates - Mean: Sampled mean value of the reversible-jump (RJ) rates hyperparameter (“RJRates”) reported by BayesTraits at each recorded iteration.
Genomic DNA was extracted from individual specimens using the Qiagen DNeasy Blood & Tissue Kit, applying destructive or non-destructive extraction depending on specimen preservation and voucher needs. Whole-genome shotgun sequencing (150 bp paired-end; ~5 Gb per specimen) was generated by an external provider. Raw reads were quality-filtered with Trimmomatic, and mitochondrial genomes were de novo assembled using NOVOPlasty v4.3 with a partial COI fragment as the seed. Assemblies were curated in Geneious Prime (v2024.1.1) via manual inspection, gap correction, and coverage assessment; when assemblies were fragmented or incomplete, reads were mapped to conspecific or congeneric reference mitogenomes (“Map to Reference”) to recover missing regions. Gene annotation (13 protein-coding genes, rrnL, rrnS, and tRNAs) was performed with MITOS/MITOS2 and validated by comparison to published Chalcidoidea mitogenomes.
To broaden taxon sampling, publicly available raw reads (e.g., WGS or UCE-derived) were downloaded from the NCBI SRA and processed with the same assembly, curation, and annotation workflow to ensure methodological consistency. For downstream comparative analyses, gene regions were aligned per locus using MAFFT (L-INS-i), ambiguously aligned segments were trimmed in MEGA7, and curated alignments were concatenated in SequenceMatrix to generate final matrices (including a mitogenome-only dataset and an expanded matrix incorporating COI-only terminals, where non-COI partitions were coded as missing). If ecological/trophic trait data are included, trait states were compiled from rearing observations and literature/database sources; taxa lacking reliable information were coded as missing (not assigned to an “unknown” state).
