Data and code from: Genomic contingence beneath ecological convergence: The tempo and mode of gene loss in parasitic bilaterians
Data files
Jul 10, 2026 version files 79.60 MB
-
DataFile1_SequenceData.zip
78.67 MB
-
DataFile2_MolClockAlignment.nexus
644.15 KB
-
DataFile3_AncestralStates.zip
54.96 KB
-
DataFile4_Code_For_Analyses.zip
212 KB
-
README.md
13.91 KB
Abstract
This dataset contains all data and code necessary to reproduce the results of the paper "Genomic contingence beneath ecological convergence: The tempo and mode of gene loss in parasitic bilaterians." Specifically, this dataset contains a master summary file (also included as a supplement to the main paper), curated miRNA and protein sequence data, reconstructed ancestral gene complements, and a sequence alignment for molecular clock calculation. These data are free for reuse without restrictions and do not contain any sensitive or legally restricted information.
Parasitism has independently evolved hundreds of times among metazoans. Nonetheless, parasites have explored only a limited range of ecologies, and they display frequent convergence in morphological, behavioral, and life-history traits. Although gene loss in parasitic species is well documented, it is not known if gene loss converges along the same lines as other traits. To test for convergent gene loss, we characterized the housekeeping, regulatory, and DNA-repair complements of 48 bilaterian species, including 20 parasites belonging to six different bilaterian phyla. We found that the canonical parasitic strategies do not display characteristic tempos or modes of gene loss. Further, the accelerated rates of gene loss seen in some parasites were almost always shared with their free-living relatives, indicating that the rate of loss increased before parasitism arose. Therefore, the convergent ecological strategies and adaptations that have arisen in distantly related parasitic lineages overlay contingent gene losses which largely reflect their phylogenetic history. These results have important implications for how ecologists and evolutionary biologists should model the acquisition of parasitism, especially regarding the long-held assumption that reversion from a parasitic to a free-living state is impossible.
Dataset DOI: 10.5061/dryad.4j0zpc8t9
Description of the data and file structure
These are the data and code underlying the paper "Genomic Contingence Beneath Ecological Convergence: The Tempo and Mode of Gene Loss in Parasitic Bilaterians" by A. Clarke, L. Bozek, A.I. Babalola, L. Wyrozemski, V.M. Paynter, J. Vinther, A. Donath, J. Olesen, H. Glenner, M.A. McPeek, T.M. Farrell, R.R. Fitak, P. Martinez, B. Fromm, and K.J. Peterson.
Files and variables
File: DataFile1_SequenceData.zip
Description: An archive containing the curated sequence data for this project. These are organized thus:
DataFile1.1 contains miRNA data in two sub-folders:
- DataFile_1.1.1 contains a series of .csv files (labeled 1.1.1a through 1.1.1q) that represent the annotated complement of each newly curated species. The columns of these .csv files are the name of the gene, the complete precursor sequence, the mature sequence, the star (or co-mature) sequence, whether there is a 3' non-templated uridine added to the 3p arm, and any important comments about that microRNA gene.
- DataFile_1.1.2 contains the newly sequenced sRNAs for Raillietiella orienatlis (1.1.2a), Vargula norvegica (1.1.2b), Sacculina carcini (1.1.2c), Pollicipes pollicipes (1.1.2d), Stylops ovinae (1.1.2e).
DataFile1.2 contains four sub-folders of .txt files with fasta-formatted protein sequences:
- DataFile_1.2.1 contains the master files used to query the transcription factor (1.2.1a), RNA-binding protein (1.2.1b), and DNA repair protein (1.2.1c) complements from each of the sampled species. It also contains the query file used to recover the sequences of the 19 genes underlying our molecular clock analysis (1.2.1d).
- DataFile_1.2.2 contains the newly curated transcription factor sequences of the sampled species. Each file in this folder is named "TXF_Genus," where Genus is the genus name of the relevant species.
- DataFile_1.2.3 contains the newly curated RNA binding protein sequences of the sampled species. Each file in this folder is named "RBP_Genus," where Genus is the genus name of the relevant species.
- DataFile _1.2.4 contains the newly curated DNA repair protein sequences of the sampled species. Each file in this folder is named "DNARP_Genus," where Genus is the genus name of the relevant species.
The member .txt files of DataFile_1.2.2, DataFile_1.2.3, and DataFile_1.2.4 identify duplicated sequences in each species' complement using either a number or a lowercase letter, depending on whether the name of the relevant gene ends with a letter or a number. Thus, duplicates of ALDOB, which ends with a letter, would be labeled "ALDOB1," "ALDOB2," and so on, and duplicates of CDK7, which ends with a number, would be labeled "CDK7a," "CDK7b," and so on.
File: DataFile2_MolClockAlignment.nexus
Description: A Nexus file containing the protein alignment used to perform a molecular clock analysis of the selected species.
File: DataFile3_AncestralStates.zip
Description: An archive containing .csv files with the reconstructed BUSCO (DataFile3.1.1 for raw BUSCO presence/absence data and DataFile3.1.2 for these raw data converted to percentage terms), transcription factor (DataFile3.2), RNA binding protein (DataFile3.3), miRNA (DataFile3.4), and DNA repair protein (DataFile3.5) complements present at each of the ancestral nodes in the phylogeny of our sampled species. The columns in these .csv files are the ancestral nodes uniting our sampled species; the rows are individual genes. This archive also contains the spreadsheet used to calculate phylogenetic independent contrast and ancestral GPC1 values (DataFile3.6). Rows in this file are species or ancestral nodes, and the variables are the fraction of losses of BUSCO, transcription factor, RNA-binding protein, DNA repair protein, and microRNA genes; average non-BUSCO loss; the raw and scaled branch lengths connecting nodes and species; reconstructed GPC1 values (for ancestral nodes only); differences in the BUSCO and average non-BUSCO losses between the descendants of each node (for ancestral nodes only); the sum of the scaled branch lengths (for ancestral nodes only); and the phylogenetic independent contrast in BUSCO and average non-BUSCO losses between the descendant species of each ancestral node (for ancestral nodes only).
Where a well-established clade name exists (e.g., Deuterostomia, Mollusca, etc), ancestral nodes are identified by that name (e.g., LCDeuterostomia for the ancestral node representing the last common ancestor of all Deuterostomes). Where such names do not exist, the nodes are identified using the names of their descendants, with descendant clades that have well-established names identified by name and descendant species identified by the three-letter code used by MirGeneDB. Thus, the node LCBmaDmd corresponds to the last common ancestor of Brugia malayi and Dracunculus med*inensis*; the node LCAvaSnePla corresponds to the last common ancestor of Adineta vaga, Seison nebaliae, and Pomphorhynchus laevis; the node LCSnePla corresponds to the last common ancestor of Seison nebaliae and Pomphorhynchus laevis; the node LCLpoIsc corresponds to the last common ancestor of Limulus polyphemus and Ixodes scapularis; the node LCLanMollusca corresponds to the last common ancestor of Lingula anatina and the molluscs; and the node LCCelTylenchida corresponds to the last common ancestor of Caenorhabditis elegans and the Tylenchida.
File: DataFile4_Code_For_Analyses.zip
Description: An archive containing the codes used to run the various statistical analyses performed in this study, organized in six sub-folders. The data tables in this archive do not contain any novel information; they merely re-organize various other tables from throughout the supplementary material for convenient analysis with R. These tables are usually in .csv format; the few tables in .xlsx format can be accessed without Excel itself because they are designed to be imported directly into R via the readxl package, and can be reviewed and edited therein. Each of the analyses contained in a sub-folder is self-contained; the outputs of one are not needed to run the next.t
- DataFile4.1 contains the code (Confirmation.r) and formatted data (Data.xlsx) needed to perform our tests for systematic biases in the patterns of losses across species. The variables in Data.xlsx are the percent gains of transcription factor, RNA-binding protein, miRNA, and DNA repair genes; the percent losses of these genes; the percent losses of BUSCO genes; the average percent losses of non-BUSCO genes; FractionFN (the percent of false-negative microRNAs in a species; WGD (= 1 if a species went through a whole-genome duplication event, = 0 otherwise); miRNABurst (= 1 if a species experienced a burst of miRNA family-level innovation); the N50 value of the genome assembly; and NosRNAs (= 1 if no sRNA libraries were available for a species, = 0 otherwise). The rows of Data.xlsx are species.
- DataFile4.2 contains the code (SlopeTest.r) and formatted data (Data.xlsx) needed to test for differences in the relative extent of losses in different gene categories between species with different ecological properties. The variables in Data.xlsx are the percent loss of BUSCO genes (or the contrast in percent BUSCO loss, for ancestral nodes), the average percent loss of non-BUSCO genes (or the contrast in average percent non-BUSCO loss, for ancestral nodes), New (= 1 if a species is newly sampled, = 0 otherwise), Parasite (= 1 if a species is a parasite, = 0 otherwise, Strategy (parasitoid vs castrator vs direct vs trophic vs vector vs micropredator vs free), Stage (adult vs juvenile vs permanent vs free), Definitive_Host (chordate vs arthropod vs plant vs free), and Contrast (= 1 if a row contains phylogenetic-indpendent contrast information, = 0 otherwise). The rows of Data.xlsx are species and ancestral nodes (identified with the same names as in the contents of DataFile3). The contrast values are derived from the calculations shown in DataFile3.6.
- DataFile4.3 contains the code (GOEnrichment.r) and formatted data (Data.csv) needed to test for differential enrichment of losses within Gene Ontology classes of BUSCO genes across species. The columns in Data.csv are the individual genes (identified as either present = 1 or absent = 0 in each species); the rows are the species. The R package used in this step required independent variables to be read in separately from the main data, so they are organized in a series of additional .csv files (GOMap.csv, Coding.csv). GOMap.csv maps BUSCO genes to their gene ontology terms. This file has columns BUSCOID (the OrthoDB v10 ID assigned to each metazoan BUSCO gene), name (a descriptive name of the gene), annotation (which indicates whether a gene has been removed from our analysis due to being a transcription factor, regulatory RNA-binding protein, or DNA repair protein), and CC_GO_Term (the GO term or terms associated with a given gene); the rows are the BUSCO genes. Coding.csv has the variables Parasite (parasite vs free), Strategy (parasitoid vs castrator vs direct vs trophic vs vector vs micropredator vs free), Stage (adult vs juvenile vs permanent vs free), Definitive_Host (chordate vs arthropod vs plant vs free), and Phylum (phylum of a given species); the rows are the species. GeneGroup.csv is used in the visualization of our results to group GO terms associated with similar organelles in the final heatmap figure; its variables are GO (a given GO term) and Type (the organelle).
- DataFile4.4 contains the code (PCoA.r) and formatted data (All_presence.csv, All_absence.csv) needed to perform our principal coordinates analyses of gene presences and absences across species. The variables in each of these spreadsheets are the genes; the rows are each species. A "1" value in All_presence.csv indicates that a given gene is present in a given species; a "1" value in All_absence.csv indicates that the gene has been lost in that species. The R package used in this analysis required independent variables to be read in separately from the raw data. These variables are contained in the Coding.csv file, which is identical to the spreadsheet of the same name in DataFile4.3. This folder also contains the code (PCoA.m) needed to convert the results of this analysis into the figure used in the main paper.
- DataFile4.5 contains the code (GWAS.r) and formatted data (Data.csv) needed to perform our repeated Fisher tests of the rate of loss of each sampled gene across species with different ecologies. The columns in Data.csv are the individual genes (identified as either present or absent in each species); the rows are the species. The R package used in this analysis required independent variables to be read in separately from the raw data. These variables are contained in the Coding.csv file, which is identical to the spreadsheets of the same name in DataFile4.3 and DataFile4.4. Map.csv is used to organize genes for visualization; its variables are the gene name, psChrom (and indicator of a gene's functional category; 1 = transcription factor, 2 = RBP, 3 = miRNA, 4 = DNARP, 5 = BUSCO), and psPos (an arbitrary spacer used to ensure genes are adequately separated in the final Manhattan plots; equal to 100 times the ordinal position of a gene within its group; e.g., the first transcription factor has a value of 100, and the third miRNA has a value of 300).
- DataFile4.6 contains the code (Shifts.r), time-scaled phylogeny (Final.tre), and appropriately formatted data (GPC1.csv) needed to perform our assessment of changes in the rate of loss of genes across species with different ecologies and phylogenetic positions. The variables of GPC1.csv are Species (the names of the species) and GPC1 (the species' positions on the first principal components axis). This folder also contains the MATLAB code (PrincipalComponents.m) needed to perform the principal components analysis that generates the GPC1 values.
File: SupplementalFiles.zip
Description: An archive containing the two supplemental files submitted alongside, and also available with, our main paper. The SuppFile1 archive contains a number of .csv spreadsheets containing summary data describing each species' gene complements, as well as the detailed statistical results tables produced by our analyses. These .csv files are named and organized to follow, as closely as is possible, the original .xlsx version of this file available with the main paper. SuppFile2 is a Word document with a small number of supplementary figures depicting the results of the intermediate processing steps in our analyses.
Code/software
All scripts were written in MATLAB or R. MATLAB scripts used base MATLAB; R scripts used the car, devtools, readxl, fastDummies, dplyr, tiddyr, ggplot2, BiocManager, clusterProfiler, GO.db, pheatmap, ape, vegan, phytools, and nmle packages.
Access information
Other publicly accessible locations of the data:
- miRNA data have also been deposited in MirGeneDB (link: https://mirgenedb.org/) and will appear on their main page following their next update (expected January 2027). Until that time, they can also be retrieved upon communication with MirGeneDB web managers. The contents of the SupplementalFiles.zip archive are also available from the main article published in Evolution.
Data were derived from the following sources:
- These data were not derived from other sources.
