Data from: A step in the deep evolution of Alvinellidae (Annelida: Polychaeta): A phylogenomic comparative approach based on transcriptomes
Data files
Aug 02, 2023 version files 11.47 MB
-
Data.zip
10.51 MB
-
README.md
11.89 KB
-
supplementary_data.pdf
948.86 KB
Aug 24, 2026 version files 149.28 MB
-
Data.zip
148.92 MB
-
README.md
18.82 KB
-
supplementary_figures.pdf
342.24 KB
Abstract
The Alvinellidae are a family of worms that are endemic to deep-sea hydrothermal vents in the Pacific and Indian Oceans. These annelid worms, a sister group to the Ampharetidae, occupy a wide range of thermal habitats. The family includes the most thermotolerant marine animals described to date, such as the Pompeii worm Alvinella pompejana, and other species living at much lower temperatures such as Paralvinella grasslei or Paralvinella pandorae. The phylogeny of this family has not been studied extensively. It is, however, a complex case where molecular phylogenies give conflicting results, especially concerning the monophyletic or polyphyletic character of the genus Paralvinella.
We carried out a comprehensive study of the phylogeny of this family using the best molecular data currently available from RNAseq datasets. The study is based on the assembly of several hundred transcripts for 11 of the 14 species currently described or in description. The results obtained by the most popular phylogenetic inference models (gene concatenation with maximum likelihood, or coalescent-based methods from gene trees) are compared using a series of ampharetid and terebellid outgroups.
Our study shows that a high number of gene trees support the hypothesis of the monophyly of the Paralvinella genus, as initially proposed by Desbruyères and Laubier, in which the species Paralvinella pandorae and Paralvinella unidentata are more closely related within the subgenus Nautalvinella. However, the global phylogenetic signal favors the hypothesis of paraphyly for this genus, with P. pandorae being sister species of the other Alvinellidae. Gene trees separated equally between these two hypotheses, making it difficult to draw conclusions about the initial split of the MCRA as different genomic regions seem to have very different phylogenetic stories. According to molecular dating, the radiation of the Alvinellidae was rapid and took place in a short period of time between 70 and 80 million years ago. This is reflected at the genomic scale by high rates of incomplete lineage sorting between the first ancestral lineages with probable gene transfers between the ancestors of Alvinella, Nautalvinella, and the rest of the Paralvinella lineages.
Brief summary of dataset contents.
Data archive for:
A step in the deep evolution of Alvinellidae (Annelida: Polychaeta): a phylogenomic comparative approach based on transcriptomes
by Pierre-Guillaume Brun, Stéphane Hourdez, Marion Ballenghien, Yadong Zhou, Jean Mary, Didier Jollivet
The Alvinellidae are a family of worms that are endemic to deep-sea hydrothermal vents in the Pacific and Indian Oceans. These annelid worms, a sister group to the Ampharetidae, occupy a wide range of thermal habitats. The family includes the most thermotolerant marine animals described to date, such as the Pompeii worm Alvinella pompejana, and other species living at much lower temperatures such as Paralvinella grasslei or Paralvinella pandorae. The phylogeny of this family has not been studied extensively. It is, however, a complex case where molecular phylogenies give conflicting results, especially concerning the monophyletic or polyphyletic character of the genus Paralvinella. We carried out a comprehensive study of the phylogeny of this family using the best molecular data currently available from RNAseq datasets. The study is based on the assembly of several hundred transcripts for 11 of the 14 species currently described or in description. The results obtained by the most popular phylogenetic inference models (gene concatenation with maximum likelihood, or coalescent-based methods from gene trees) are compared using a series of ampharetid and terebellid outgroups. Our study shows that a high number of gene trees support the hypothesis of the monophyly of the Paralvinella genus, as initially proposed by Desbruyères and Laubier, in which the species Paralvinella pandorae and Paralvinella unidentata are more closely related within the subgenus Nautalvinella. However, the global phylogenetic signal favors the hypothesis of paraphyly for this genus, with P. pandorae being sister species of the other Alvinellidae. Gene trees separated equally between these two hypotheses, making it difficult to draw conclusions about the initial split of the MCRA as different genomic regions seem to have very different phylogenetic stories. According to molecular dating, the radiation of the Alvinellidae was rapid and took place in a short period of time between 70 and 80 million years ago. This is reflected at the genomic scale by high rates of incomplete lineage sorting between the first ancestral lineages with probable gene transfers between the ancestors of Alvinella, Nautalvinella, and the rest of the Paralvinella lineages.
Alvinellid species were collected during several oceanic cruises from 2004 to 2019 using the N/O L’Atalante and either the ROV Victor6000 or the manned submersible Nautile with the exception of Paralvinella palmiformis which was sampled during the jdfR cruise on Juan de Fuca. Alvinella caudata, Paralvinella grasslei and Paralvinella pandorae irlandei were collected from hydrothermal vents on the EPR à 2,550m depth during the Mescal Oeanographic cruise in 2010. Other alvinellid species and the vent terebellid were sampled on different vent sites of the western Pacific back-arc basins during the Chubacarc 2019 cruise with the tele-manipulated arm of the ROV Victor6000 and brought back to the surface in an insulated basket. Paralvinella hessleri was sampled from Fenway in the Manus basin, Paralvinella fijiensis from Big Papi (Manus Basin) and Tu’i Malila (Lau Basin), Paralvinella unidentata from Fenway (Manus Basin), Paralvinella sp. nov. from the volcano South Su (Manus Basin) and the vent terebellid Terebellidae gen. sp. (not yet described) at Snow Cap (Manus Basin). Finally, Paralvinella mira was sampled from the Wocan vent field in the northwest Indian Ocean on the Carlsberg Ridge by HOV Jialong during the DY38 cruise in March 2017 (Han et al., 2021). Melinna palmata and Neoamphitrite edwardsii, which are shallow-water species, were respectively sampled in the bay of Morlaix and Roscoff, France. Total RNA extraction from flash-frozen tissue were performed with Trizol. RNAseq libraries were produced at Genome Qu ́ebec following a polyA purification of mRNAs, and sequenced using the Novaseq technology.
Schema of data archive
Data.zip contains:
genes_alignments
|--- nucleotides
| |--- raw: contains initial gene alignments (fasta format) and script align.py, used to align genes with the java MACSE software.
| |--- filtered: contains gene alignments (fasta format) checked for composition bias (IQ-TREE symtest). Also contains log files for symtest.
| --- filtered_reduced: same alignments as filtered (fasta format) but sequences for Pfijiensis3, Pfijiensis4 and Punidentata4 are removed.
--- amino_acids
|--- raw: contains initial gene alignments (fasta format) and script align.py, used to align genes with the python PROBCONS software.
|--- filtered: contains gene alignments (fasta format) checked for composition bias (IQ-TREE symtest). Also contains log files for symtest.
--- filtered_reduced: same alignments as filtered (fasta format) but sequences for Pfijiensis3, Pfijiensis4 and Punidentata4 are removed.
datation
|--- nt
| |--- T6: contains nucleotide supergene concatenation (fasta format), T6 tree topology file (trees format), and calibration/outgroup files used for datation. The scripts to run the datation in Phylobayes are also provided (.sh files for SLURM job). *cir* files refer to CIR model results, *ln* files to log-normal results, and *ugam* to uncorrelated gamma model. script, concatenation file, tree, calibration, output, cir ln ugam. *chronogram files give the results of Phylobayes' datation.
| |--- T7: contains nucleotide supergene concatenation (fasta format), T7 tree topology file (trees format), and calibration/outgroup files used for datation. The scripts to run the datation in Phylobayes are also provided (.sh files for SLURM job). *cir* files refer to CIR model results, *ln* files to log-normal results, and *ugam* to uncorrelated gamma model. script, concatenation file, tree, calibration, output, cir ln ugam. *chronogram files give the results of Phylobayes' datation.
| --- T9: contains nucleotide supergene concatenation (fasta format), T9 tree topology file (trees format), and calibration/outgroup files used for datation. The scripts to run the datation in Phylobayes are also provided (.sh files for SLURM job). *cir* files refer to CIR model results, *ln* files to log-normal results, and *ugam* to uncorrelated gamma model. script, concatenation file, tree, calibration, output, cir ln ugam. *chronogram files give the results of Phylobayes' datation.
--- aa
|--- T6: contains nucleotide supergene concatenation (fasta format), T6 tree topology file (trees format), and calibration/outgroup files used for datation. The scripts to run the datation in Phylobayes are also provided (.sh files for SLURM job). *cir* files refer to CIR model results, *ln* files to log-normal results, and *ugam* to uncorrelated gamma model. script, concatenation file, tree, calibration, output, cir ln ugam. *chronogram files give the results of Phylobayes' datation.
|--- T7: contains nucleotide supergene concatenation (fasta format), T7 tree topology file (trees format), and calibration/outgroup files used for datation. The scripts to run the datation in Phylobayes are also provided (.sh files for SLURM job). *cir* files refer to CIR model results, *ln* files to log-normal results, and *ugam* to uncorrelated gamma model. script, concatenation file, tree, calibration, output, cir ln ugam. *chronogram files give the results of Phylobayes' datation.
--- T9: contains nucleotide supergene concatenation (fasta format), T9 tree topology file (trees format), and calibration/outgroup files used for datation. The scripts to run the datation in Phylobayes are also provided (.sh files for SLURM job). *cir* files refer to CIR model results, *ln* files to log-normal results, and *ugam* to uncorrelated gamma model. script, concatenation file, tree, calibration, output, cir ln ugam. *chronogram files give the results of Phylobayes' datation.
candidate_trees: contains trees' topologies used for constrained topologies optimizations and scoring
scripts: contains scripts used for model testing and the generation of figures
results: contains results presented in the main text
|--- * Table_ML_Astral_result: summary of ML scores and quartet scores for the different topologies, with detailed p-values according to the AU test compared to the best scoring topology (Numbers format). See below for details.
|--- Coalescent_constraint_test: species trees' topologies' AU test with constrained gene trees' topologies
| |--- * ASTRAL_constraint_test.txt: AUtest pvalues and RSS for each topology (log from the script AUtest_contraint.py)
| |--- * ASTRAL_matrix_quartet.txt: temporary file giving the number of quartet in each topology
| --- * ASTRAL_matrix.txt: temporary file giving the number of identical quartets in pairs of topologies
|--- gene_trees
| |--- trees_nucleotides
| | |--- constrained: contains evaluation of the 15 topologies for all genes
| | | |--- * sitesinformatifs.csv: number of parcimony-informative sites for all genes, distinguishing between 15 topologies (csv format). topologiesML.csv and topologiesPP.csv: ML scores and posterior probability of 15 topologies for all genes (csv format).
| | | |--- * adaptingtrees.py: script used by iqtree_genes.sh
| | | |--- * iqtree_genes.sh: ML scoring of topologies
| | | --- iqtreelogs
| | | --- *.iqtree: IQ-TREE logs (containing ML scores) of 15 topologies for all genes
| | --- freerunning: gene tree optimization without constraints
| | |--- * gene_trees_nucleotides.tre: all trees obtained from free-running gene trees optimization
| | |--- * gene_trees_nucleotidesBS10.tre: all trees obtained from free-running gene trees optimization with branches below 10% bootstrap confidence collapsed
| | |--- * Astral_nucleotides_free_running.tre: ASTRAL species tree optimized from gene trees
| | |--- * Astral_log.txt: terminal output of ASTRAL optimization
| | --- gene_trees
| | |--- * model_count.txt: counting the number of times a model was used for gene_tree optimization (txt format)
| | |--- * bestmodel.txt: model used for each gene tree
| | |--- * log: IQ-TREE log files
| | --- * treefile: optimized ML tree for each gene
| --- trees_amino_acids
| |--- constrained
| | |--- * sitesinformatifs.csv: number of parcimony-informative sites for all genes, distinguishing between 15 topologies (csv format). topologiesML.csv and topologiesPP.csv: ML scores and posterior probability of 15 topologies for all genes (csv format).
| | |--- * adaptingtrees.py: script used by iqtree_genes.sh
| | |--- * iqtree_genes.sh: ML scoring of topologies
| | --- iqtreelogs
| | --- *.iqtree: IQ-TREE logs (containing ML scores) of 15 topologies for all genes
| --- freerunning
| |--- * gene_trees_amino_acids.tre: all trees obtained from free-running gene trees optimization
| |--- * gene_trees_amino_acidsBS10.tre: all trees obtained from free-running gene trees optimization with branches below 10% bootstrap confidence collapsed
| |--- * Astral_amino_acids_free_running.tre: ASTRAL species tree optimized from gene trees
| |--- * Astral_log.txt: terminal output of ASTRAL optimization
| --- gene_trees
| |--- * model_count.txt: counting the number of times a model was used for gene_tree optimization (txt format)
| |--- * bestmodel.txt: model used for each gene tree
| |--- * log: IQ-TREE log files
| --- * treefile: optimized ML tree for each gene
|--- Coalescent_freerunning_and_test
| |--- nucleotides
| | |--- * TX.log: terminal output for 15 species tree topologies, based on free-running gene trees
| | |--- * AUtestcoalescent.py: AUtest for coalescent-resolved species tree, to compare candidate topologies scores with a reference topology score. requires Newick utility (python script, usage: python AUtest_ASTRAL.py folder_gene_trees reference_tree folder_candidate_trees)
| | |--- * AUtest.txt: AUtest results for coalescent species topology with free-running gene trees
| | --- treesAU
| | * gene trees and gene trees with branches under 10 bootstrap collapsed (BS10)
| --- amino_acids
| |--- * TX.log: terminal output for 15 species tree topologies, based on free-running gene trees
| |--- * AUtestcoalescent.py: AUtest for coalescent-resolved species tree, to compare candidate topologies scores with a reference topology score. requires Newick utility (python script, usage: python AUtest_ASTRAL.py folder_gene_trees reference_tree folder_candidate_trees)
| |--- * AUtest.txt: AUtest results for coalescent species topology with free-running gene trees
| --- treesAU
| * gene trees and gene trees with branches under 10 bootstrap collapsed (BS10)
--- Concatenation_free_running_and_test
|--- nucleotide
| |--- * partition.txt: initial partition file used for partitioned phylogeny (txt format)
| |--- * .fa: concatenated supergene (fasta format)
| |--- * .treefile: IQ-TREE optimized ML tree in partition model
| |--- * .log: IQ-TREE log file
| |--- * AUtestlog.txt: AUtest results
| --- AU_test
| |--- * .iqtree, .log: IQ-TREE output for ML optimization from concatenated supergene (non partition model) for topologies T1-T15 and free-running
| |--- * candidate_trees.tree: contains 15 topologies, with first topology being the reference topology for AU test
| --- * AUtestcontraint.sh: script for IQ-TREE AU test
--- amino_acids
|--- * partition.txt: initial partition file used for partitioned phylogeny (txt format)
|--- * .fa: concatenated supergene (fasta format)
|--- * .treefile: IQ-TREE optimized ML tree in partition model
|--- * .log: IQ-TREE log file
|--- * AUtestlog.txt: AUtest results
--- AU_test
|--- * .iqtree, .log: IQ-TREE output for ML optimization from concatenated supergene (non partition model) for topologies T1-T15 and free-running
|--- * candidate_trees.tree: contains 15 topologies, with first topology being the reference topology for AU test
--- * AUtestcontraint.sh: script for IQ-TREE AU test
Table_ML_Astral_result:
Each line refers to a tree topology (from 1 to 15)
2 first columns: nt and aa refer to the likelihood of the optimized tree under the constrained topology from nucleotide or amino acid sequences.
Next 2 columns: astral nt and astral aa give the number of quartets agreeing between individual gene topologies and the constrained species topologies.
The last line gives the maximum of each column, which corresponds to the best score.
last 8 columns: nt, aa, astral nt and astral aa give the difference between the topology and the best scoring topology in each case. the pval columns give the p-value that the topology is as good as the best topology according to the AU test. a p-value > 5% indicates that the topology is able to perform better than the reference best-scoring topology in the case of site or gene resampling/bootstraping.
Supplementary figures
- supplementary_figures.pdf: supplementary figures referred to in the main text (pdf format)
Scripts details
- missing.py: remove genes with less than 19 species, and not containing at least Apompejana or Acaudata, 1 Punidentata, and Ppandorae. Warning: the script actually deletes files. This was applied at the very beginning of the process, to obtain genes in gene_alignments/-/raw
- filter.py: removes gene fragments that are too short after alignments, coded as taillefragmin in the script. Also removes alignments containing too few sequences (<nbseqmin) after filtering.
- testcomposition.sh: filtering gene alignments that violate phylogenetic assumptions (composition homogeneity, stationarity). Loops symtest implemented in IQ-TREE, and removes the most biased sequences until the test is a succeed (p-value>5%). It cannot remove sequences of P. pandorae, P. unidentata, P. caudata or P. pompejana. if the remaining sequences are too few (<20), the gene is discarded.
- concatenation.py: used to concatenate gene alignments into 1 supergene. Produced the concatenated alignments and the partition file giving the first and last residue position of each gene in the supergene
- removefijuni.py: removes Pfijiensis3, Pfijiensis4 and Punidentata4 from input fastas. Used to remove very noisy quartets between closely-related individuals to focus quartet scoring variation by ASTRAL on the root of the Alvinellidae
- adaptingtrees_fijiuni.py : removes Pfijiensis3, Pfijiensis4 and Punidentata4 from input tree.
- AUtest_ASTRAL.py: AU test on species trees obtained from free-running gene trees with ASTRAL. Same as AUtestcoalescent.py
- toposintesinf.py: used to produce figure 3, contains 3 parts: ML counting (writes topologiesML.csv and topologiesPP.csv), Parsimony sites counting (writes sitesinformatifs.csv), Figures (uses the generated CSV files to produce figure 3)
- AUtest_contraint.py: performs AU test from Euclidean distances, as shown in Figure 4. needs a folder containing candidate trees. The reference topology for AU test (T7), and the paths to topologiesPP.csv (nucleotide and amino acids) are hard-coded in the script. Also produces the actual Figure 4.
- countbias.py: counts the proportion of different residues (amino acids or nucleotides) to produce figS6. The script will work (produce a table) both on protein or DNA alignments, but the results are only valid for the corresponding nature residue (e.g., A for proteins is Alanine, but Adenosine for DNA)
- iqtree_genes.sh: phylogeny model testing for each gene, and ML scoring of candidate trees. Needs adaptingtrees.py and selectmodel.py
- selectmodel.py: extracts from IQ-TREE logs the best phylogeny model, according to the BIC
- adaptingtrees.py: utility script that writes a phylogenetic tree containing only a subset of species. Input: phylogenetic tree, fasta file containing a subset of species of the phylogenetic tree.
- tree.py: used by other scripts, extracts the structure (topology and branch lengths) of an input tree file
- scriptagecir.sh/scriptsageln.sh/scriptagegamma.sh: SLURM jobs to use Phylobayes datation models, either CIR model, log-normal, or uncorrelated gamma.
Alvinellid species were collected during several oceanic cruises from 2004 to 2019 using the N/O L’Atalante and either the ROV Victor6000 or the manned submersible Nautile with the exception of Paralvinella palmiformis which was sampled during the jdfR cruise on Juan de Fuca. Alvinella caudata, Paralvinella grasslei and Paralvinella pandorae irlandei were collected from hydrothermal vents on the EPR à 2,550m depth during the Mescal Oeanographic cruise in 2010. Other alvinellid species and the vent terebellid were sampled on different vent sites of the western Pacific back-arc basins during the Chubacarc 2019 cruise with the tele-manipulated arm of the ROV Victor6000 and brought back to the surface in an insulated basket. Paralvinella hessleri was sampled from Fenway in the Manus basin, Paralvinella fijiensis from Big Papi (Manus Basin) and Tu’i Malila (Lau Basin), Paralvinella unidentata from Fenway (Manus Basin), Paralvinella sp. nov. from the volcano South Su (Manus Basin) and the vent terebellid Terebellidae gen. sp. (not yet described) at Snow Cap (Manus Basin). Finally, Paralvinella mira was sampled from the Wocan vent field in the northwest Indian Ocean on the Carlsberg Ridge by HOV Jialong during the DY38 cruise in March 2017 (Han et al., 2021). Melinna palmata and Neoamphitrite edwardsii, which are shallow-water species, were respectively sampled in the bay of Morlaix and Roscoff, France. Total RNA extraction from flash-frozen tissue were performed with Trizol. RNAseq libraries were produced at Genome Qu ́ebec following a polyA purification of mRNAs, andsequenced using the Novaseq technology.
Changes after Aug 2, 2023:
Added all scripts used to perform the study, including AU test after reviewer request and molecular datation. Added raw genes alignments. Added logs of IQ-tree for ML tree search, composition test, AU test, Model finder, in free running and partitionned models. Added ASTRAL logs. Added final supplementary figures. Translated some comments in the scripts.
