Data from: Contrasting dispersal histories shape distinct evolutionary trajectories between Malesian and pantropical Talipariti (Malvaceae)
Data files
Jul 17, 2026 version files 17.53 MB
-
Dispersal_shapes_Talipariti_evolution_dataset.zip
17.53 MB
-
README.md
3.35 KB
Abstract
The historical biogeography of Talipariti and allied species within the Hibiscus section Azanzae provides a unique system to test dispersal and diversification drivers across Malesia since the Miocene. From this framework, we aim to test whether species distributions were shaped primarily by stochastic long-distance dispersal (LDD) or by stepping-stone dispersal associated with tectonic collisions, by comparing two lineages with contrasting geographic ranges and dispersal abilities: a pantropical sea hibiscus group and an endemic Papuan group.
Authors: FVE, HN, TK & KT
Summary of experimental efforts underlying this datasets and file description:
We combined newly generated sequences with publicly available data from NCBI, the PAFTOL project, the European Nucleotide Archive, and OneKP.
Description of data and file structure
Dispersal_shapes_Talipariti_evolution_dataset.zip
File Samples_working_codes.txt:
Metadata describing the working code and origin of each species in the dataset.
Folder CDS_dataset_for_Time_Calibration
After genome skimming NGS and reads QC, plastomes were assembled (map to reference) and annotated in GENEIOUS PRIME by transferring annotations from the reference plastome (70% similarity threshold), with manual adjustment of coding regions to preserve reading frames. Unassembled positions were coded as “N”.The inverted repeat A (IRA) regions were removed, and the small single-copy (SSC) region was reverse-complemented when necessary. Alignments were generated using MAFFT with the "auto" strategy and filtered with TRIMAL (-gt 0.1 -cons 35). After manual inspection and removal of ambiguous regions, parsimony-informative sites were extracted using AMAS.
File Alignment.phy:
The CDS dataset comprising 75 concatenated plastid coding regions, extracted and filtered using the same pipeline described above.
File Tree.nw:
Maximum likelihood (ML) inference with 1,000 ultrafast bootstrap replicates were performed using IQ-TREE. This plastome CDS alignment and tree was then used for time calibration analysis.
Folder Nuclear_concatenated_dataset
After Target Capture with Angiosperms353, NGS and reads QC, reads were assembled using HYBPIPER with the mega353.fasta target file and a minimum coverage of 3. Paralog detection and recovery statistics were evaluated within HYBPIPER. To reduce alignment noise, we followed the filtering strategy of Pokorny et al. (2024). Phylogenetic inference was conducted using both concatenation and coalescent-based approaches.
Files Nuclear_concatenated_alignment.fasta and concatenated_dataset_partitions.txt:
For concatenation, loci were combined into a supermatrix using AMAS.
File RAxML_nuclear_concatenated_tree.nw:
Maximum-likelihood trees were inferred with RAXML-NG using the best-fit substitution model determined by MODELTEST-NG.
File 267_singlecopy_gene-trees_for_coalescence.nw:
For coalescent analyses, individual gene trees were inferred with IQ-TREE, with branch support assessed from 1,000 bootstrap replicates.
File ASTRALpro2_nuclear_coalescent_tree.nw:
Gene trees were then processed to reduce noise: outlier branches were removed with TREESHRINK, and branches with very low support (<20%) were collapsed using NEWICK UTILITIES. The processed gene trees were used for species tree inference with ASTRAL-PRO2.
File structure:
|- CDS_dataset_for_Time_Calibration/
| |- Alignment.phy
| |- Tree.nw
|- Nuclear_concatenated_dataset/
| |- Nuclear_concatenated_alignment.fasta
| |- concatenated_dataset_partitions.txt
| |- RAxML_nuclear_concatenated_tree.nw
| |- 267_singlecopy_gene-trees_for_coalescence.nw
| |- ASTRALpro2_nuclear_coalescent_tree.nw
