Data from: Advancing phylogenomics in Amaranthaceae sensu stricto: Development and application of a new nuclear target enrichment bait set
Data files
May 05, 2026 version files 39.09 MB
Abstract
Premise: Current phylogenies of Amaranthaceae are inadequately sampled and resolved to reflect the entire evolutionary history of the lineage, which is likely complex due to at least three whole-genome duplication events, occasionally followed by subsequent additional polyploidization events and rapid diversification of individual sub-lineages. To overcome these challenges when reconstructing a phylogeny, we designed a new target enrichment bait set and demonstrated its applicability to the entire Amaranthaceae s.s. lineage.
Methods: We analyzed 12,775 orthologous and low-copy genes from a previous comprehensive transcriptomic study for marker selection. Following a newly developed approach that allows the selection of long exons and thus avoids the assembly of chimeric loci, we selected 1,000 orthologous exons for phylogenomic analyses.
Results: Our in vivo application showed a high locus recovery rate across all major clades of Amaranthaceae s.s., generated a robust phylogenetic tree, and clarified previously ambiguous relationships of the genera Bosea and Charpentiera. Gene tree conflict analysis revealed mainly high levels of gene tree concordance within the lineage, with a couple of notable exceptions.
Discussion: The Amaranthaceae1000 kit will provide the basis for an Amaranthaceae s.s.-wide phylogenetic tree, facilitating future studies on systematics, diversification, and genome evolution within in this economically important lineage.
Dataset DOI: 10.5061/dryad.k3j9kd5m6
Description of the data and file structure
We analyzed a total of 12,775 orthologous and low-copy genes from a previous comprehensive transcriptomic study for marker selection. Following a newly developed approach that allows the selection of long exons and thus avoids the assembly of chimeric loci, we selected 1,000 orthologous exons for phylogenomic analyses. The final bait set targets a total of 1.29 Mbp.
Files and variables
File: Amaranthaceae1000_Design_Files_input-seq.fasta
Description: 2,000 target sequences with a total target size of 2,571,494nt and an average GC content of 42.6%. Targets were softmasked for simple and low complexity repeats against the dicot plant database. Strings of Ns 1-10nt were replaced with T.
File: Amaranthaceae1000_Design_Files_baits-Moderate-25RMpc-noChloro.fasta
Description: 37,727 baits passed Moderate BLAST filtration, <=25% repeat masked, and no BLAST hits to the chloroplast genome.
File: Amaranthaceae1000_Design_Files_bait-120-60.fasta
Description: 39,091 baits designed using 120nt probes and 2x tiling (probe every ~60nt).
File: Amaranthaceae1000_target_file.fas.fasta
Description: Target file for loci extraction, including 29 transcriptomes of Amaranthaceae s.s.
Code/software
All of the code used to select the target loci for bait design can be found at https://github.com/tinakiedaisch/bait_design_from_orthologs
Access information
Other publicly accessible locations of the data:
- The sequence data generated with the bait can be found under the BioProject ID: PRJNA1209683.
