Skip to main content
Dryad

Benchmarking assembly-free k-mer methods for species identification in complex plant groups: A case study in Populus

Abstract

This dataset contains the analytical data associated with the study evaluating assembly-free k-mer methods for species identification in the genus Populus. Using whole-genome resequencing data from 235 Populus individuals across 36 species, we established a SNP-based phylogenetic species framework as a benchmark and compared the identification performance of complete chloroplast genomes and an assembly-free k-mer workflow (Bindash). The dataset includes: (1) the filtered and LD-pruned SNP dataset used to construct the benchmark phylogeny; (2) genetic distance matrices derived from SNPs, k-mer sketches, and chloroplast genomes; and (3) Neighbor-Joining tree files in Newick format. These data support the finding that k-mer-based identification achieves 91 % monophyly and 99 % nearest-neighbor accuracy at sequencing depths as low as 0.2×, providing a resource-efficient alternative to traditional barcoding methods for taxonomically complex plant groups.