Data from: Southern Iberia as a hotspot of wild grapevine genetic diversity
Data files
Aug 07, 2026 version files 237.16 KB
-
Admixture_Vsylvestris.zip
4.72 KB
-
Heterozygosity_Vsylvestris.zip
3.91 KB
-
Phylogenetic_tree_Vsylvestris.zip
1.97 KB
-
README.md
2.47 KB
-
Sequences_Vsylvestris.zip
1.46 KB
-
Treemix_Vsylvestris.zip
222.63 KB
Abstract
The commercial interest of grapevines (Vitis vinifera L.) has prompted numerous studies on their origin and genetic resources in the context of global change. However, genomic-scale information on diversity patterns and genetic structure in southwestern Europe remains scarce. This study infers the genetic structure, gene flow events between genetic groups, and genetic refugia of Vitis vinifera spp. sylvestris in the Iberian Peninsula. We reanalyzed a set of 137 complete genomes of V. vinifera spp. sylvestris. After variant calling, validation and annotation, we obtained a high-quality SNP dataset. Using these markers, we performed phylogenetic and population structure analyses to determine the number and spatial distribution of genetic groups and their contact zones. Next, we inferred the timing and directionality of gene flow events between groups. Finally, heterozygosity and allele rarity were estimated to identify populations with high conservation value. We detected three major ancestral populations and four putative genetic refugia in the south of the Iberian Peninsula. Demographic analyses indicate sustained gene flow between ~21,000 and ~7,000 years ago from a North African ancestral group into Iberian wild populations in the south. Heterozygosity and allele rarity analyses identified populations of high conservation value in a variety of areas within the Iberian Peninsula.We identify the biogeographical factors behind the long-known singularity of wild Iberian grapevines. The southern Iberian Peninsula is a hotspot of genetic diversity for wild grapevines, hosting three ancestral populations and multiple contact zones that acted as micro‑refugia. The current genetic variability of Iberian wild grapevines is best explained by natural, climate-driven gene flow between African lineages with Middle Eastern origin and Iberian groups. These contacts were favored by climatic conditions during the late Pleistocene (~21,000 years) and early Holocene (~8,300 years). Our results dismiss a significant anthropogenic influence during Neolithic domestication for explaining the genetic composition of Iberian wild grapevine genotypes.
This file includes 5 subfolders containing files and code generated during the development of this work. Individual distribution data and some code are not publicly available due to privacy or ethical restrictions.
Description of the data and file structure
The Sequences_Vsylvestris.zip subfolder includes: the filtered_sequences_Vsylvestris.csv file. This file includes the following variables: sample (ID of the filtered sequences), Mapped_Reads(%) (mapping percentage), Duplicate(%) (duplication rate), and Avg_Depth(rmdup) (sequencing depth).
The Phylogenetic_tree_Vsylvestris.zip subfolder includes: the Phylogenetic_tree_snphylo_Vsylvestris.raxml.support file, which is used to visualize the obtained phylogeny.
The Admixture_Vsylvestris.zip subfolder includes: the Admixture_Vsylvestris file. This file includes the following variables: ID (ID of each sequence), K1, K2, K3 (the values in each of these variables correspond to the genetic ancestry of each sequence to each of these ancestral populations (K)). The files K_Vsylvestris.3.Q, and K_Vsylvestris.fam are the result of the ADMIXTURE analysis. Finally, the file Admixture_Vsylvestrys_code.txt includes the R code for graphing the analysis.
The Treemix_Vsylvestris.zip subfolder includes: the treemix70_runs_Vsylvestris and treemix75_runs_Vsylvestris folders, for each of the scenarios evaluated (ancestry greater than 70 % or 75 %). Each folder contains the .cov, .covse, .edges, .llik, .modelcov, .treeout, and .vertices files generated for each of the migration events evaluated. Similarly, the OpTM_Vsylvestris_code.txt file includes the R code for determining the optimal number of migrations. Finally, the Treemix_Vsylvestris_code.txt file is included with the R code for graphing the phylogenies obtained from Treemix.
The Heterozygosity_Vsylvestris.zip subfolder includes: the rarity_Vsylvestris.profile file. This file includes the variables FID and IID (ID of each sequence), PHENO (phenotype), CNT (non-missing alleles used for calculation), CNT2 (effective alleles for calculating rarity), and SCORESUM (sum of alleles per individual). Similarly, the file Heterozygosity_Vsylvestris.het.txt is included with the variables INDV (ID of each sequence), O (HOM) (observed homozygotes), E(HOM) (expected homozygotes), N_SITES (number of variants (SNPs) used per individual) and F (inbreeding coefficient).
