Data from: Genome-wide SNP data reveal population boundaries, gene flow, and diversification patterns in the Western Banded Gecko (Coleonyx variegatus)
Data files
Jul 21, 2026 version files 284.96 MB
-
azm_ns.txt
14.35 MB
-
bcn_bcs.txt
12.90 MB
-
bf_bidirectional.py
2.39 KB
-
Col.vari_enm-genetic.csv
959 B
-
coleonyx_all.phy
16.91 MB
-
coleonyx_all.vcf
17.48 MB
-
coleonyx_assembly.str
7.20 MB
-
coleonyx_ibd.R
8.44 KB
-
coleonyx_ibe.R
2.27 KB
-
coleonyx_only.vcf
25.29 MB
-
Davis_et_al_-_Summary_of_variable_contribution.csv
5.59 KB
-
migration_plot.py
2.35 KB
-
mo_azm.txt
19.67 MB
-
mo_dv.txt
14.09 MB
-
mo_st.txt
14.47 MB
-
ns_ss.txt
9.36 MB
-
README.md
8.05 KB
-
sample_coordinates.csv
8.73 KB
-
ss_bcn.txt
6.44 MB
-
st_azm.txt
14.77 MB
-
st_bcn.txt
11.56 MB
-
st_ns.txt
9.66 MB
-
st_ss.txt
8.66 MB
-
variegatus_A00.txt
39.98 MB
-
variegatus_BPP.txt
42.14 MB
Abstract
This dataset contains ddRADseq and environmental data associated with the Western Banded Gecko (Coleonyx variegatus), a species distributed across the Sonoran and Mojave deserts and the Baja California Peninsula. The dataset includes genome-wide RADseq loci for input into the program BPP, a concatenated phylip file for phylogenetic reconstruction, genome-wide SNP data in a VCF file for population structure analyses, and ecological niche model outputs representing environmental suitability across the species’ range. Associated metadata provide detailed information on sampling localities and specimen identifiers.
These data were used to investigate phylogeographic structure, patterns of gene flow, and population differentiation among C. variegatus populations spanning distinct desert ecoregions. The dataset enables reanalysis using alternative population structure methods, integration into multispecies comparative phylogeographic studies, or refinement of ecological niche models across southwestern North America. Geographic coordinates are coarsened to protect sensitive localities, and all specimens were collected under appropriate state and federal permits.
Dataset DOI: 10.5061/dryad.41ns1rnvn
Description of the data and file structure
We demultiplexed and assembled the data using ipyrad. We applied a sequence similarity threshold of 50% to cluster reads within samples and loci between samples. We removed consensus sequences with low coverage (< 6 reads), excessive undetermined or heterozygous sites (> 5%), too many alleles for a sample (> 2 for diploids), an excess of shared heterozygosity among samples (paralog filter = 0.5), and any samples with < 500,000 raw reads. We generated one assembly with all individuals, including outgroups requiring a minimum 50% of individuals to share each locus. This assembly contains 179 individuals and 1,177 loci (coleonyx_all.phy & coleonyx_all.vcf). From this assembly, we used the branching function in ipyrad to remove outgroup samples (variegatus_only.vcf). Finally, we created additional branches for conducting species tree inference and gene flow estimation (azm_ns.txt, mo_st.txt, st_bcn.txt, variegatus_BPP.txt, bcn_bcs.txt, ns_ss.txt, st_ns.txt, mo_azm.txt, ss_bcn.txt, st_ss.txt, mo_dv.txt, st_azm.txt, variegatus_A00.txt).
Population acronyms used throughout files:
- DV = Death Valley
- MO = Mojave Desert
- AZM = Arizona/New Mexico Mountains
- ST = Salton Trough
- NS = northern Sonora Desert
- SS = southern Sonora Desert
- BCN = Baja California Norte
- BCS = Baja California Sur
Files and variables
File: sample_coordinates.csv
Description: Sample localities and identifiers for all samples used in the study.
Variables
- no: Arbitrary number from 1–N samples
- sampleID: Formatted as State, Location, Museum/Field Voucher
- Pop: Acronym for the population name
- lon: Longitude
- lat: Latitude
File: coleonyx_all.vcf
Description: VCF file containing all samples used in the study, ingroup (Coleonyx variegatus) and outgroup (C. brevis).
File: coleonyx_all.phy
Description: PHYLIP file for the full dataset used to reconstruct concatenated tree.
File: azm_ns.txt
Description: Population comparison for the MSCM model in BPP between AZM and NS.
Variables
- Number of samples, sequence length, sample name, sequence data
File: coleonyx_only.vcf
Description: VCF file containing only samples of the ingroup (Coleonyx variegatus).
File: bcn_bcs.txt
Description: Population comparison for the MSCM model in BPP between BCN and BCS.
Variables
- Number of samples, sequence length, sample name, sequence data
File: mo_azm.txt
Description: Population comparison for the MSCM model in BPP between MO and AZM.
Variables
- Number of samples, sequence length, sample name, sequence data
File: mo_dv.txt
Description: Population comparison for the MSCM model in BPP between MO and DV.
Variables
- Number of samples, sequence length, sample name, sequence data
File: ns_ss.txt
Description: Population comparison for the MSCM model in BPP between NS and SS.
Variables
- Number of samples, sequence length, sample name, sequence data
File: mo_st.txt
Description: Population comparison for the MSCM model in BPP between MO and ST.
Variables
- Number of samples, sequence length, sample name, sequence data
File: ss_bcn.txt
Description: Population comparison for the MSCM model in BPP between SS and BCN.
Variables
- Number of samples, sequence length, sample name, sequence data
File: st_ns.txt
Description: Population comparison for the MSCM model in BPP between ST and NS.
Variables
- Number of samples, sequence length, sample name, sequence data
File: st_azm.txt
Description: Population comparison for the MSCM model in BPP between ST and AZM.
Variables
- Number of samples, sequence length, sample name, sequence data
File: st_bcn.txt
Description: Population comparison for the MSCM model in BPP between BCN and ST.
Variables
- Number of samples, sequence length, sample name, sequence data
File: st_ss.txt
Description: Population comparison for the MSCM model in BPP between ST and SS.
Variables
- Number of samples, sequence length, sample name, sequence data
File: variegatus_A00.txt
Description: Dataset for estimating population genetic parameters
Variables
- Number of samples, sequence length, sample name, sequence data
File: variegatus_BPP.txt
Description: Dataset to infer the species tree.
Variables
- Number of samples, sequence length, sample name, sequence data
File: Col.vari_enm-genetic.csv
Description: A dataset containing pairwise population comparisons including identity indices, genetic differentiation (FST), and PGD values.
Variables
- pops: Two populations being compared.
- identity: A similarity index used to quantify the degree of niche overlap between the ecological models of the two populations.
- fst: FST value between the two populations being compared.
- pgd: Pairwise genetic distance between the two populations being compared.
File: Davis_et_al_-_Summary_of_variable_contribution.csv
Description:
MaxEnt ecological niche model evaluation metrics and bioclimatic variable contributions for each population.
Variables
- Population: MaxEnt evaluation metric name or WorldClim bioclimatic variable name.
- AZM: Results for the specific population.
- BCN: Results for the specific population.
- BCS: Results for the specific population.
- DV: Results for the specific population.
- MO: Results for the specific population.
- NS: Results for the specific population.
- SS: Results for the specific population.
- ST: Results for the specific population.
Key Rows Included
-
Training samples: Number of occurrence points used to train each model.
-
Training AUC: Area Under the Curve metric indicating model accuracy.
-
bioXX contribution: Percentage contribution of bioclimatic variable XX during model construction.
-
bioXX permutation importance: Percent decrease in AUC when variable XX values are randomly permuted.
-
Thresholds (Cloglog): Standard decision thresholds (e.g., 10th percentile, Max Sensitivity + Specificity) used to convert predicted probabilities into binary presence/absence maps.
File: coleonyx_ibd.R
Description: R script for estimating isolation-by-distance for the species and for each population.
File: coleonyx_assembly.str
Description: Structure file containing only samples of the ingroup (Coleonyx variegatus). File used as the input for the coleonyx_ibd.R and coleonyx_ibe.R scripts.
File: coleonyx_ibe.R
Description: Extracts WorldClim data and performs Mantel tests to analyze Isolation-by-Environment.
File: migration_plot.py
Description: Generates error bar plots to visualize migration rates (m) and confidence intervals between population pairs.
File: bf_bidirectional.py
Description: Calculates Bayes Factors using the Savage-Dickey density ratio to test for significant migration from BPP output.
Code/software
Format: File, Language, Description
bf_bidirectional.py, Python, Calculates Bayes Factors using the Savage-Dickey density ratio to test for significant migration from BPP output.
migration_plot.py, Python, Generates error bar plots to visualize migration rates ($m$) and confidence intervals between population pairs.
coleonyx_ibd.R, R, Performs Mantel tests to analyze Isolation by Distance using genetic (Edward's distance) and geographic distances.
coleonyx_ibe.R, R, Extracts WorldClim data and performs Mantel tests to analyze Isolation by Environment.
Access information
Other publicly accessible locations of the data:
- N/A
Data was derived from the following sources:
- N/A
We extracted genomic DNA from tissue samples using a salt extraction protocol, and conducted double digest restriction-site associated DNA sequencing (ddRADseq). We digested each sample using the digestion enzymes SbfI and MspI in CutSmart Buffer (New England Biolabs) for 7 hours at 37 ºC. For fragment purification, we used Sera-Mag SpeedBeads. We then ligated eight distinct barcodes with unique molecular identifiers (UMI) to the cut sites of the fragmented DNA. After barcode ligation, we size-selected each library between 415 and 515 base pairs (bp) on a Blue Pippin Prep size fractionator (Sage Science). For final library amplification, we used Phusion Hi-Fidelity DNA Polymerase and Illumina index primers. We determined the concentration and size distribution of each pool using an Agilent 2200 TapeStation. Lastly, we sent the quantified libraries to QB3-Berkeley Genomics, UC Berkeley for qPCR and sequencing on two Illumina NovaSeq 6000 lanes (100-bp, single-end reads; 28 pools containing 8 samples). The demultiplexed data are deposited in the NCBI Sequence Read Archive (SRA; BioProject ID: PRJNA 1235118).
