Data from: fine-scale population structure and relatedness of Argali (Ovis ammon) in Kyrgyzstan revealed by high-density SNP data
Data files
Apr 26, 2026 version files 609.70 MB
-
Argali_DNA.xlsx
17.91 KB
-
data100.dsf
37.93 MB
-
data11.dsf
31.07 MB
-
data11.dsf.cache
66 B
-
data14.dsf
45.65 MB
-
data14.dsf.cache
174 B
-
data17.dsf
27.29 MB
-
data17.dsf.cache
18.27 MB
-
data2.dsf
27.29 MB
-
data2.dsf.cache
18.27 MB
-
data23.dsf
31.48 MB
-
data23.dsf.cache
228 B
-
data27.dsf
31.48 MB
-
data27.dsf.cache
226 B
-
data30.dsf
31.48 MB
-
data30.dsf.cache
226 B
-
data34.dsf
15.52 MB
-
data34.dsf.cache
382 B
-
data38.dsf
32.77 KB
-
data38.dsf.cache
257 B
-
data41.dsf
32.77 KB
-
data41.dsf.cache
30 B
-
data46.dsf
32.77 KB
-
data46.dsf.cache
174 B
-
data49.dsf
32.77 KB
-
data49.dsf.cache
174 B
-
data5.dsf
27.29 MB
-
data5.dsf.cache
18.27 MB
-
data52.dsf
32.77 KB
-
data52.dsf.cache
174 B
-
data55.dsf
32.77 KB
-
data55.dsf.cache
174 B
-
data58.dsf
32.77 KB
-
data58.dsf.cache
174 B
-
data67.dsf
32.77 KB
-
data67.dsf.cache
614 B
-
data70.dsf
57.34 KB
-
data70.dsf.cache
8.88 KB
-
data74.dsf
46.26 MB
-
data74.dsf.cache
382 B
-
data78.dsf
32.77 KB
-
data78.dsf.cache
174 B
-
data8.dsf
27.29 MB
-
data81.dsf
32.77 KB
-
data81.dsf.cache
228 B
-
data87.dsf
44.06 MB
-
data87.dsf.cache
382 B
-
data91.dsf
32.77 KB
-
data91.dsf.cache
110 B
-
data94.dsf
37.93 MB
-
data94.dsf.cache
18.27 MB
-
data97.dsf
37.93 MB
-
map20.dsm
31.35 MB
-
map20.dsm.cache
4.85 MB
-
README.md
13.46 KB
Abstract
This dataset contains high-density SNP genotyping data from 72 Argali sheep (Ovis ammon) sampled across Kyrgyzstan and Tajikistan. Genotyping was performed using the Illumina Ovine High Density SNP array, and after quality control, 135,242 SNP markers were retained. The dataset supports analyses of population structure, subspecies delineation, and relatedness within the Tian Shan region. Results derived from this dataset indicate no genomic distinction between O. a. polii and O. a. karelini, suggesting they form a single genetic unit. The data also reveal moderate genetic substructure and relatedness patterns consistent with dominant male breeding systems. This dataset provides a genomic baseline for conservation planning and transboundary management of Argali populations in Central Asia.
Dataset Title
Data associated with: Fine-Scale Population Structure and Relatedness of Argali (Ovis ammon) in Kyrgyzstan Revealed by High-Density SNP Data
Dataset DOI: 10.5061/dryad.76hdr7t7t
1. General Description
This dataset contains genetic and associated metadata used to examine fine-scale population structure and relatedness of Argali wild sheep (Ovis ammon) in Kyrgyzstan using high-density single nucleotide polymorphism (SNP) genotyping.
The dataset includes:
- Sample-level metadata (sample identifiers and geographic coordinates)
- Raw SNP genotype data generated using the Illumina HD Ovine SNP array
- SNP mapping information aligned to the Ovis aries reference genome
The data are intended for reuse in population genetics, conservation biology, and spatial or landscape genetic analyses. This README is written to support users who may be new to SNP array data formats or Argali biology.
2. File List and Organization
2.1 Metadata File
File: Argali_DNA.xlsx
Description
This file contains sample‑level metadata and DNA quality metrics for Argali (Ovis ammon) samples used in this study. The file includes two sample sets corresponding to two separate mailings of biological samples to the laboratory. The sample sets differ only by shipment timing and handling logistics; all samples were processed using the same laboratory protocols and quality control procedures.
Sample Sets
• Argali DNA Sample Set One: First shipment of samples received and processed.
• Argali DNA Sample Set Two: Second shipment of samples received and processed.
Sample set designation reflects mailing batches only and does not indicate biological, geographic, or analytical differences.
Variable Definitions
Mapping ID
• Description: Internal identifier used to link samples in this metadata file to corresponding genotype files (.dsf).
• Format: Numeric (integer or decimal for replicate extractions; e.g., 14.1, 14.2).
• Notes: Decimal values indicate replicate DNA extractions or repeated measurements from the same biological sample.
Sample ID
• Description: Field‑assigned sample identifier.
• Format: Numeric.
• Notes: Sample IDs may repeat across sample sets; the combination of Sample ID and Sample Set uniquely identifies each record.
Species
• Description: Species identification.
• Format: Text.
• Values: Ovis ammon (abbreviated as O. ammon).
Notes: Blank or NA values indicate species was assumed based on collection context but not explicitly recorded in the original metadata.
Sample Type
• Description: Biological material used for DNA extraction.
• Format: Text.
• Values include: skin, hair, skin + hair, skin + hide, blood + bone, meal/fur.
• Notes: “NA” indicates sample type was not recorded.
Sex
• Description: Sex of the sampled individual.
• Format: Text.
• Values: M (male), F (female), Unknown, NA.
Age
• Description: Estimated age of the individual at the time of sampling.
• Format: Numeric (years).
• Notes: “NA” indicates age was not available or could not be reliably estimated.
GPS
• Description: Geographic coordinates of sample collection location.
• Format: Text.
• Values: Coordinates recorded either in degrees/minutes/seconds (DMS) or projected coordinate format, depending on field conditions.
• Notes: Coordinate format varies among records due to multiple field teams and collection periods.
Place of Collection
• Description: Text description of the collection location.
• Format: Text.
• Notes: Includes region, area, and local geographic feature names.
Conc. (ng/µL) Test 1, 2
• Description: DNA concentration measured during quality control.
• Format: Numeric.
• Units: nanograms per microliter (ng/µL).
• Notes: When multiple measurements were taken, values represent replicate tests from the same extraction or replicate extractions.
Quality (260/280) Test 1, 2
• Description: DNA purity ratio measured as absorbance at 260 nm divided by 280 nm.
• Format: Numeric.
• Notes: Values near ~1.8 indicate high‑quality DNA; lower values may indicate protein or reagent contamination.
Date Tested
• Description: Date DNA concentration and quality were measured.
• Format: Date (DD/MM/YYYY)
Relationship to Genotype Files
Records in Argali_DNA.xlsx correspond to genotype data stored in proprietary .dsf / .dsm / .cache files. The Mapping ID column provides the link between metadata records and genotype files. Replicate Mapping IDs (e.g., 32.1 and 32.2) indicate multiple extractions or measurements from the same biological sample.
Missing Values
• NA or n/a indicates information that was not available, not recorded, or not applicable at the time of data collection or processing.
• For some records in Sample Set One, Species or Mapping ID values are missing due to incomplete metadata received with the original biological samples or clerical omissions during early data intake.
• Mapping ID entries with decimal values (e.g., 73.1) denote replicate DNA extractions or repeated measurements from the same biological sample. If a Mapping ID cell is blank, the corresponding sample did not receive a unique mapping identifier at the time of genotyping.
• Missing values do not imply exclusion from analysis unless explicitly stated in downstream analytical workflows.
2.2 Genotype Data Files
Files: data2.dsf – data100.dsf (30 files total; non-consecutive numbering)
Format: Illumina Data Storage Format (.dsf)
Description:
Each .dsf file contains raw SNP genotype data produced by the Illumina HD Ovine SNP array. These files store genotype calls and associated signal intensity information for one or more Argali samples.
Associated auxiliary files:
Each .dsf file has a corresponding .dsf.cache file that contains indexing and cache information required by Illumina software. The .cachefiles do not contain independent biological data but are necessary for efficient file access and interpretation.
2.3 SNP Map Files
File: map20.dsm
Format: Illumina SNP map (.dsm)
Description:
Provides SNP mapping information, including chromosomal positions of SNPs aligned to the Ovis aries reference genome assembly. This file is required to associate genotype calls with genomic coordinates.
Associated auxiliary file:
-
map20.dsm.cache: Cache file required by Illumina software.2.4 Genotype Data File List
The following Illumina genotype data files are included in this dataset. Each
.dsffile contains SNP genotype data generated using the Illumina HD Ovine SNP array. Each.dsffile has a corresponding.dsf.cachefile that contains auxiliary indexing information required by Illumina software.2.4.1 Genotype Data Files (
.dsf)data2.dsf data5.dsf data8.dsf data11.dsf data14.dsf data17.dsf data23.dsf data27.dsf data30.dsf data34.dsf data38.dsf data41.dsf data46.dsf data49.dsf data52.dsf data55.dsf data58.dsf data67.dsf data70.dsf data74.dsf data78.dsf data81.dsf data87.dsf data91.dsf data94.dsf data97.dsf data100.dsfDescription:
Each.dsffile stores raw SNP genotype calls and associated signal intensity data for one or more Argali (Ovis ammon) samples. Sample identifiers embedded within these files correspond to entries inArgali_DNA.xlsx.2.4.2 Auxiliary Cache Files (
.dsf.cache)data2.dsf.cache data5.dsf.cache data11.dsf.cache data14.dsf.cache data17.dsf.cache data23.dsf.cache data27.dsf.cache data30.dsf.cache data34.dsf.cache data38.dsf.cache data41.dsf.cache data46.dsf.cache data49.dsf.cache data52.dsf.cache data55.dsf.cache data58.dsf.cache data67.dsf.cache data70.dsf.cache data74.dsf.cache data78.dsf.cache data81.dsf.cache data87.dsf.cache data91.dsf.cache data94.dsf.cacheDescription:
These files are automatically generated cache and index files associated with each.dsfgenotype file. They are required for efficient data access by Illumina software but do not contain independent biological or metadata information.2.4.3 SNP Mapping Files
map20.dsm map20.dsm.cacheDescription:
Themap20.dsmfile contains SNP identifiers and genomic positions aligned to the Ovis aries reference genome. The corresponding.cachefile is required by Illumina software.4. Relationship Between Files (Clarification for Reuse)
- The list of
.dsffiles above constitutes the full set of genotype data used in this study. - Sample identifiers embedded within these
.dsffiles correspond directly to theSample_IDcolumn inArgali_DNA.xlsx. - The
.cachefiles are auxiliary and should be retained with their corresponding primary files. - The
map20.dsmfile provides the genomic coordinate reference required to interpret SNP locations.
5. Data Formats and Software Requirements
The genotype and SNP mapping files included in this dataset are provided in proprietary Illumina formats generated by the Illumina HD Ovine SNP array platform.
5.1 Genotype Data Files (
.dsfand.dsf.cache).dsffiles contain raw SNP genotype calls and associated signal intensity data..dsf.cachefiles are auxiliary cache and index files automatically generated by Illumina software and are required for efficient file access.
These files are not readable using standard spreadsheet or text-based software. They are intended to be accessed using Illumina-compatible genotyping and analysis software, including but not limited to:- Illumina GenomeStudio (Illumina, Inc.)
- Other bioinformatics tools or pipelines capable of importing Illumina SNP array data after appropriate format conversion
5.2 SNP Mapping Files (
.dsmand.dsm.cache)- The
.dsmfile contains SNP identifiers and genomic positions aligned to the Ovis aries reference genome. - The accompanying
.cachefile is required by Illumina software.
These files are used in conjunction with genotype data to assign genomic coordinates to SNPs and are accessed through the same Illumina-compatible software environments.
5.3 Metadata File (
Argali_DNA.xlsx)The metadata file is provided in Microsoft Excel format and can be opened using standard spreadsheet software or imported into statistical environments such as R or Python. This file enables users to link genotype data to sample identifiers and geographic metadata.
5.4 Notes for Data Reuse
Users unfamiliar with Illumina SNP array formats are encouraged to consult Illumina documentation or population genetics resources describing SNP array data workflows prior to reanalysis. No data files in this repository are corrupted; inability to open
.dsf,.dsm, or.cachefiles using general-purpose software reflects expected behavior for these proprietary formats.5.5 Conversion to Analysis‑Ready Formats
While the genotype data are provided in Illumina proprietary formats (
.dsfand.dsm), these files can be converted to commonly used population genetics formats (e.g., PLINK or text‑based genotype matrices) using Illumina-compatible software or downstream bioinformatics workflows.Such conversions are typically performed after importing the data into Illumina-supported environments (e.g., GenomeStudio), followed by export to analysis-ready formats supported by population genetics software. Users interested in reanalysis are encouraged to consult documentation for their preferred analysis tools to determine appropriate conversion workflows.
No converted or derived genotype files are included in this repository; the raw data are provided to maximize flexibility for reuse.
6. Reuse Notes
This dataset supports analyses including:
- Population structure and clustering
- Relatedness and kinship estimation
- Conservation and landscape genetics
Because Ovis ammon is a conservation-relevant species, results should be interpreted within appropriate ecological and management contexts.
7. Citation and Contact
Please cite this dataset using the DOI listed above, along with the associated publication.
Questions regarding data structure or interpretation should be directed to the corresponding author of the associated manuscript.
- The list of
We analyzed genomic data from 88 Argali sheep (Ovis ammon) sampled in Kyrgyzstan and Tajikistan. Samples included blood, tissue, bone, horn, and desiccated skin, collected from legally harvested individuals and natural mortalities. DNA extraction protocols varied by sample type, using Maxwell and Qiagen kits, with specialized procedures for horn and bone following Harper et al. (2013).
Genotyping was performed using the Illumina High Density Ovine SNP array, originally developed for domestic sheep (Ovis aries), yielding 606,006 SNPs. After quality control—including filtering for call rate, mapping accuracy, minor allele frequency, and Hardy–Weinberg equilibrium—72 samples and 135,242 informative markers were retained. Linkage disequilibrium pruning was applied prior to principal component analysis (PCA).
Genomic analyses included PCA to assess population structure, identity-by-descent (IBD) to estimate relatedness, and inbreeding coefficient (Fis) calculations. All analyses were conducted using Golden Helix software. Raw genotypes and metadata are included in this dataset.
