Ancestry and adaptation of resident, adfluvial, and anadromous rainbow trout (Oncorhynchus mykiss) in Putah Creek, an ecotone crossing watershed
Data files
Jul 23, 2026 version files 247.35 MB
-
cdfw_meta_2share.csv
84.14 MB
-
fw_M1329_corrected.csv
12.67 KB
-
fw_M1371_corrected.csv
11.62 KB
-
gtseq1-6_13-16_observed_unfiltered_haplotypeshare.csv
73.56 MB
-
gtseq20_observed_unfiltered_haplotypeshare.csv
6.72 MB
-
lab_pop_meta.csv
750.41 KB
-
miseq_ids.csv
97.47 KB
-
observed_unfiltered_haplotype_CDFW-4-30-2020share.csv
82.04 MB
-
README.md
10.80 KB
-
repository_M1371.csv
8.23 KB
-
respository_M1329.csv
8.92 KB
Abstract
Evaluating patterns in the distribution of genetic variation across heterogeneous landscapes provides perspectives on how environmental features contribute to molecular evolution, and how human activities influence populations. Impassable barriers, such as hydroelectric dams, and human-directed movement or propagation of individuals across riverscapes have had strong impacts on the genetic diversity observed in natural populations. We considered the population structure and distribution of adaptive genetic variation in the salmonid species Oncorhynchus mykiss throughout Putah Creek, a watershed near Napa Valley, California that was transformed by the 1950’s construction of multiple dams and a large artificial reservoir (Lake Berryessa). In addition, hatchery-raised O. mykiss have been introduced throughout the watershed over many years, which may have introduced genetic variation into existing Putah Creek populations. To explore the population structure and distribution of adaptive genetic variation in Putah Creek, we analyzed microhaplotypes from 196 individuals from eight O. mykiss populations above Lake Berryessa and two populations below Lake Berryessa, and compared them with 2,075 samples from 40 reference populations from California Central Valley, coastal, inland, and hatchery rainbow trout lineages. We found distinct patterns of neutral and adaptive variation between populations above Lake Berryessa and those below. Populations below Lake Berryessa resembled various Central Valley populations and hatchery rainbow trout strains, while those above were more similar to coastal O. mykiss populations. Additionally, Putah Creek populations below Lake Berryessa possessed significantly different proportions of adaptive variants associated with life-history compared to populations above Lake Berryessa, consistent with studies in other populations located above and below barriers to migration.
Dataset DOI: 10.5061/dryad.1c59zw49w
Description of the data and file structure
196 rainbow trout were sampled from ten Putah Creek locations for comparison with 2,075 rainbow trout from 40 reference populations. Ten Californian regions were considered for this study, with at least two O. mykiss populations represented per region (Table 1; Table S1; Figure 1), including coastal and Central Valley lineages, anadromous steelhead, rainbow and redband trout populations, and five strains of hatchery rainbow trout commonly used for stocking in California (Coleman, Shasta, Pit, Kamloops, and Hot Creek). All genetic data collected for this study were obtained by sequencing a panel of ‘microhaplotype’ loci (Le Gall et al. 2024), each containing one or more single nucleotide polymorphisms (SNPs). We followed sequencing protocols outlined by Le Gall et al. (2024). To characterize genomic variation associated with migratory life-histories, the Le Gall et al (2024) genotyping panel includes multiple microhaplotypes in both the GREB1L/ROCK1 genomic region and across the chromosomal inversion on Omy05. To characterize frequencies of the inversion haplotypes on Omy05, we targeted a single SNP within the Omy05 chromosomal inversion (omy5_9_54854574-19) and estimated the frequencies of homozygous and heterozygous inversion haplotypes based on the anadromous (AA), resident (RR), and heterozygous (AR) genotypes at that locus. Similarly, to characterize the distribution of variation in the GREB1L/ROCK1 genetic region, we focused on a single SNP (mhap8_71, pos. 11667915) that has been used in many recent studies of this gene region (e.g. Collins et al. 2020; reviewed by Waples et al. 2022). Genotypes at this SNP are associated with the early- (EE) and late- (LL) migratory life-histories known as summer- and winter-run steelhead, respectively, and individuals can also be heterozygous (EL).
Files and variables
File: fw_M1329_corrected.csv
Description: Sampling file for Putah Creek samples, pt1
Variables
- NMFS_DNA_ID: Lab-assigned sample ID
- STATE_F: US state of sampling location
- COUNTY_F: US county of sampling location
- WATERSHED: Name of watershed
- TRIB_1: Nearest tributary #1
- TRIB_2: Nearest tributary #2
- WATER_NAME: General sampling location name
- REACH_SITE: Specific sampling location name
- HATCHERY: Name of hatchery sampled (if applicable)
- STRAIN: Name of strain sampled (if applicable)
- LATITUDE_F: Latitude of sampling location
- LONGITUDE_F: Longitude of sampling location
- LOCATION_COMMENTS_F: General comments for sampling location
File: fw_M1371_corrected.csv
Description: Sampling file for Putah Creek samples pt2
Variables
- NMFS_DNA_ID: Lab-assigned sample ID
- STATE_F: US state of sampling location
- COUNTY_F: US county of sampling location
- WATERSHED: Name of watershed
- TRIB_1: Nearest tributary #1
- TRIB_2: Nearest tributary #2
- WATER_NAME: General sampling location name
- REACH_SITE: Specific sampling location name
- HATCHERY: Name of hatchery sampled (if applicable)
- STRAIN: Name of strain sampled (if applicable)
- LATITUDE_F: Latitude of sampling location
- LONGITUDE_F: Longitude of sampling location
- LOCATION_COMMENTS_F: General comments for sampling location
File: lab_pop_meta.csv
Description: Metadata file for reference population samples downloaded from lab wiki. Cells are either blank or contain 'NA' for null values.
Variables
- NMFS_DNA_ID: Lab-assigned sample ID
- BOX_ID: Box id where samples are stored
- BOX_POSITION: DNA sample well in box (e.g. 1A, 1B, 1C, etc.)
- SAMPLE_ID: Original sample ID provided with samples when sent to lab
- BATCH_ID: Batch id that includes all samples sequenced
- PROJECT_NAME: Name of project for sampling
- GENUS: Species genus
- SPECIES: Species name
- LENGTH: Fork length of individual fish
- WEIGHT: Weight of individual fish
- SEX: Recorded phenotypic sex of individual fish
- AGE: Estimated age of individual fish (different than calculated age at spawning)
- REPORTED_LIFE_STAGE: If adult or juvenile
- PHENOTYPE: Any phenotypic data
- HATCHERY_MARK: If adipose fin is intact or not - YES means hatchery staff removed adipose fin before release from hatchery program
- TAG_NUMBER: Each fish is checked for external or internal tags during sampling - if a fish was tagged the tag number is in this column.
- COLLECTION_DATE: Date fish was sampled
- ESTIMATED_DATE: Estimated sample date if collection_date is missing
- PICKER: Which lab member cut a piece of fin clip tissue and placed in labeled tube for sequencing
- PICK_DATE: When sample was prepared for sequencing
- LEFTOVER_SAMPLE: If all of tissue sample was used or not
- SAMPLE_COMMENTS: Any notes about tissue sample from Picker
- NMFS_DNA_ID.1: Duplicate column with lab-assigned ID
- STATE_F: US state of sampling location
- COUNTY_F: US county of sampling location
- WATERSHED: Name of watershed
- TRIB_1: Nearest tributary #1
- TRIB_2: Nearest tributary #2
- WATER_NAME: General sampling location name
- REACH_SITE: Specific sampling location name
- HATCHERY: Name of hatchery sampled (if applicable)
- STRAIN: Name of strain sampled (if applicable)
- LATITUDE_F: Latitude of sampling location
- LONGITUDE_F: Longitude of sampling location
- LOCATION_COMMENTS_F: General comments for sampling location
File: miseq_ids.csv
Description: Two column file with Lab-assigned ID and sequencing ID. Allows us to connect metadata to microhaplotype sequences.
Variables
- NMFS_DNA_ID: Lab-assigned sample ID
- MISEQ_ID: Microhaplotype sequencing-assigned sample ID
File: repository_M1371.csv
Description: Metadata file for Putah Creek samples downloaded from lab wiki, pt1
Variables
- NMFS_DNA_ID: Lab-assigned sample ID
- BOX_ID: Box id where samples are stored
- BOX_POSITION: DNA sample well in box (e.g. 1A, 1B, 1C, etc.)
- SAMPLE_ID: Original sample ID provided with samples when sent to lab
- BATCH_ID: Batch ID that includes all samples sequenced
- PROJECT_NAME: Name of project for sampling
- GENUS: Species genus
- SPECIES: Species name
- LENGTH: Fork length of individual fish (mm)
- WEIGHT: Weight of individual fish (g)
- SEX: Recorded phenotypic sex of individual fish
- AGE: Estimated age of individual fish (different than calculated age at spawning)
- REPORTED_LIFE_STAGE: If adult or juvenile
- PHENOTYPE: Any phenotypic data
- HATCHERY_MARK: If adipose fin is intact or not - YES means hatchery staff removed adipose fin before release from hatchery program
- TAG_NUMBER: Each fish is checked for external or internal tags during sampling - if fish was tagged the tag number is in this column
- COLLECTION_DATE: Date fish was sampled
- ESTIMATED_DATE: Estimated sample date if collection_date is missing
- PICKER: Which lab member cut a piece of fin clip tissue adn placed in labeled tube for sequencing
- PICK_DATE: When sample was prepared for sequencing
- LEFTOVER_SAMPLE: If all of tissue sample was used or not
- SAMPLE_COMMENTS: Any notes about tissue sample from Picker
File: respository_M1329.csv
Description: Metadata file for Putah Creek samples downloaded from lab wiki, pt2
Variables
- NMFS_DNA_ID: Lab-assigned sample ID
- BOX_ID: Box id where samples are stored
- BOX_POSITION: DNA sample well in box (e.g. 1A, 1B, 1C, etc.)
- SAMPLE_ID: Original sample ID provided with samples when sent to lab
- BATCH_ID: Batch ID that includes all samples sequenced
- PROJECT_NAME: Name of project for sampling
- GENUS: Species genus
- SPECIES: Species name
- LENGTH: Fork length of individual fish (mm)
- WEIGHT: Weight of individual fish (g)
- SEX: Recorded phenotypic sex of individual fish
- AGE: Estimated age of individual fish (different than calculated age at spawning)
- REPORTED_LIFE_STAGE: If adult or juvenile
- PHENOTYPE: Any phenotypic data
- HATCHERY_MARK: If adipose fin is intact or not - YES means hatchery staff removed adipose fin before release from hatchery program
- TAG_NUMBER: Each fish is checked for external or internal tags during sampling - if fish was tagged the tag number is in this column
- COLLECTION_DATE: Date fish was sampled
- ESTIMATED_DATE: Estimated sample date if collection_date is missing
- PICKER: Which lab member cut a piece of fin clip tissue adn placed in labeled tube for sequencing
- PICK_DATE: When sample was prepared for sequencing
- LEFTOVER_SAMPLE: If all of tissue sample was used or not
- SAMPLE_COMMENTS: Any notes about tissue sample from Picker
File: cdfw_meta_2share.csv
Description: Metadata and microhaplotypes for samples received from CDFW
Variables
- row_num: Row number
- group: Population name
- indiv.ID: Individual sample ID
- locus: Name of locus sequenced
- haplo: Haplotype
- depth: Read depth
- allele.balance: Haplotype read depth
- rank: Number of sequence
File: gtseq20_observed_unfiltered_haplotypeshare.csv
Description: Microhaplotypes for Putah and lab-provided reference samples, pt1
Variables
- row_num: Row number
- group: Population name
- indiv.ID: Individual sample ID
- locus: Name of locus sequenced
- haplo: Haplotype
- depth: Read depth
- allele.balance: Haplotype read depth
- rank: Number of sequence
File: gtseq1-6_13-16_observed_unfiltered_haplotypeshare.csv
Description: Microhaplotypes for Putah and lab-provided reference samples, pt2
Variables
- row_num: Row number
- group: Population name
- indiv.ID: Individual sample ID
- locus: Name of locus sequenced
- haplo: Haplotype
- depth: Read depth
- allele.balance: Haplotype read depth
- rank: Number of sequence
File: observed_unfiltered_haplotype_CDFW-4-30-2020share.csv
Description: Microhaplotypes for CDFW-provided samples
Variables
- row_num: Row number
- group: Population name
- indiv.ID: Individual sample ID
- locus: Name of locus sequenced
- haplo: Haplotype
- depth: Read depth
- allele.balance: Haplotype read depth
- rank: Number of sequence
Code/software
Scripts are available at https://github.com/lauracgoetz/OmykissPutah
Scripts were written for use with R statistical programming.
Data can be viewed using Excel or any other software that can handle .csv files.
Access information
Other publicly accessible locations of the data:
Data were derived from the following sources:
- See Le Gall et al. 2024 Conservation Genetics Resources
