Data from: Replicate geographic transects across a hybrid zone reveal parallelism and differences in the genetic architecture of reproductive isolation
Data files
Sep 30, 2025 version files 9.22 GB
-
DRYAD_Submission.zip
9.22 GB
-
README.md
4.59 KB
Abstract
Determining the genetic architecture of traits involved in adaptation and speciation is one of the key components of understanding the evolutionary mechanisms behind biological diversification. Hybrid zones provide a unique opportunity to use genetic admixture to identify traits and loci contributing to partial reproductive barriers between taxa. Many studies have focused on temporal dynamics of hybrid zones, but geographical variation in hybrid zones that span distinct ecological contexts has received less attention. We address this knowledge gap by analyzing hybridization and introgression between black-capped and Carolina chickadees in two geographically remote transects across their extensive hybrid zone, one located in eastern and one in central North America. Previous studies demonstrated that this hybrid zone is moving northward as a result of climate change, but is staying consistently narrow due to selection against hybrids. In addition, the hybrid zone is moving ~5x slower in central North America compared to more eastern regions, reflecting continent-wide variation in the rate of climate change. We use whole genome sequencing of 259 individuals to assess whether variation in the rate of hybrid zone movement is reflected in patterns of hybridization and introgression, and which genes and genomic regions show consistently restricted introgression in distinct ecological contexts. Our results highlight substantial similarities between geographically remote transects and reveal large Z-linked chromosomal rearrangements that generate measurable differences in the degree of gene flow between transects. We further use simulations and analyses of climatic data to examine potential factors contributing to continental-scale nuances in selection pressures. We discuss our findings in the context of speciation mechanisms and the importance of sex chromosome inversions in chickadees and other species.
Dataset DOI: 10.5061/dryad.fn2z34v68
Description of the data and file structure
ASAPH
"ASAPH" directory contains R script and input files for analysis of chromosomal rearrangements. ASAPH plotting.R script specifies input files for each analysis. All input files for these analyses are located in the ASAPH directory. There are two types of files with .eigneval and .eigenvec extensions used for each analysis of chromosomal rearrangements (please refer to ASAPH manual for further details). "allo_bcch" = comparison between allopatric population of black-capped chickadees. "allo_cach" = comparison between allopatric populations of Carolina chickadees. Further, four comparisons are made for PCA ASAPH analysis: "bcch_pca" = comparison between allopatric black-capped chickadees, "cach_pca" = comparison between allopatric Carolina chickadees, "mo_pca" = comparison for the Missouri transect. "pa_pca" = comparison for the Pennsylvania transect.
BGC analysis
"BGC analysis" directory contains input files and R scripts for Bayesian Genomic Cline analysis. MO.big.code.R and PA.big.code.R specify steps for the analysis of Missouri and Pennsylvania transects respectively. MO.big.Rdata and PA.big.Rdata contain input files for Missouri and Pennsylvania transects respectively.
MO.big.Rdata file structure:
- gc1 -- the gghybrid output object containing results of Bayesian Genomic Cline analysis for Missouri transect's test run
- gc.MO.long -- the gghybrid output object containing results of Bayesian Genomic Cline analysis for Missouri transect's final run
- dat -- gghybrid object containing information on input SNP table for hybrid index estimates
- prepdata -- an intermediate gghybrid object used to prepare input data for analysis (see gghybrid manual)
- hindlabel -- gghybrid object containing posterior hybrid index estimates used for plotting
PA.big.Rdata file structure:
- gc1 -- the gghybrid output object containing results of Bayesian Genomic Cline analysis for Pennsylvania transect's test run
- gc.PA.long -- the gghybrid output object containing results of Bayesian Genomic Cline analysis for Missouri transect's final run
- dat -- gghybrid object containing information on input SNP table for hybrid index estimates
- prepdata -- an intermediate gghybrid object used to prepare input data for analysis (see gghybrid manual)
- hind label -- gghybrid object containing posterior hybrid index estimates used for plotting
Climate
"Climate" directory contains R script Climate analysis.R used for analysis of climatic differences between Missouri and Pennsylvania transects.
Hybrid index
"Hybrid index" directory contains R script and input files for analysis of heterozygosity and hybrid index. MO.new.ref.hybrid.index.heteroz.Rand PA.new.ref.hybrid.index.heteroz.R outline analysis steps for Missouri and Pennsylvania transects and specify input files. All input files for these analyses are located in the Hybrid index directory. All files that have MO and PA in their naming refer to Missouri and Pennsylvania transects respectively.
- meta.hiho -- data tables containing metadata (Order in VCF file, ID, sampling region)
- hybrid.index.and.hiho.*.W.mtdna - output files containing information on hybrid index and heterozygosity for each transect and information on MtDNA haplotype (0=Black-capped chickadee , 1=Carolina chickadee)
- for.hybr.index.012* -- input file in 012 VCFtools format with SNPs.
PCA
"PCA" directory contains R script and input data for the Principal Component Analysis of genomic data. pca.both.transect.new.reference.R specifies input files used for Missouri and Pennsylvania transects.
- pca.meta.new.ref.txt -- metadata (Order in VCF file, ID, sampling region)
- final.filtered.combined.new.reference..thin10000.012 -- input file in 012 VCFtools format with SNPs filtered to retain sites fixed between allopatric Black-capped and Carolina chickadees and thinned to retain only 1 SNP per 10,000bp window.
VCF table
"VCF table" directory contains filtered panel of SNPs final.filtered.combined.new.reference.recode.vcf used in all downstream analyses:
final.filtered.combined.new.reference.recode.vcf -- Filtered dataset in VCFTools format containing SNP table used in all downstream analyses
DRYAD_Submission.zip contains all directories in the compressed format
