Data for: Ancient sedimentary DNA shows more than 5000 years of continuous beaver occupancy in Grand Teton National Park
Data files
Sep 03, 2025 version files 56.85 MB
-
README.md
4.19 KB
-
SI_tableS4_trnl_rawdata.xlsx
10.61 MB
-
SI_tableS5_P007_rawdata.xlsx
46.23 MB
Abstract
Beaver-based restoration is emerging as a cost-effective conservation and climate adaptation strategy, but efforts are constrained by limited knowledge of pre-colonial beaver distribution and their long-term ecosystem impacts. Here, we apply sedimentary ancient DNA (sedaDNA) techniques to investigate the history of beaver occupancy at three lakes in Grand Teton National Park, Wyoming, over the last ~10 ka, as well as interactions with the local plant community. To investigate change in the vascular plant community, extracts were PCR amplified using barcode primers targeting the trnL P6 loop of the plant chloroplast genome with five replicates per DNA extract. A barcode targeting the 16S region of the mammalian mitochondrial genome (16SmammP007) was also amplified and sequenced from a subset of samples to validate beaver presence results from the species-species qPCR assay. trnL metabarcoding showed differing plant communities in the two lower-elevation lakes as compared to the higher elevation lakes, with mid-Holocene shifts in the plant community coinciding with both local beaver establishment (as indicated by species-specific qPCR results) and regional climatic changes trending towards wetter conditions. 16S metabarcoding yielded sporadic detections of a limited number of mammalian taxa, but confirmed beaver detection from the species-species qPCR assay.
Dataset DOI: 10.5061/dryad.cc2fqz6j9
Description of the data and file structure
Dataset includes two Microsoft Excel files (SI_tableS4_trnl_rawdata.xlsx and SI_tableS5_P007_rawdata.xlsx) summarizing raw metabarcode sequencing data
generated from sedimentary ancient DNA associated with the Ecosphere publication:
Ancient sedimentary DNA shows more than 5000 years of continuous beaver occupancy in Grand Teton National Park. 2025. D. Nevé Baker, Darren J. Larsen, Emily Fairfax, Amelia P. Muscott, Beth Shapiro, Sarah E. Crump.
Each data file contains sequencing reads that have been trimmed and processed with the Anacapa QC pipeline (https://github.com/limey-bean/Anacapa), then clustered as Amplicon Sequence Variants (ASVs). Taxonomy was then assigned by aligning ASV clusters to the NCBI nr/nt databases with the Anacapa CRUX pipeline (https://github.com/limey-bean/CRUX_Creating-Reference-libraries-Using-eXisting-tools). ASV assignments with a 60% Bayesian Confidence Cutoff were retained.
SI_tableS4_trnl_rawdata.xlsx contains vascular plant sequences generated using barcode primers targeting the trnL P6 loop of the plant chloroplast genome.
SI_tableS5_P007_rawdata.xlsx contains sequences generated using barcode primers targeting the 16S region of the mammalian mitochondrial genome (16SmammP007). All reads assigned as human have been removed.
Each file contains a metadata sheet ("metadata") and multiple data matrix sheets ("{primer}_ASV_taxonomy_detailed_X"), linked to the metadata by the "sum.taxonomy" field. A more complete description of the data structure can be found at https://github.com/limey-bean/Anacapa.
Data files can be visualized and analyzed in R using the Ranacapa package or with the Ranacapa online shiny application (https://shiny.eeb.ucla.edu/ranacapa/)
Metadata fields:
sum.taxonomy: sample replicate name
sum.tax.order: sample order
extraction_ID: DNA extraction ID
type: sample or negative
lake: lake ID
core_section: sediment core section ID
core_depth_mid: depth of sample (cm) on core section
total_depth: depth of sample (cm) for complete core
batch: metabarcode lab processing batch
"na" fields indicate non-applicable data fields (e.g., for negatives).
data matrix description:
{primer}_seq_number: unique sequence identifier
sequence: raw sequence after trimming and merging
sequenceF: forward sequence
sequenceR: reverse sequence
forward_{primer}_seq_number: number of sequences assigned to sequenceF
reverse_{primer}_seq_number: number of sequences assigned to sequenceR
merged_{primer}_seq_number: number of sequences assigned to merged sequence
variable number of columns beginning with either trnl_ (SI_tableS4_trnl_rawdata.xlsx) or X16SmammP007_ (SI_tableS5_P007_rawdata.xlsx), indicating sample replicate name (linked to metadata via sum.taxonomy field). Values reflect the number of assigned sequences
unmerged_{primer}_seq_number: number of sequences assigned to both sequenceF and sequenceR (unmerged)
single_or_multiple_hit: single or multiple taxonomic assignment with BLAST
end_to_end_or_local: BLAST assignment mode
max_percent_id: BLAST maximum taxonomic assignment likelihood
input_sequence_length: sequence length in base pairs
taxonomy: BLAST taxonomic assignment at each level
taxonomy_confidence: BLAST confidence at each taxonomic level
accessions: NCBI accession numbers for taxonomic assignments
Empty cells indicate non-applicable fields. Empty cells will be correctly ignored by the R package Ranacapa during downstream processing.
Code/software
Microsoft Excel, LibreOffice.
Data files can be visualized and analyzed in R using the Ranacapa package or with the Ranacapa online shiny application (https://shiny.eeb.ucla.edu/ranacapa/)
