Data from: Structure of bee communities in marginal lands of the Puget Sound, USA
Data files
Aug 19, 2025 version files 1.62 MB
-
composition_analyses.Rmd
13.89 KB
-
genus_plots_v2.Rmd
23.41 KB
-
MarginalBees.csv
1.50 MB
-
README.md
3.90 KB
-
ridge_plots.Rmd
32.65 KB
-
SAC_and_Chao.Rmd
43.76 KB
Abstract
Wild bee communities in urban ecosystems are often challenged by habitat fragmentation and low floral diversity. In such settings, marginal land surrounding airports or in power line corridors may support bees, even with small habitat patches. However, temporal surveys of wild bees are lacking for many urban areas such as the Puget Sound region of western Washington State, USA. Here, we conducted wild bee surveys at three peri-urban sites in the Puget Sound over 7 years. Specifically, a standardized protocol was used to sample wild bee communities monthly from April to October at two sites associated with airports and one site in a power line corridor. In total, our surveys collected 25,441 specimens representing 118 confirmed species within 24 genera, with individual subsites having between 15 and 35 species in any year. The Halictidae represented the most individuals collected, with 47% of specimens. By genus, Lasioglossum was the most speciose (n = 21), with Bombus, Osmia, and Andrena also ubiquitous and diverse. Bee diversity was high across spring and summer, and our surveys resolved the presumptive overlap of parasites with their hosts. Our study shows that marginal lands requiring little management can support diverse wild bee communities in urban areas. Our work also provides a baseline for future evaluations of wild bee communities in the Puget Sound and broader Pacific Northwest.
Dataset DOI: 10.5061/dryad.bg79cnpnx
Description of the data and file structure
Bees were collected from marginal lands of the Puget Sound region in western WA (USA) from 2014-2020. Trap and net collection methods were used at 3 sites: Port of Seattle (POS), Boeing Paine Field (BPF), and Seattle City Light (SCL) power line corridor. All 25,000 bee specimens collected have been identified to genus, and 98.3% have been identified to the species level.
Files and variables
File: MarginalBees.csv
Description:
Variables
- Unique Specimen Number: A unique identifier for each row, which may be a single specimen, or many specimens of the same species for that collection event. Each row represents a single species and collection event.
- Label Date: The date for each collection event as a string following the format “yyyymmdd”.
- Site: Three letter abbreviation for the collection site. There are three sites, POS, BPF, and SCL. May contain NA values.
- Display Date: Same as Label Date but in the format “DD-MMM-YY”.
- Station: The subsite “station”. There are multiple stations per site. May contain NA values.
- Collection Method: The method by which the specimen(s) was/were collected, either T for “trap” or N for “net”. May contain NA values.
- Specimen Kind: How the specimen is stored, either P for “pinned” or V for “vial”.
- Genus: Genus identification.
- Subgenus: Subgenus identification. May contain NA values.
- Species: Species identification.
- Subspecies: Subspecies identification. May contain NA values.
- New record: Binary Y/N whether the specimen is a new record (state, county, etc.). See main manuscript for more details about state/county records.
- Label Name: A combination of Genus, Subgenus (if any), and Species.
- Short label name: Genus, species nomenclature.
- Sex: M for “male”, and F for “female”. May contain NA values.
- Voucher Specimen: Binary Y/N describing whether the specimen is vouchered in the curated collection. May contain NA values.
- Questionable or Incomplete Identification: Binary Y/N whether the specimen identification is incomplete. May contain NA values.
- Broken Specimen: Binary Y/N whether the specimen was damaged. May contain NA values.
- Entered by: Initials of the person who entered the specimen data.
- Entered Date: Date at which the data was entered.
- NOTES: Notes about the specimens. May contain NA values.
- Count: This column denotes the number of individuals in each record. Most rows are 1, but some rows represent hundreds of bees of the same species from the same trap.
Code/software
The data is saved in .csv (comma separated values) format and can be viewed many types of software. Code is written in R, with 4 attached R markdown files (.Rmd)—composition_analyses.Rmd,genus_plots_v2.Rmd,ridge_plots.Rmd, and SAC_and_Chao.Rmd. Each markdown file contains all data preparation necessary to generate all analyses and figures reported in the manuscript and supplemental materials. These scripts were written within an R project so the working directory is set to find the root directory. If this is an unfamiliar workflow, simply save the data to a local folder, and select that folder as the working directory when running these scripts. All code was written in R V4.2.3. You will need rmarkdown_2.27, knitr_1.47, caret_6.0-94, randomForest_4.7-1.1, vegan_2.6-6.1, lattice.0.20-45, permute.0.9-7, cowplot.1.1.3, tidyverse.2.0.0, ggridges.0.5.6, sjPlot.2.8.16, fossil_0.4.0.
Access information
Other publicly accessible locations of the data:
Data was derived from the following sources:
- NA
