Data accompanying: Spatial and temporal structure of environmentally-acquired Caballeronia symbionts of a leaffooted bug
Data files
Jul 30, 2026 version files 959.27 KB
-
Dryad_Ravenscraft2025.zip
947 KB
-
README.md
12.27 KB
Abstract
Animals that acquire beneficial microbial symbionts from their environment risk acquiring a sub-optimal partner, or no partner at all. The leaffooted bug Leptoglossus zonatus (Coreidae) acquires its Caballeronia (Burkholderiaceae) bacterial symbiont from the environment, presumably from local soil. Despite large contributions to the bug’s fitness, young nymphs must re-acquire the symbiont every generation. To understand how the environmental reservoir of symbiont lineages shapes the insect’s biology, we examined the role of space and time in the distribution of Burkholderia sensu lato (including Caballeronia) strains in bugs and soils. We compared samples within trees, within plots, within cities and among different cities in the Southwest USA. We also sampled Caballeronia in L. zonatus within a pomegranate orchard over two years. We found high Caballeronia diversity both in soils (32 lineages) and in bugs (29 lineages). Caballeronia lineages were spatially structured among soils and bugs, with fewer shared as distance between samples increased. Where a bug develops, therefore, influences the symbiont strain it acquires, consistent with a process of passive spatial turnover. Also, while some Caballeronia subclade frequencies in bugs approximated frequencies in soils, the Coreiodea-associated subclade of Caballeronia (SBE𝛿) was enriched in bugs. There was only slight turnover of strains in insects over time, suggesting that geographic variation is much more important than temporal change. Ultimately, understanding how symbiont strains of varying local benefit are distributed in space and time will help us predict how geography and seasonality are related to host fitness in environmentally acquired symbioses.
https://doi.org/10.5061/dryad.fj6q5743s
Description of the data and file structure
Paper: Spatial and temporal structure of environmentally-acquired Caballeronia symbionts of a leaffooted bug
Authors: Alison Ravenscraft, Suzanne E. Kelly, David Haviland, Johnathan Adamson, Martha S. Hunter
Description: This README describes the data package accompanying the above publication. Files 1-6 are related to placement of the sequence variants onto a well-supported reference phylogeny via SEPP. These files are included in hopes that they will facilitate future studies that require phylogenetic placement of Burkholderia sensu lato sequences from amplicon data or Sanger sequencing. Researchers wishing to perform their own phylogenetic placement should start with the instructions (file #1). File 7 reports the partitions used to build the phylogeny presented in the main text. File 8 contains the representative sequences for the Caballeronia ASVs and lineages. Files 9-11 are newick phylogenies and File 12 provides the associated metadata. Files 13-17 provide the OTU table resulting from the amplicon sequencing and its associated metadata. Files 19 and 21 provide the same data packaged as R phyloseq objects. File 18, 20, and 22 are R scripts that analyze the data and plot the figures.
Files in Dryad_Ravenscraft2025.zip:
Phylogenetic placement of ASVs & phylogenies
- Ravenscraft2025_amplicon_phylogentic_placement_instructions.rtf: Instructions and QIIME 2 commands to perform SEPP placement of amplicon sequence variants onto a custom-built phylogeny. Uses a reference phylogeny (e.g. “Ravenscraft2025_full_reference_phylogeny.newick”), a RAxML info file (e.g. “Ravenscraft2025_full_RAXML_info.txt”), and an alignment (e.g. “Ravenscraft2025_full_reference_alignment.fasta”) to build a SEPP reference database (e.g. “Ravenscraft2025_full_SEPP_database.qza”). This database is then used for phylogenetic placement via QIIME 2.
- Ravenscraft2025_full_reference_alignment.fasta: DNA alignment of reference sequences. Five full-length genes (rpoB, rpoC, rplA, recA, and 16S rRNA) were extracted from Burkholderiaceae reference genomes. These data were supplemented with additional 16S rRNA sequences from isolates sequenced for this study. For the associated accession numbers, see “Ravenscraft2025_phylogeny_metadata.csv.” We performed protein alignments on the four protein-coding genes and a nucleotide alignment of the 16S using MAFFT.
- Ravenscraft2025_full_reference_alignment_AMPLICONREGION.fasta: DNA alignment of just the amplicon region for the reference sequences.
- Ravenscraft2025_full_reference_phylogeny.newick: A maximum-likelihood phylogeny in Newick format. Tree was built from the alignment in file ”Ravenscraft2025_full_reference_alignment.fasta” using RAxML with the GTR+Gamma model of nucleotide substitution, with separate partitions for each coding position of each protein coding gene plus one partition for the 16S rRNA. Partitions are provided in “Ravenscraft2025_phylogeny_partitions.txt.” The reference tree’s node support values were calculated using rapid bootstrapping which was halted automatically based on the MRE criterion. The tree was rooted with the outgroups Pandoraea pulmonicola and Pandoraea oxalativorans.
- Ravenscraft2025_full_RAXML_info.txt: Information file output by RAxML. This file was generated by extracting the aligned amplicon region from the aligned reference sequences (“Ravenscraft2025_full_reference_alignment_AMPLICONREGION.fasta”) and running RAxML with the GTR+Gamma model of nucleotide substitution.
- Ravenscraft2025_full_SEPP_database.qza:**** QIME-formatted SEPP reference database built from files “Ravenscraft2025_full_reference_alignment.fasta”, “Ravenscraft2025_full_reference_phylogeny.newick” and “Ravenscraft2025_full_RAXML_info.txt”.
- Ravenscraft2025_phylogeny_partitions.txt: The partitions used build the phylogeny in “Ravenscraft2025_full_reference_phylogeny.newick.”
- Ravenscraft2025_raw_repset.fasta:**** Representative sequences of the raw ASVs. To generate these, priming sites and poor-quality bases were removed from the 5’ and 3’ ends of the raw Illumina sequences using the program cutadapt. We discarded all reads that contained any unassigned bases (Ns) or had an expected error score greater than 2. We truncated the remaining reads at the first instance of a quality score less than 2. For each sequencing run (spatial and temporal), we used the DADA2 R package to merge paired ends and infer the bacterial strains present. Data from the two runs were then merged by sequence using the DADA2 package’s mergeSequenceTables() command. We performed de novo chimera checking and removal on the combined data. Note1: Since the tip-glommed lineages presented in the manuscript are named with the identifier of the raw ASV that was most abundant within the lineage prior to collapsing the tree tips, this file also provides the reference sequences for the collapsed lineages. Note2: This file contains sequence variants from the other bug species (L. brevirostris, L. fulvicornis, L. oppositus, L. phyllopus and Anasa sp - likely Anasa tristis) displayed in Fig S3.
- Ravenscraft2025_SEPP_tree.nwk: Phylogeny with the placement of the raw ASVs. “Ravenscraft2025_amplicon_phylogentic_placement_instructions.rtf” describes how this was generated.
- Ravenscraft2025_phylogeny_common_lineages.newick: Phylogeny from Fig 1 in the manuscript. Lineages detected in this study (names starting “sv”) were placed onto the reference phylogeny (file “Ravenscraft2025_ampunique_reference_phylogeny.newick”) with SEPP, and therefore have no bootstrap support values. The R script drops lineages that did not account for at least 1% of the reads in at least 2 samples from this tree prior to plotting it. Reference sequences that did not have close relatives to the amplicon linages were also dropped. An equivalent phylogeny with all lineages and all references is provided in file **“**Ravenscraft2025_phylogeny_all_lineages.newick.”
- Ravenscraft2025_phylogeny_all_lineages.newick: Phylogeny from Fig S1 in the manuscript. This is equivalent to file “Ravenscraft2025_phylogeny_common_lineages.newick,” except that all lineages (including those that did not account for at least 1% of the reads in at least 2 samples) and all reference sequences are included.
- Ravenscraft2025_phylogeny_metadata.csv: Metadata for the phylogenies in files “Ravenscraft2025_phylogeny_common_ASVs.newick” and “Ravenscraft2025_phylogeny_all_ASVs.newick” (Figures 1 and S1 in the manuscript).
- fulllab: Full name of the sequence
- newlab: Shortened name of the sequence (omits the strain)
- finallab: Shortest labels (genus is abbreviated, strain ID retained when needed for disambiguation.)
- genus: Genus epithet of the sequence.
- species: Species epithet of the sequence.
- strain: Strain ID of the sequence.
- hostbug: Insect species from which the isolate was derived (if any).
- accession: GenBank accession number.
- newclade: Genus or, in the case of Caballeronia, subclade of the sequence.
- notes: Notes.
Microbiome data in open access format
- Ravenscraft2025_raw_ASV_table.csv: The count (number of reads) of each raw bacterial ASV in each sample prior to rarefying and prior to collapsing tips of the phylogeny. Columns are raw bacterial ASVs, rows are samples (insects or soils). Note that the ASV identifiers here are not equivalent to the identifiers used in the manuscript and figures, because the manuscript presents lineages after collapsing the tips of the phylogeny; see also the note for the next file.
- Ravenscraft2025_raw_ASV_metadata.csv: Metadata for the raw ASVs in “Ravenscraft2025_raw_ASV_table.csv.” The “fasta” column provides the representative sequence for each ASV. (These match the sequences in “Ravenscraft2025_repset.fasta.”)
- Ravenscraft2025_tipglom_lineage_table.csv: The count (number of reads) of the bacterial lineages in each sample prior to rarefaction, but after collapsing the tips of the phylogeny at a cophenetic distance of 0.1 (using the “tip_glom” function from the R package “phyloseq”). Columns are bacterial lineages (which were interchangeably called “ASVs” or ‘lineages” in the manuscript), rows are samples (insects or soils). Note that each lineage is named with the identifier of the raw ASV that was most abundant within that lineage prior to collapsing the tree tips.
- Ravenscraft2025_lineage_metadata.csv: Metadata for the microbial lineages in “Ravenscraft2025_tipglom_ASV_table.csv” (after collapsing the tips of the phylogeny at a cophenetic distance of 0.1). “clade” is the genus or, in the case of Caballeronia, subclade to which the lineage belongs.
- Ravenscraft2025_sample_metadata.csv: Sample (insect and soil) metadata.
- seqname: Identifier of the sequencing data from a sample.
- run: Specifies which run the sample was sequenced on.
- project: The project a sample belongs to. “USDAgeo” is insects from the spatial survey, “USDAsoil” is soils from the spatial survey, and “temporal” is bugs from the temporal survey.
- specid: The identifier of a bug or soil sample. For the spatial survey, the first two characters indicate site, the third character indicates a tree, and the final digit indicates the individual bug collected from that tree and site.
- sampletype: Bug or soil.
- collectiondate: Date the bug or soil was collected.
- site: Location code where the sample was collected.
- tree: Identifies which tree the sample was collected from (or under, in the case of soils).
- patchtree: Unique identifiers for each tree.
- bugnum: The last digit of the specid; identifies a bug from a given site and tree.
- sexinstar: The sex (if an adult) or instar (if a nymph) of a bug.
- instar: The install of the bug, if a nymph.
- sex: The sex of a bug, if an adult.
- region: City where the sample was collected (Fresno CA, Bakersfield CA, or Tucson AZ).
- coordnorth: Latitude where sample was collected.
- coordwest: Longitude where sample was collected.
- notes: Notes.
- bugspecies: Insect species.
- hostplant: Plant on which the bug was collected.
- city: City where the sample was collected (Fresno CA, Bakersfield CA, or Tucson AZ).
- state: State where the sample was collected.
- mergegroup: Used to merge a couple samples that had multiple replicates on the run.
Microbiome data in R format & R scripts
- Ravenscraft2025_functions.R: R script with some convenient functions for phyloseq objects.
- Ravenscraft2025_raw_phyloseq.Rdata: An R phyloseq object including the raw OTU table (“Ravenscraft2025_raw_ASV_table.csv”), sample metadata (“Ravenscraft2025_sample_metadata.csv”), ASV metadata (“Ravenscraft2025_raw_ASV_metadata.csv”), and a phylogeny of the ASVs (extracted from “Ravenscraft2025_phylogeny_all_ASVs.newick”).
- Ravenscraft2025_collapseTips_assignClades.R: R script that collapses the raw ASVs from “Ravenscraft2025_raw_phyloseq.Rdata” into lineages with cophenetic distance = 0.1 and assigns clades to the resulting lineages. Generates “Ravenscraft2025_lineage_phyloseq.Rdata”
- Ravenscraft2025_lineage_phyloseq.Rdata: An R phyloseq object including the tip-glommed OTU table (“Ravenscraft2025_tipglom_lineage_table.csv”), sample metadata (“Ravenscraft2025_sample_metadata.csv”), ASV metadata (“Ravenscraft2025_lineage_metadata.csv”), and a phylogeny of the lineages (extracted from “Ravenscraft2025_phylogeny_all_ASVs.newick”).
- Ravenscraft2025_analysis.R: The main R script for data analysis. Requires “Ravenscraft2025_lineage_phyloseq.Rdata,” “Ravenscraft2025_functions.R”, “Ravenscraft2025_phylogeny_all_lineages.newick,” and “Ravenscraft2025_phylogeny_metadata.csv.”
Code/software
This Dryad package includes scripts in the R programming language and data in R format.
Files 1-6 are related to placement of the sequence variants onto a well-supported reference phylogeny via SEPP. These files are included in hopes that they will facilitate future studies that require phylogenetic placement of Burkholderia sensu lato sequences from amplicon data or Sanger sequencing. Researchers wishing to perform their own phylogenetic placement should start with the instructions (file #1). File 7 reports the partitions used to build the phylogeny presented in the main text. File 8 contains the representative sequences for the Caballeronia ASVs and lineages. Files 9-11 are newick phylogenies and File 12 provides the associated metadata. Files 13-17 provide the OTU table resulting from the amplicon sequencing and its associated metadata. Files 19 and 21 provide the same data packaged as R phyloseq objects. File 18, 20 and 22 are R scripts that analyze the data and plot the figures.
