Variation in the resource environment affects patterns of seasonal adaptation at phenotypic and genomic levels in Drosophila melanogaster
Data files
Sep 05, 2025 version files 3.97 GB
-
bray_curtis_distance_matrix.qza
144.22 KB
-
bray_curtis_pcoa_results.qza
415.76 KB
-
correcttaxonomy.qza
571.79 KB
-
demultiplexed_fastq_gz.zip
2.63 GB
-
evenness_vector.qza
113.22 KB
-
faith_pd_vector.qza
113.36 KB
-
glm.A.TP1_5.RData
26.86 MB
-
glm.AppleVBloom.T1_5.RData
86.37 MB
-
glm.B.TP1_5.RData
26.88 MB
-
observed_features_nw.csv
4.27 KB
-
observed_features_vector.qza
111.25 KB
-
orch2020_Baseline_filtered.RData
53.70 MB
-
orch2020_filtered.RData
1.14 GB
-
orchard2016smetadata.csv
116.48 KB
-
README.md
22.95 KB
-
rooted-tree.qza
489.70 KB
-
shannon_vector.qza
113.11 KB
-
table-no-wolbachia-exact.qza
558.16 KB
-
unweighted_unifrac_distance_matrix.qza
505.01 KB
-
unweighted_unifrac_pcoa_results.qza
366.20 KB
-
weighted_unifrac_distance_matrix.qza
554.55 KB
-
weighted_unifrac_pcoa_results.qza
348.35 KB
Abstract
Natural populations often experience heterogeneity in the quality and abundance of environmentally acquired resources across both space and time, and this variation can influence population demographics and evolutionary dynamics. In this study, we directly manipulated diet in replicate populations of Drosophila melanogaster cultured in experimental mesocosms in the field. We found no significant effect of resource variation on estimates of adult census size. Resource variation altered patterns of phenotypic and genomic evolution across replicate populations; however, we find that this effect is secondary to selection driven by the fluctuating seasonal environment. Seasonal adaptation was observed for all traits assayed and elicited genome-wide signatures of selection. In contrast, adaptation to the resource environment was trait-specific and exhibited an oligogenic architecture. This illustrates the capacity of populations to adapt to a specific axis of variation (the resource environment) without hindering the adaptive response to seasonal change. This, in turn, suggests that resource variation may be an important force driving fluctuating selection across natural populations, ultimately contributing to the maintenance of genetic and phenotypic variation.
Dataset DOI: 10.5061/dryad.t1g1jwtfr
Description of the data and file structure
This database contains raw and analyzed data related to a manuscript published in Evolution Letters (EVL3-25-0011.R1) titled "Variation in the resource environment affects patterns of seasonal adaptation at phenotypic and genomic levels in Drosophila melanogaster." Any questions or concerns should be directed to the corresponding authors, Jack Beltz (jkbeltz@sas.upenn.edu), Mark Bitter (mcbitter@stanford.edu), and Paul Schmidt (schmidtp@sas.upenn.edu).
This study aimed to investigate how variation in the resource environment (food supply) influenced the adaptive outcomes of replicate populations of Drosophila melanogaster as they expanded and evolved across seasonal time, as well as to describe the variation in microbial associations formed. 18 populations of Drosophila were founded in large outdoor mesocosms, and adults from both treatment groups (Apple and Bloom, two diets) were sampled for whole genome and 16s rRNA sequencing.
This contains the raw demultiplexed fastq files ("demultiplexed_fastq_gz.zip") associated with 16s V1V2 sequencing of Drosophila used in this experiment, as well as all QIIME2 files generated after analysis following the pipeline (https://github.com/jbisanz/16Spipelines.git), and using "orchard2016smetadata.csv" as the metadata file, which contains multiple variables that describe the treatment conditions applied to each sample.
Additionally, this database contains large RData files generated during the analysis of WGS data. WGS FASTA reads can be found here (https://www.ncbi.nlm.nih.gov/sra/PRJNA1306087), and the associated bioinformatics analysis, including scripts that utilized these RData files and further description of how to conduct this, can be found here (https://github.com/jkbeltz/EVL3-25-0011.R1.git).
Files and variables
File: demultiplexed_fastq_gz.zip
Description: Compressed folder containing all 16s reads needed to reproduce the analysis. When unzipped, each file represents all sequence reads and quality scores identified in that sample. File name corresponds to the sample ID, which is further described in the metadata file "orchard2016smetadata.csv".
File: glm.A.TP1_5.RData
Description: A series of files viewable in the environment tab when loaded in R Studio. Files generated following code at https://github.com/jkbeltz/EVL3-25-0011.R1.git , using whole genome sequence reads available https://www.ncbi.nlm.nih.gov/sra/PRJNA1306087. Specifically, this data used a generalized linear model to identify regions of the Drosophila genome which show consistent directional changes across all sampling timepoints (1-5), for populations fed a semi-natural apple-based diet (n=9). Contains the following files, which describe the genomic region identified.
-samps: sample information whereby row order corresponds to column order of the allele frequency matrix and effective coverage matrix.
-afmat: numeric data frame containing all haplotype-derived allele frequencies
-eec: numeric data frame containing estimated effective coverage for each sample/site
-sites: dataframe containing chromosome and site information (corresponding to rows of afmat and eec_
File: glm.B.TP1_5.RData
Description: A series of files viewable in the environment tab when loaded in R Studio. Files generated following code at https://github.com/jkbeltz/EVL3-25-0011.R1.git , using whole genome sequence reads available https://www.ncbi.nlm.nih.gov/sra/PRJNA1306087. Specifically, this data used a generalized linear model to identify regions of the Drosophila genome that show consistent directional changes across all sampling timepoints (1-5), for populations fed a lab-utilized molasses-based diet (n=9). Contains the following files, which describe the genomic region identified across samples.
-samps: sample information whereby row order corresponds to column order of the allele frequency matrix and effective coverage matrix.
-afmat: numeric data frame containing all haplotype-derived allele frequencies
-eec: numeric data frame containing estimated effective coverage for each sample/site
-sites: dataframe containing chromosome and site information (corresponding to rows of afmat and eec_
File: glm.AppleVBloom.T1_5.RData
Description: A series of files viewable in the environment tab when loaded in R Studio. Files generated following code at https://github.com/jkbeltz/EVL3-25-0011.R1.git , using whole genome sequence reads available https://www.ncbi.nlm.nih.gov/sra/PRJNA1306087. Specifically, this data used a generalized linear model to identify regions of the Drosophila genome which show consistent directional changes across all sampling timepoints (1-5), across populations fed a lab-utilized molasses-based diet (n=9), as well as populations fed a lab-utilized molasses-based diet (n=9). Contains the following files, which describe the genomic region identified across samples.
-samps: sample information whereby row order corresponds to column order of the allele frequency matrix and effective coverage matrix.
-afmat: numeric data frame containing all haplotype-derived allele frequencies
-eec: numeric data frame containing estimated effective coverage for each sample/site
-sites: dataframe containing chromosome and site information (corresponding to rows of afmat and eec_
File: orch2020_Baseline_filtered.RData
Description: A series of files viewable in the environment tab when loaded in R Studio. Files generated following code at https://github.com/jkbeltz/EVL3-25-0011.R1.git , using whole genome sequence reads available https://www.ncbi.nlm.nih.gov/sra/PRJNA1306087 . Specifically, this data contains all allelic information across all samples in the founder population before Drosophila evolved in the field. Contains the following files, which describe the genomic region identified across samples.
-samps: sample information whereby row order corresponds to column order of the allele frequency matrix and effective coverage matrix.
-afmat: numeric data frame containing all haplotype-derived allele frequencies
-eec: numeric data frame containing estimated effective coverage for each sample/site
-sites: dataframe containing chromosome and site information (corresponding to rows of afmat and eec_
File: orch2020_filtered.RData
Description: A series of files viewable in the environment tab when loaded in R Studio. Files generated following code at https://github.com/jkbeltz/EVL3-25-0011.R1.git , using whole genome sequence reads available https://www.ncbi.nlm.nih.gov/sra/PRJNA1306087 . Specifically, this data contains all allelic information across all samples in all field-evolving populations (n=18) across five timepoints. Contains the following files, which describe the genomic region identified across samples.
-samps: sample information whereby row order corresponds to column order of the allele frequency matrix and effective coverage matrix.
-afmat: numeric data frame containing all haplotype-derived allele frequencies
-eec: numeric data frame containing estimated effective coverage for each sample/site
-sites: dataframe containing chromosome and site information (corresponding to rows of afmat and eec_
File: bray_curtis_distance_matrix.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a matrix that can be used to plot samples in multi-dimensional space, using the Bray-Curtis distance method to estimate the relative similarity between samples based on the relative abundance of each major taxonomic group. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: evenness_vector.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a vector describing the "evenness" of the microbial communities in each sample, which is a metric that describes how evenly distributed the individuals of a particular taxonomic group are in each sample. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: bray_curtis_pcoa_results.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for coordinates used to plot samples in multi-dimensional space, using the Bray-Curtis distance method to estimate the relative similarity between samples based on the relative abundance of each major taxonomic group. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: faith_pd_vector.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a vector describing the faith phylogenetic distance of the microbial communities in each sample, which is a metric that describes the relative diversity of the microbial communities based on phylogeny, as described by Faith (1992). This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: observed_features_nw.csv
Description: This is a CSV file that describes the number of unique microbial ASVs (a proxy for species) found in each sample. This file was originally generated as a .qza but was extracted for convenience as a .csv. Also, this file does not include any ASVs assigned to the genus Wolbachia. This file is generated and utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: observed_features_vector.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a vector describing the total number of unique sequence variants (ASVs, a proxy for species) found in each sample. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: correcttaxonomy.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a vector describing the closest taxonomic assignment (KPCOFGS) for each unique ASV identified across all samples. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: orchard2016smetadata.csv
Description: This CSV is the table that describes each sample, which was 16s rRNA sequenced and is used as the critical second file for the Qiime2 analysis. This table contains many variables that are largely irrelevant to the analysis that describe the primer sequence used, location in the sequencing plate, project identifiers, among others. Variables relevant to analysis include "SampleID," which is the unique alphanumeric identifier used throughout the analysis to refer to a particular sample. "SampleType" describes whether this sample contains genetic material from a fly population or environmental control (further described in "sample_type_detail) which is mirrored by "host_species*".* "study_day", "date_collected" and "time_point" are three distinct manners to describe when in the course of the experiment the samples were collected. "cage_treatment" is a categorical variable that describes the treatment given to each population of drosophila as they are evolving in the field environment, Apple is the semi-natural apple based diet, Bloom is the molasses based modified "bloomington" lab diet, Apple_AT is the apple diet enriched with an endemic acetobacter species, AppleLB is the apple diet enriched with an endemic lactobacillus species, and axenic_AT and axenic_LB refer to populations from the field which were decorionated and reared a single generation in lab before sequencing. "study group" is the variable that identifies each unique population moving through time AT / LB populations are identified by Cage # and AT or LB, while Apple vs Bloom populations are identified Cage A/B-#, and founder populations are proceeded with the word "baseline". "sample_type_detail" and "collection_type" describe how the samples were collected, "flies" refers to wild aspirated flies directly from the cages, "dry stored f1/f2s" refers to flies reared one or two generations in lab environment and collected using co2 anesthesia, and finally "cage food" are samples scraped from the surface of the food supply of the corresponding cage. Finally, "final_library_conc_ng_ul" is the only metric of this type that should be utilized and describes the concentration of RNA found in the sample after PCR amplification and before sequencing in nanograms /microliter and is utilized as a means to correct for variation in concentration when determining the relative abundance of read counts in a sample.
File: shannon_vector.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a vector describing the diversity of the microbial communities found in each sample as calculated using the Shannon diversity index, which quantifies species diversity in a community by considering both the number of species present (richness) and their relative abundance (evenness). This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: rooted-tree.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a vector describing the relative phylogenetic distance between the ASVs described in correcttaxonomy.qza, and can be utilized to construct a rooted tree using Bayesian statistics. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: unweighted_unifrac_pcoa_results.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for coordinates used to plot samples in multi-dimensional space, using the unweighted unifrac distance method (which does not account for the relative abundance of a particular taxon) to estimate the relative similarity between samples based on the relative abundance of each major taxonomic group. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: table-no-wolbachia-exact.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a table which identifies each unique taxanomic group (ASV) by an alphanumeric identifier (rows) which corresponds to the taxonomic description found in correcttaxonomy.qza, and describes the number of sequence reads of that particular ASV which are found in each sample (columns) which are further described in the orchard2016smetadata.csv. Also, this file does not include any ASVs assigned to the genus Wolbachia. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: weighted_unifrac_pcoa_results.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for coordinates used to plot samples in multi-dimensional space, using the weighted unifrac distance method (which accounts for the relative abundance of a particular taxon) to estimate the relative similarity between samples based on the relative abundance of each major taxonomic group. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: weighted_unifrac_distance_matrix.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a matrix used to compare samples in multi-dimensional space, using the weighted unifrac distance method (which accounts for the relative abundance of a particular taxon) to estimate the relative similarity between samples based on the relative abundance of each major taxonomic group. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
File: unweighted_unifrac_distance_matrix.qza
Description: Output file generated in Qiime2 analysis of 16s rRNA reads found in "demultiplexed_fastq_gz.zip", using analytic pipeline described at https://github.com/jbisanz/16Spipelines.git. This is an archive containing data, metadata, and provenance information from a QIIME 2 analysis, which specifically contains information for a matrix used to compare samples in multi-dimensional space, using the unweighted unifrac distance method (which does not account for the relative abundance of a particular taxa) to estimate the relative similarity between samples based the relative abundance of each major taxonomic group. This archive is utilized in the full 16s rRNA analysis, code for which is found at https://github.com/jkbeltz/EVL3-25-0011.R1.git.
Code/software
This data requires Python and R, and utilized Qiime2 to recreate (https://docs.qiime2.org/2024.10/install/index.html).
All code necessary to reproduce the analysis for this manuscript can be found at https://github.com/jkbeltz/EVL3-25-0011.R1.git. 16s rRNA qiime2 analysis pipeline found at https://github.com/jbisanz/16Spipelines.git, and whole genome sequences found at https://www.ncbi.nlm.nih.gov/sra/PRJNA1306087.
