Data and code from: Odor and microbial analysis of human participants in association with mosquito attraction rates
Data files
Jul 14, 2026 version files 3.53 MB
-
Attraction-Analysis.zip
61.99 KB
-
Demographic-Analysis.zip
14.87 KB
-
Linear-Mixed-Model.zip
37.72 KB
-
Microbiome-Analysis.zip
3.37 MB
-
Odor-Analysis.zip
27.76 KB
-
README.md
21.45 KB
Abstract
Humans are not equally attractive to mosquitoes, leaving some individuals more vulnerable to mosquito-borne illnesses than others. The risk of a mosquito biting a human depends on multiple cues, with body odor amongst the most crucial. In this study, we measured the attraction of Aedes aegypti, Aedes albopictus, and Culex quinquefasciatus mosquitoes in a uniport olfactometer to determine species-specific attraction rates for each of our 119 human participants. Mosquito species typically did not favor the same participants. Each species attraction to participants was associated with distinct odor and bacterial signatures. For example, Ae. aegypti and Cx. quinquefasciatus attraction was associated with the absence of repulsive odors like cyclic alcohols and monoterpenes, while Ae. albopictus attraction was associated with the presence of attractive odors like ketones. In addition, Ae. aegypti, but not other species tested, were slightly more attracted to male than female participants. We show that mosquito species respond differently to individual humans and highlight cues that may regulate mosquito behavior.
Dataset DOI: 10.5061/dryad.vdncjsz97
Description of the data and file structure
Humans are not equally attractive to all mosquito species, but taking a human blood meal is integral for anthropophilic mosquitoes. Species such as Aedes aegypti, Aedes albopictus, and Culex quinquefasciatus inhabit the Americas and spread diseases like dengue, yellow fever, and West Nile fever. Mosquitoes use a range of cues to find a suitable host for a blood meal, including human body odor. This is largely derived from the transformations of skin secretions by the human skin microbiome into volatile organic compounds. Previous studies have shown that ratios of certain human skin bacteria can influence some mosquito species’ host-seeking behavior both in vivo and in vitro. Yet, cross-comparison of individual human volatilomes and microbiomes to the attraction behaviors of multiple mosquito species has yet to be examined.
In this study, we assessed the attraction rates of 119 individuals to Ae. aegypti, Ae. albopictus, and Cx. quinquefasciatus female mosquitoes. We obtained the volatilome of the arm of each participant through dynamic headspace sampling. Finally, we performed full-length 16S rRNA sequencing to detect microbial communities in humans that varied in attraction to each mosquito species. We found that attraction is associated with different odor and bacterial signatures for each mosquito species. The dataset contains the odor compounds and microbes detected for our deidentified participants.
Files and variables
File Structure
All files are provided in .zip format and named according to the collected data (attraction, odor, bacteria, or age). Additional Each .zip includes a menu.xlsx describing its internal files and purpose.
Input Files
File: Attraction-Analysis.zip
Description: Input files used for statistical analysis of attraction to Aedes aegypti, Aedes albopictus, or Culex quinquefasciatus for our human participants. Please note that ethnicity and date data have been removed from these file,s and ages have been changed to ranges to meet Dryad requirements for approval, but are available upon request. Contains 8 files:
· Aegypti_Samples.csv: Contains data for attraction of Ae. aegypti to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The attraction of each participant was assessed three times, which are labelled “Trial 1,” “Trial 2,” and “Trial 3.” Using these trial data, the mean (“Mean”), standard deviation (“Standard Deviation”), coefficient of variation (“CoefficientOfVariation”), and order norm (“OrderNorm”) were calculated for each participant. Participants that qualified for the high attraction group are labelled “HA” in the final column, while those that qualified for the low attraction group are labelled “LA” in the final column. Cells left blank indicated the participant qualified for the “Mid” attraction group.
· Albopictus_Samples.csv: Contains data for attraction of Ae. albopictus to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The attraction of each participant was assessed three times, which are labelled “Trial 1,” “Trial 2,” and “Trial 3.” Using these trial data, the mean (“Mean”), standard deviation (“Standard Deviation”), coefficient of variation (“CoefficientOfVariation”), and order norm (“OrderNorm”) were calculated for each participant. Participants that qualified for the high attraction group are labelled “HA” in the final column, while those that qualified for the low attraction group are labelled “LA” in the final column. Cells left blank indicated the participant qualified for the “Mid” attraction group.
· Culex_Samples.csv: Contains data for attraction of Cx. quinquefasciatus to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The attraction of each participant was assessed three times, which are labelled “Trial 1,” “Trial 2,” and “Trial 3.” Using these trial data, the mean (“Mean”), standard deviation (“Standard Deviation”), coefficient of variation (“CoefficientOfVariation”), and order norm (“OrderNorm”) were calculated for each participant. Participants that qualified for the high attraction group are labelled “HA” in the final column, while those that qualified for the low attraction group are labelled “LA” in the final column. Cells left blank indicated the participant qualified for the “Mid” attraction group.
· Attraction-Data-w-Sex.csv: Contains data for attraction of all three mosquito species to our participant. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. Additional demographic data includes age ranges, self-reported sex, race, and ethnicity. The attraction of each participant was assessed three times for each mosquito species as indicated in the “trial” and “attraction” columns. Finally, the temperature of each participant is listed in the “Temperature” column.
· Supplemental-data-2.csv/supplemental-data-2.xlsx: CSV and Excel files showing the results of the average attraction and standard deviation of each participant, as well as their cohort classification (high or low attraction). The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The mean and standard deviation for each participant’s attraction to each of the three mosquitoes is labelled (“Aegypti Mean,” “Aegypti Standard Deviation,” “Albopictus Mean,” “Albopictus Standard Deviation,” “Culex Mean,” and “Culex Standard Deviation”). The attraction group for each participant is labelled in the columns “Aegypti Attraction,” “Albopictus Attraction,” and “Culex Attraction,” where HA means high attraction, LA means low attraction, and Mid means in the middle.
· Uniport-attraction-data.csv: Contains all data from uniport attraction trials including participant ID (listed as “SUBJECT”), number of attracted mosquitoes for each species (listed as “ATTRACTED”), number of activated mosquitoes for each species (listed as “ACTIVATED”), number of unactivated mosquitoes for each species (listed as “UNACTIVATED”), number of dead mosquitoes, number of total alive mosquitoes, and the percent of attracted mosquitoes for each species.
· blank-data-all-spp.csv: Contains data for attraction rates of each mosquito species in the uniport olfactometer when only carbon dioxide and airflow were present. Each cell under the column “SUBJECT” is listed as “Blank” to indicate there was no participant assessed. The number of attracted mosquitoes for each species (listed as “ATTRACTED”), number of activated mosquitoes for each species (listed as “ACTIVATED”), number of unactivated mosquitoes for each species (listed as “UNACTIVATED”), number of dead mosquitoes, number of total alive mosquitoes, and the percent of attracted mosquitoes for each species. Each blank was run prior to each participant trial. For any attraction 15% or greater, the uniport olfactometer was cleaned and re-run until the attraction rate was under 15%. See Methods for more details.
· attraction-groups-script. R: An R/R Studio script containing all statistical analysis and visualization of the .csv files.
File: Odor-Analysis.zip
Description: Input files used for statistical analysis of odor profiles for our human participants. Contains 3 files:
· sample-attraction-all-mosquitoes.csv: Metadata describing the level of attraction for all three mosquito species to each of the participants. The first column labelled “Sample_ID” refers to the de-identified four-digit ID given to each participant in the study, with the word “sample” in front. Attraction levels are abbreviated as HA for high attraction, LA for low attraction, and Mid for middle attraction group.
· normalized-odors-samples.csv: Normalized odors detected via dynamic headspace sampling and identified via gas chromatography-mass spectrometry. The first column labelled “Sample_ID” refers to the de-identified four-digit ID given to each participant in the study, and the remaining columns are the chemicals detected.
· arm-odors-script. R: An R/R Studio script containing all statistical analysis and visualization of the .csv files.
File: Microbiome-Analysis.zip
Description: Input files used for statistical analysis of the microbiomes for our human participants. Contains 17 files:
· taxonomyfixed_otu_table.biom: Contains taxonomic and read count information for 16S sequencing for all participants.
· aegypti-attraction-metadata.txt: Metadata describing the level of attraction for Aedes aegypti mosquitoes to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. Attraction levels in the “attraction.by.mean” column are abbreviated as HA for high attraction, LA for low attraction, and Mid for middle attraction group. Two controls are noted as “Control.”
· albopictus-attraction-metadata.txt: Metadata describing the level of attraction for Aedes albopictus to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. Attraction levels in the “attraction.by.mean” column are abbreviated as HA for high attraction, LA for low attraction, and Mid for middle attraction group. Two controls are noted as “Control.”
· culex-attraction-metadata.txt: Metadata describing the level of attraction for Culex quinquefasciatus to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. Attraction levels in the “attraction.by.mean” column are abbreviated as HA for high attraction, LA for low attraction, and Mid for middle attraction group. Two controls are noted as “Control.”
· MVF_Metadata.csv: Metadata describing the self-reported sex of our participants. The first column labelled “#NAME” refers to the de-identified four-digit ID given to each participant in the study. Attraction levels in the “SampleType” column are the self-reported sex of the participant. Two controls are noted as “NA” in the “SampleType” column.
· phylotree_mafft_rooted.nwk: Phylogenetic tree for creation of our phyloseq object.
· microbiome-analysis-script. R: An R/R Studio script containing all statistical analysis and visualization of the microbiome data.
· supplemental_data_1.csv/supplemental_data_1.xls: CSV and Excel files showing the results of the Wilcoxon rank-sum test used for the male/female heat tree. The taxonomic assignment for the 16S results is listed in “Taxonomy Name.” The columns labelled “Group 1” and “Group 2” indicate the groups being compared in the Wilcoxon rank-sum test. Here, “Group 1” is “Male” and “Group 2” is “Female.” The remaining three columns are results from the test, including the log2 median ratio, median difference, and Wilcox p-value. Significant features are highlighted in orange in the Excel file.
· supplemental_data_3.csv/supplemental_data_3.xls: CSV and Excel files showing all 246 operational taxonomic units listed as “OTU” and the respective taxonomic assignment listed as “Taxa” detected on our 119 participants.
· supplemental_data_4.csv/supplemental_data_4.xls: CSV and Excel files showing the results of the Wilcoxon rank-sum test used for the Ae. aegypti heat tree. The taxonomic assignment for the 16S results is listed in “Taxonomy Name.” The columns labelled “Group 1” and “Group 2” indicate the groups being compared in the Wilcoxon rank-sum test. Here, “Group 1” is “HA,” meaning high attraction a,nd “Group 2” is “LA,” meaning low attraction. The remaining three columns are results from the test, including the log2 median ratio, median difference, and Wilcox p-value. Significant features are highlighted in orange in the Excel file.
· supplemental_data_5.csv/supplemental_data_5.xls: CSV and Excel files showing the results of the Wilcoxon rank-sum test used for the Ae. albopictus heat tree. The taxonomic assignment for the 16S results is listed in “Taxonomy Name.” The columns labelled “Group 1” and “Group 2” indicate the groups being compared in the Wilcoxon rank-sum test. Here, “Group 1” is “HA,” meaning high attraction, and “Group 2” is “LA,” meaning low attraction. The remaining three columns are results from the test, including the log2 median ratio, median difference, and Wilcox p-value. Significant features are highlighted in orange in the Excel file.
· supplemental_data_6.csv/supplemental_data_6.xls: CSV and Excel files showing the results of the Wilcoxon rank-sum test used for the Cx. quinquefasciatus heat tree. The taxonomic assignment for the 16S results is listed in “Taxonomy Name.” The columns labelled “Group 1” and “Group 2” indicate the groups being compared in the Wilcoxon rank-sum test. Here, “Group 1” is “HA,” meaning high attraction, and “Group 2” is “LA,” meaning low attraction. The remaining three columns are results from the test, including the log2 median ratio, median difference, and Wilcox p-value. Significant features are highlighted in orange in the Excel file.
File: Demographic-Analysis.zip
Description: Input files used for statistical analysis of age distribution for our human participants. Please note that ethnicity and date data have been removed from these files, and ages have been changed to ranges to meet Dryad requirements for approval, but are available upon request. Contains 3 files:
· HA_LA_Aegypti_Super_Table.csv: Metadata containing demographic, attraction, odor, and microbial information. The first column labelled “Unique ID” refers to the de-identified four-digit ID given to each participant in the study. The attraction of each participant was assessed three times, which are labelled either “Aegypti (1/2/3)”, “Albo (1/2/3)”, or “Culex (1/2/3).” The columns “Age range,” “Sex,” “Race,” and “Ethnicity” are all self-reported demographic information about our participants. Temperature was measured and recorded for each participant during their visit. The columns starting with “1-Penten-3-ol” and ending with “Norbourbonone” are normalized odors from gas chromatography-mass spectrometry. The remaining columns are operational taxonomic units (OTUs) from the 16S results. The matching taxonomic assignment for these OTUs can be found in “supplemental_data_3.csv” or “supplemental_data_3.xls.”
· age-counts.csv: Metadata describing the self-reported ages of our human participants. The “Age” is the reported age, and the “Num” column is an abbreviation for “number,” indicating the number of participants of the age in the first column.
· age-script. R: An R/R Studio script containing all statistical analysis and visualization of the age distribution data.
File: Linear-Mixed-Model.zip
Description: Input files used to create the linear mixed model of attraction for each mosquito species. Please note that ethnicity and date data have been removed from these file,s and ages have been changed to ranges to meet Dryad requirements for approval, but are available upon request. Contains 7 files.
· Aegypti_Samples.csv: Contains data for attraction of Ae. aegypti to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The attraction of each participant was assessed three times, which are labelled “Trial 1,” “Trial 2,” and “Trial 3.” Using these trial data, the mean (“Mean”), standard deviation (“Standard Deviation”), coefficient of variation (“CoefficientOfVariation”), and order norm (“OrderNorm”) were calculated for each participant. Participants that qualified for the high attraction group are labelled “HA” in the final column, while those that qualified for the low attraction group are labelled “LA” in the final column. Cells left blank indicated the participant qualified for the “Mid” attraction group.
· Albopictus_Samples.csv: Contains data for attraction of Ae. albopictus to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The attraction of each participant was assessed three times, which are labelled “Trial 1,” “Trial 2,” and “Trial 3.” Using these trial data, the mean (“Mean”), standard deviation (“Standard Deviation”), coefficient of variation (“CoefficientOfVariation”), and order norm (“OrderNorm”) were calculated for each participant. Participants that qualified for the high attraction group are labelled “HA” in the final column, while those that qualified for the low attraction group are labelled “LA” in the final column. Cells left blank indicated the participant qualified for the “Mid” attraction group.
· Culex_Samples.csv: Contains data for attraction of Cx. quinquefasciatus to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The attraction of each participant was assessed three times, which are labelled “Trial 1,” “Trial 2,” and “Trial 3.” Using these trial data, the mean (“Mean”), standard deviation (“Standard Deviation”), coefficient of variation (“CoefficientOfVariation”), and order norm (“OrderNorm”) were calculated for each participant. Participants that qualified for the high attraction group are labelled “HA” in the final column, while those that qualified for the low attraction group are labelled “LA” in the final column. Cells left blank indicated the participant qualified for the “Mid” attraction group.
· Aegypti_TableClean1.csv: Contains data for attraction, sex, age, race, ethnicity, and temperature for each participant. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. The column labelled “trial” indicates the attraction of the listed mosquito species to that participant, and each participant was assessed three times. The column labelled “Attraction” was the calculated attraction level as a percent. Columns labelled “dim(1-12)” are each of the 12 dimensions from the PCoA using the odors from the gas chromatography-mass spectrometry. Columns labelled as “nnmfdim(1-4)” are each of the 4 dimensions from the NMF using the bacteria from the 16S sequencing results.
· Attraction-Data-w-Sex.csv: Contains data for attraction of all three mosquito species to our participants. The first column labelled “Subject” refers to the de-identified four-digit ID given to each participant in the study. Additional demographic data includes age ranges, self-reported sex, race, and ethnicity. The attraction of each participant was assessed three times for each mosquito species as indicated in the “trial” and “attraction” columns. Finally, the temperature of each participant is listed in the “Temperature” column
· Study-participant-test-dates-and-time.csv: The column labelled “Unique ID” refers to the de-identified four-digit ID given to each participant in the study. The column labelled “Column1” indicates the time of day of testing, where a number 1 indicates a morning visit and a number 2 indicates an afternoon visit. Exact dates are considered direct identifiers, so we have removed this information from this public dataset.
· Linear-mixed-model-script. R: An R/R Studio script containing all information about the creation and evaluation of three linear mixed models of mosquito attraction.
Access information
Other publicly accessible locations of the data:
- The 16S rRNA sequencing data files are available for download on NCBI Sequence Read Archive with Bioproject ID PRJNA1415315.
Human subjects data
We have received explicit consent from each participant to publish this de-identified data in the public domain. We have randomly generated four-character ID's for each participant to remove all identifying information.
