Data from: Host phylogeny and microbiome composition predict gut nematode community composition within a diverse assemblage of African herbivores
Data files
Jul 09, 2026 version files 101.73 MB
-
Diet_OTU_RRA.csv
1.70 MB
-
Diet_OTU_taxonomy.csv
30.65 KB
-
Diet_sample_metadata.csv
364.78 KB
-
Host_species_trait_data.csv
9.48 KB
-
Microbiome_ASV_RRA.csv
89.79 MB
-
Microbiome_ASV_taxonomy.csv
8.21 MB
-
Microbiome_sample_metadata.csv
364.19 KB
-
Nemabiome_OTU_RRA.csv
248.74 KB
-
Nemabiome_OTU_taxonomy.csv
73.09 KB
-
Nemabiome_sample_metadata.csv
97.36 KB
-
Nematode-microbiome_code.R
205.52 KB
-
Pathogenic_microbe_sample_summary.csv
9.06 KB
-
Pathogenic_microbe_spp_RRA.csv
594.59 KB
-
Pathogenic_microbe_spp_taxonomy.csv
8.11 KB
-
README.md
24.31 KB
Abstract
Gut nematodes influence animal health and fitness, with effects shaped by their abundance and community composition. While closely related host species tend to harbor similar nematodes, the relative roles of host ecology and phylogeny in structuring nematode communities remain unclear. Here, we assess how host holobiont traits—body size, diet, space use, potentially pathogenic gut microbes, and overall gut microbiome composition—predict gut nematode abundance and composition, while controlling for host phylogenetic relatedness. We jointly analyze DNA metabarcoding data on nematodes, microbiomes, and diets from 17 free-ranging African herbivore species. Host phylogeny and microbiome composition were the strongest predictors of nematode community structure. Physical proximity and diet also contributed, though to lesser extents, whereas body size did not. Nematode abundance correlated positively with richness of putative pathogenic bacteria, which in turn increased with diet richness. Nematode presence/absence covaried with microbiome and diet composition, and we identified pairwise associations between nematodes, putative pathogenic bacteria, and diet plants. Our findings illustrate that host ecology and phylogeny jointly influence gut nematode communities. In particular, the gut microbiome is a key predictor of nematode communities, even after accounting for host phylogeny, emphasizing the ecological interconnectedness of these gut constituents.
Dataset DOI: 10.5061/dryad.sxksn03k0
Description of the data and file structure
DNA metabarcoding data on diet, microbiome, and gut nematodes were generated from fecal samples collected between 2013–2017 in Laikipia County, Kenya. Across datasets, data were available for 17 herbivore species (14 wild, 3 domestic): from these species, 1,075 samples were analyzed for diet (trnL-P6 marker), 354 for microbiome (16S-V4 rRNA marker), and 266 for nematode composition (ITS-2 rRNA marker). An additional 457 samples were tested for nematode presence via qPCR; primarily those that tested positive for nematodes (qPCR cycle threshold [Ct]<35) were sequenced for nematode composition. All data types were available for 87 samples.
Data were generated via Illumina sequencing and processed using consistent pipelines. Sequences were then clustered into mOTUs for diet (N=213) and nematode data (N=96), and into ASVs for microbiome data (N=29,308). Putative pathogenic bacterial taxa within microbiome data were identified using the 16SPIP bioinformatics pipeline; 144 species of putative pathogenic bacteria were identified within microbiome data.
See associated publication for complete details on how data were generated and processed.
Files and variables
File: Nemabiome_sample_metadata.csv
Description: Sample metadata for the fecal samples in 'Nemabiome_OTU_RRA.csv'. Within this file, ‘NA’ refers to ‘not available’, indicating that the corresponding information was not available for that particular sample.
Variables
- Sample_ID: the unique ID of the fecal sample, corresponding to those in the published herbivore gut nematode dataset (Titcomb et al. 2022)
- Location: location from which the fecal sample was collected (all 'Mpala', corresponding to the Mpala Research Centre')
- Species: common name of the herbivore species from which the fecal sample was derived
- MSW93_Order: the mammalian order to which the herbivore species belongs, per the 1993 Mammal Species of the World taxonomy
- MSW93_Family: the mammalian family to which the herbivore species belongs, per the 1993 Mammal Species of the World taxonomy
- MSW93_Genus: the mammalian genus to which the herbivore species belongs, per the 1993 Mammal Species of the World taxonomy
- MSW93_Species: the mammalian species to which the herbivore species belongs, per the 1993 Mammal Species of the World taxonomy
- MSW93_Binomial: the scientific name (genus followed by species) of the herbivore species, per the 1993 Mammal Species of the World taxonomy
- Population_Size: estimated population size of the herbivore species
- Period: the collection period during which the fecal sample was collected, in the format 'YYYYMMM'
- BM_KG: the body mass of the herbivore species, in kilograms
- RS_KM2: the range size of the herbivore species, in kilometers squared
- GS: the gut size of the herbivore species
- GUT: the gut type of the herbivore species, where 'FG' corresponds to foregut-fermenting and 'HG' to hindgut-fermenting
- UNDERSTORY_SP_MEAN: the average relative read abundance (RRA) of understory plant mOTUs in the diet of the herbivore species from which the fecal sample was derived
- UNDERSTORY_SP_SD: the standard deviation in the relative read abundance (RRA) of understory plant mOTUs in the diet of the herbivore species from which the fecal sample was derived
- Sample.date: the date the fecal sample was collected, in the format 'DD/MM/YYYY'
- Longitude: the longitude at which the fecal sample was collected
- Latitude: the latitude at which the fecal sample was collected
- Rain90: the amount of rain in the 90 days preceding the collection of the fecal sample, in millimeters
- UNDERSTORY_PROP: the total relative read abundance (RRA) of understory plant mOTUs in the particular fecal sample
- rd_2_qpcr_conducted: whether or not qPCR was conducted on the sample
- rd_2_qpcr_mean_ct: the mean qPCR cycle threshold (Ct) for that sample; Ct > 35 were taken to indicate that there was no nematode DNA in the sample
- rd_2_qpcr_present: whether or not the sample tested positive for nematode DNA (Ct < 35)
- rd2_sequenced: whether or not the amplified DNA was sequenced for nemabiome composition
- note: any relevant notes about the fecal sample
File: Nematode-microbiome_code.R
Description: Code to conduct all analyses and generate all figures for the paper associated with this dataset.
File: Pathogenic_microbe_sample_summary.csv
Description: Summary data on the richness and relative abundances of putative pathogenic bacteria species within each sample. Within this file, ‘NA’ refers to ‘not available’, indicating that data on that species was not available from that particular sample.
Variables
- Sample_ID: the unique ID of the fecal sample, corresponding to those in the published herbivore microbiome dataset (Kartzinel et al. 2019)
- Pathogenic_spp_N: the number of unique putative pathogenic bacteria species identified within the fecal sample
- Pathogenic_spp_RRA: the total relative read abundance (RRA) of putative pathogenic bacteria species within the fecal sample
File: Pathogenic_microbe_spp_taxonomy.csv
Description: The taxonomic information for all the putative pathogenic bacteria species in 'Pathogenic_microbe_spp_RRA.csv'. Within this file, 'NA' refers to 'not available', indicating that the corresponding taxonomic information was not available for that putative pathogenic bacteria species, as it was not sufficiently taxonomically resolved.
Variables
- Kingdom: the kingdom to which the putative pathogenic bacteria species belongs (all 'Bacteria')
- Genus: the genus to which the putative pathogenic bacteria species belongs
- Species: the species to which the putative pathogenic bacteria species belongs
- Scientific_name: the scientific name (genus followed by species) of the putative pathogenic bacteria
File: Diet_OTU_taxonomy.csv
Description: The taxonomic information for all the plant mOTUs in 'Diet_OTU_RRA.csv'. Within this file, 'NA' refers to 'not available', indicating that the corresponding taxonomic information was not available for that plant mOTU, as it was not sufficiently taxonomically resolved.
Variables
- ID: the ID of the plant mOTU, corresponding to those in the published herbivore diet dataset (Kartzinel et al. 2019)
- Order: the plant order to which the plant mOTU belongs
- Family: the plant family to which the plant mOTU belongs
- Genus: the plant genus to which the plant mOTU belongs
- Species: the plant species to which the plant mOTU belongs
- Best Match(es): list of all potential species to which the plant mOTU is matched, where potential species are separated by semicolons
- Library: whether the best match for the mOTU was derived from the global plant reference library ('global') or local plant reference library ('local'; Gill et al. 2019)
- Sequence: the DNA sequence of the plant mOTU
File: Host_species_trait_data.csv
Description: Trait data for the host herbivore species from which fecal samples were derived. Within this file, ‘NA’ refers to ‘not available’, indicating that the corresponding information was not available for that particular species.
Variables
- Common_name: the common name of the herbivore species
- Scientific_name: the scientific name (genus followed by species) of the herbivore species
- Family: the mammalian family to which the herbivore species belongs
- Genus: the mammalian genus to which the herbivore species belongs
- Species: the mammalian species to which the herbivore species belongs
- Ruminant: whether the species is a ruminant or non-ruminant
- Domestic: whether the species is a domesticate or not
- IUCN2020_name: the scientific name per the IUCN 2020
- IUCN: the IUCN Red List threat level, such that 'LC' corresponds to 'Least Concern', 'NT' corresponds to 'Near Threatened', 'VU' corresponds to 'Vulnerable', 'EN' corresponds to 'Endangered', 'CR' corresponds to 'Critically Endangered', and 'DD' corresponds to 'Data Deficient'.
- Body_mass: the body mass of the herbivore species
- Adult_brain_mass: the adult brain mass of the herbivore species
- Adult_body_length: the adult body length of the herbivore species
- Max_longevity: the maximum lifespan of the herbivore species
- Gestation_length: the typical length of the gestation period of the herbivore species
- Teat_number_n: the typical number of teats possessed by the herbivores species
- Litter_size: the typical number of offspring per litter of the herbivore species
- Litters_per_year_n: the typical number of litters per year had by the herbivore species
- Group_size: the typical group size displayed by the herbivore species
File: Diet_sample_metadata.csv
Description: Sample metadata for the fecal samples in 'Diet_OTU_RRA.csv'. Within this file, ‘NA’ refers to ‘not available’, indicating that the corresponding information was not available for that particular sample.
Variables
- ID: the unique ID of the fecal sample, corresponding to those in the published herbivore diet dataset (Kartzinel et al. 2019)
- Species: the common name of the herbivore species from which the fecal sample was derived
- Order: the mammalian order of the herbivore species
- Family: the mammalian family of the herbivore species
- Latin Name: the scientific name (genus followed by species) of the herbivore species
- Digestive System: the digestive system of the herbivore species
- Domestic Species: the domestication status of the herbivore species
- Sample Date: the date the fecal sample was collected
- Sample Period: the period during which the fecal sample was collected
- Longitude: the longitude at which the fecal sample was collected
- Latitude: the latitude at which the fecal sample was collected
- Rain (mm/90 d prior): the amount of rain in the 90 days preceding the collection of the fecal sample
- Diet sequence-read depth: the read depth (total number of reads resulting from sequencing) of the sample for diet DNA metabarcoding
- Microbiome sequence-read depth: the read depth (total number of reads resulting from sequencing) of the sample for microbiome DNA metabarcoding
- In overall diet analysis: whether the sample was included in overall diet analysis for the original publication from which data are derived (Kartzinel et al. 2019)
- In overall microbiome analysis: whether the sample was included in overall microbiome analysis for the original publication from which data are derived (Kartzinel et al. 2019)
- In paired diet-microbiome comparisons: whether the sample was included in paired diet-microbiome analysis for the original publication from which data are derived (Kartzinel et al. 2019)
- Dietary richness: the number of unique diet mOTUs detected within the sample
- Microbiome richness: the number of unique bacterial ASVs detected within the sample
- Dietary diversity: the Shannon diversity of diet mOTUs detected within the sample
- Microbiome diversity: the Shannon diversity of bacterial ASVs detected within the sample
- Barcode confirmation: whether the species identity of the herbivore from which the sample was derived was confirmed genetically
- N777/H16498: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- UniMinibarF1/C1N1777: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- LCO1490/C1N1777: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- bushCOIF/C1N1777: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- bushCOIF/bushCO1R: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
File: Diet_OTU_RRA.csv
Description: Relative read abundances (RRA) of plant mOTUs within herbivore fecal samples. Each row is a unique plant mOTU, with corresponding taxonomic information for each mOTU provided in 'Diet_OTU_taxonomy.csv', and each column is a unique fecal sample, with corresponding metadata for each sample provided in 'Diet_sample_metadata.csv'. Values are relative read abundances of each mOTU within each sample. Within this file, ‘NA’ refers to ‘not available’, indicating that data on that mOTU was not available from that particular sample.
File: Microbiome_sample_metadata.csv
Description: Sample metadata for the fecal samples in 'Microbiome_OTU_RRA.csv' and 'Pathogenic_microbe_spp_RRA.csv'. Within this file, ‘NA’ refers to ‘not available’, indicating that the corresponding information was not available for that particular sample.
Variables
- ID: the unique ID of the fecal sample, corresponding to those in the published herbivore microbiome dataset (Kartzinel et al. 2019)
- Species: the common name of the herbivore species from which the fecal sample was derived
- Order: the mammalian order of the herbivore species
- Family: the mammalian family of the herbivore species
- Latin Name: the scientific name (genus followed by species) of the herbivore species
- Digestive System: the digestive system of the herbivore species
- Domestic Species: the domestication status of the herbivore species
- Sample Date: the date the fecal sample was collected
- Sample Period: the period during which the fecal sample was collected
- Longitude: the longitude at which the fecal sample was collected
- Latitude: the latitude at which the fecal sample was collected
- Rain (mm/90 d prior): the amount of rain in the 90 days preceding the collection of the fecal sample
- Diet sequence-read depth: the read depth (total number of reads resulting from sequencing) of the sample for diet DNA metabarcoding
- Microbiome sequence-read depth: the read depth (total number of reads resulting from sequencing) of the sample for microbiome DNA metabarcoding
- In overall diet analysis: whether the sample was included in overall diet analysis for the original publication from which data are derived (Kartzinel et al. 2019)
- In overall microbiome analysis: whether the sample was included in overall microbiome analysis for the original publication from which data are derived (Kartzinel et al. 2019)
- In paired diet-microbiome comparisons: whether the sample was included in paired diet-microbiome analysis for the original publication from which data are derived (Kartzinel et al. 2019)
- Dietary richness: the number of unique diet mOTUs detected within the sample
- Microbiome richness: the number of unique bacterial ASVs detected within the sample
- Dietary diversity: the Shannon diversity of diet mOTUs detected within the sample
- Microbiome diversity: the Shannon diversity of bacterial ASVs detected within the sample
- Barcode confirmation: whether the species identity of the herbivore from which the sample was derived was confirmed genetically
- N777/H16498: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- UniMinibarF1/C1N1777: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- LCO1490/C1N1777: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- bushCOIF/C1N1777: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
- bushCOIF/bushCO1R: the DNA sequence of the herbivore species resulting from this particular primer set used for genetic verification
File: Nemabiome_OTU_taxonomy.csv
Description: The taxonomic information for all the nematode mOTUs in 'Nematode_OTU_RRA.csv'. Within this file, 'NA' refers to 'not available', indicating that the corresponding information was not available for that nematode mOTU.
Variables
- Taxa: the name of the mOTU, composed of the taxonomic identity followed by the unique mOTU ID
- Boot: the bootstrap support for the taxonomic identity assigned to the sample
- seq_id: a unique identified for the DNA sequence of the nematode mOTU, corresponding to those in the published herbivore nemabiome dataset (Titcomb et al. 2022)
- mOTU: the unique ID of the nematode mOTU, corresponding to those in the published herbivore nemabiome dataset (Titcomb et al. 2022)
- Level: the taxonomic level to which the nematode mOTU is resolved
- Kingdom: the consensus kingdom to which the nematode mOTU belongs, per Titcomb et al. 2022
- Other: additional taxonomic information for the nematode mOTU, indicating that all mOTUs belong to the clade 'Metazoa'
- Phylum: the consensus phylum to which the nematode mOTU belongs, per Titcomb et al. 2022
- Class: the consensus class to which the nematode mOTU belongs, per Titcomb et al. 2022
- Order: the consensus order to which the nematode mOTU belongs, per Titcomb et al. 2022
- Family: the consensus family to which the nematode mOTU belongs, per Titcomb et al. 2022
- Genus: the consensus genus to which the nematode mOTU belongs, per Titcomb et al. 2022
- Species: the consensus species to which the nematode mOTU belongs, per Titcomb et al. 2022
- dada.Kingdom: the kingdom to which the nematode mOTU belongs, as indicated by DADA2 comparing against the nemabiome database
- dada.Other: other taxonomic information about the nematode mOTU, as indicated by DADA2 comparing against the nemabiome database
- dada.Phylum: the phylum to which the nematode mOTU belongs, as indicated by DADA2 comparing against the nemabiome database
- dada.Class: the class to which the nematode mOTU belongs, as indicated by DADA2 comparing against the nemabiome database
- dada.Order: the order to which the nematode mOTU belongs, as indicated by DADA2 comparing against the nemabiome database
- dada.Family: the family to which the nematode mOTU belongs, as indicated by DADA2 comparing against the nemabiome database
- dada.Genus: the genus to which the nematode mOTU belongs, as indicated by DADA2 comparing against the nemabiome database
- dada.Species: the species to which the nematode mOTU belongs, as indicated by DADA2 comparing against the nemabiome database
- boot.Kingdom: bootstrap support value for the kingdom assignment from DADA2
- boot.Other: bootstrap support value for the additional taxonomic information from DADA2
- boot.Phylum: bootstrap support value for the phylum assignment from DADA2
- boot.Class: bootstrap support value for the class assignment from DADA2
- boot.Order: bootstrap support value for the order assignment from DADA2
- boot.Family: bootstrap support value for the family assignment from DADA2
- boot.Genus: bootstrap support value for the genus assignment from DADA2
- boot.Species: bootstrap support value for the species assignment from DADA2
- sequence: the DNA sequence of the nematode mOTU
- BLAST_ID: the taxonomic identity of the mOTU resulting from BLAST
- BLAST_QueryCover: the coverage (in percent) of the DNA sequence by the corresponding BLAST sequence match
- BLAST_PercID: the level of correspondence (in percent) between the DNA sequence and the corresponding BLAST sequence match
- BLAST_potential_species: all equally well-matched species returned by BLAST
- BLAST_notes: notes from searching the sequence in BLAST
- AverageAdultSize: length range(s) the nematode species to which the DNA sequence was matched using BLAST
- Size Citation: citations for the sources from which size estimates were derived
- Size_Class: the size class (small or large) of the nematode mOTU
- Size_Average: the average of all body length estimates associated with the nematode mOTU
- Feeding habit: the feeding habit (breach or blood) of the nematode mOTU
- Habit citation: citations for the sources from which feeding habit classifications were derived
File: Nemabiome_OTU_RRA.csv
Description: Relative read abundances (RRA) of nematode mOTUs within herbivore fecal samples. Each column is a unique nematode mOTU, with corresponding taxonomic information for each mOTU provided in 'Nemabiome_OTU_taxonomy.csv', and each row is a unique fecal sample, with corresponding metadata for each sample provided in 'Nemabiome_sample_metadata.csv'. Values are relative read abundances of each mOTU within each sample. Within this file, ‘NA’ refers to ‘not available’, indicating that data on that mOTU was not available from that particular sample.
File: Microbiome_ASV_taxonomy.csv
Description: The taxonomic information for all the bacterial ASVs in 'Microbiome_ASV_RRA.csv'. Within this file, 'NA' refers to 'not available', indicating that the corresponding taxonomic information was not available for that bacterial ASV, as it was not sufficiently taxonomically resolved.
Variables
- ID: the unique ID of the bacterial ASV, corresponding to those in the published herbivore diet dataset (Kartzinel et al. 2019)
- Phylum: the phylum to which the bacterial ASV belongs
- Class: the class to which the bacterial ASV belongs
- Order: the bacterial order to which the bacterial ASV belongs
- Family: the bacterial family to which the bacterial ASV belongs
- Genus: the bacterial genus to which the bacterial ASV belongs
- Species: the bacterial species to which the bacterial ASV belongs
- Sequence: the DNA sequence of the bacterial ASV
File: Microbiome_ASV_RRA.csv
Description: Relative read abundances (RRA) of bacterial ASVs within herbivore fecal samples. Each row is a unique bacterial ASV, with corresponding taxonomic information for each ASV provided in 'Microbiome_ASV_taxonomy.csv', and each column is a unique fecal sample, with corresponding metadata for each sample provided in 'Microbiome_sample_metadata.csv'. Values are relative read abundances of each ASV within each sample. Within this file, ‘NA’ refers to ‘not available’, indicating that data on that ASV was not available from that particular sample.
File: Pathogenic_microbe_spp_RRA.csv
Description: Relative read abundances (RRA) of putative pathogenic bacterial species within herbivore fecal samples. Each column is a unique putative pathogenic bacterial species, with corresponding taxonomic information for each putative pathogenic bacterial species provided in 'Pathogenic_microbe_spp_taxonomy.csv', and each row is a unique fecal sample, with corresponding metadata for each sample provided in 'Microbiome_sample_metadata.csv'. Values are relative read abundances of each putative pathogenic bacterial species within each sample. Within this file, ‘NA’ refers to ‘not available’, indicating that data on that species was not available from that particular sample.
Code/software
Analyses were conducted in R using publicly available R packages, as cited in the text of the manuscript associated with this dataset. The R code to reproduce the analyses in the associated manuscript is included in this data repository.
Access information
Other publicly accessible locations of the data:
- Diet and microbiome data are available in one Dryad repository: https://doi.org/10.5061/dryad.c119gm5
- Nemabiome data are available in a second Dryad repository: https://doi.org/10.25349/D96P6K
Data was derived from the following sources:
- Diet and microbiome data were derived from: Kartzinel TR, Hsing JC, Musili PM, Brown BRP, Pringle RM. 2019 Covariation of diet and gut microbiome in African megafauna. PNAS 116, 23588–23593. (doi:10.1073/pnas.1905666116)
- Nemabiome data were derived from: Titcomb GC, Pansu J, Hutchinson MC, Tombak KJ, Hansen CB, Baker CCM, Kartzinel TR, Young HS, Pringle RM. 2022 Large-herbivore nemabiomes: patterns of parasite diversity and sharing. Proceedings of the Royal Society B: Biological Sciences 289, 20212702. (doi:10.1098/rspb.2021.2702)
