Senegalese Bulinus truncatus and environmental samples bacterial microbiota
Data files
Jun 12, 2026 version files 16.85 GB
-
controls.tar
139.44 MB
-
raw_bulinus_truncatus.tar
7.91 GB
-
raw_edna.tar
8.77 GB
-
README.md
5.15 KB
-
Suppl_table_2.csv
3.14 KB
-
suppl_table_4.csv
23.49 MB
-
suppl_table_5.csv
12.27 KB
-
suppl_table_6.csv
10.49 KB
Jun 22, 2026 version files 16.85 GB
-
controls.tar
139.44 MB
-
raw_bulinus_truncatus.tar
7.91 GB
-
raw_edna.tar
8.77 GB
-
README.md
6.47 KB
-
Suppl_table_2.csv
3.14 KB
-
suppl_table_4.csv
23.49 MB
-
suppl_table_5.csv
12.27 KB
-
suppl_table_6.csv
10.49 KB
-
Suppl._table_2-_Microsatellite_primers_sequences_W1_W2.xlsx
24.48 KB
Abstract
Microbiota have emerged as fundamental regulators of host physiology, shaping both ecological interactions and evolutionary trajectories. Yet, the determinants of microbiota diversity and structure in wild populations—particularly the respective roles of host genetics and environmental context—are still poorly understood. In this study, we investigated these influences in the freshwater snail Bulinus truncatus, a key intermediate host for human and animal Schistosoma parasites, using a multifactorial approach. We developed 31 new microsatellite markers to resolve population genetic structure across nine sites in Senegal. Metabarcoding methods were then employed to profile the bacterial microbiota of individual snails and to characterize environmental bacterial assemblages from each location via environmental DNA. Shell measurements and molecular diagnostics for trematode infection status were included to assess additional potential contributors. Employing multiple regression on distance matrices (MRM), we quantified how snail population genetics, site-specific environmental bacterial communities, spatial patterns, and infection status shape microbiota composition. Our analyses reveal that snail geographic distribution and population genetic structure drive the composition of Bulinus truncatus microbiota, with environmental bacterial communities exerting a weaker but still significant effect. In contrast, neither shell size nor trematode infection status impacted microbiota structure significantly. Notably, a considerable fraction of variation remains unexplained, indicating the likely involvement of other ecological or intrinsic factors. These results advance understanding of microbiota determinants in natural populations and underscore the intricate interplay between host genetics, environment, and microbial communities.
This dataset comprises fastq files from 16S v3-v4 rDNA metabarcoding data of Bulinus truncatus and environmental DNA samples from Senegal. There is one TAR archive containing the fastq files of eDNA samples (raw_edna.tar) and another containing the fastq files of B. truncatus individual snails (raw_bulinus_truncatus.tar).
https://doi.org/10.5061/dryad.v6wwpzh5d
Description of the data and file structure
This OTUs table and fastq sequences underlie the main results of the study "Evidences that host genetic background more than the environment shapes the microbiota of the snail Bulinus truncatus, an intermediate host of Schistosoma species.".
In this study, our aim was to characterise and compare the bacterial communities found in B. truncatus snails and in the aquatic environment at 9 ecologically contrasted. The field work was conducted in February 2022 and we focused on 9 natural sites from Northern Senegal that differ in terms of habitats. At each of these sites, we sampled the water-sediment interface and all B. truncatus snails found. We also filtered comercial spring water at each site to have a field negative control. We caracterised the microbiota using v3-v4 16s rDNA metabarcoding a MiSeq amplicons sequencing approach.
Those approaches led to the processing of 198 bacterial 16S metabarcoding libraries including the DNA extracts of the 124 snails, 70 eDNA replicates consisting in 54 eDNA samples (each extraction triplicate was amplified in duplicate) and 16 technical field negative controls (each control was amplified in duplicate), and 4 negative PCR controls (milliρ water).
Files and variables
File: raw_edna.tar
Description: fastq sequences coming from the 16s rDNA sequencing of environmental samples (54 samples + 16 field negative controls)
File: raw_bulinus_truncatus.tar
Description: fastq sequences coming from the 16s rDNA sequencing of Bulinus truncatus snails
File: controls.tar
Description: fastq sequences coming from the 16s rDNA sequencing of the PCR negative controls (water)
Naming convention for FASTQ files
The FASTQ files follow an extended Illumina naming structure that encodes sample type, site, replicate information, and sequencing metadata.
General structure:
[SampleCode]_[InternalID]_L001_R[ReadNumber]_001.fastq.gz
Where:
- SampleCode = descriptive identifier indicating sample type and origin
S1-1= snail sample 1 from site 1C-S1-1= field eDNA control 1 from site 1S1-1-1= eDNA sample 1, subsample 1 from site 1neg-9D-pl3= PCR negative control (plate 3)
- InternalID = sequencing facility internal sample ID (e.g.,
S1,S101,S103,S194) - L001 = sequencing lane
- R1 / R2 = forward and reverse reads
- 001 = file index (always 001)
Examples
S1-1_S1_L001_R1_001.fastq.gz→ snail sample 1 from site 1, forward readC-S1-1_S101_L001_R1_001.fastq.gz→ field eDNA control 1 from site 1, forward readS1-1-1_S103_L001_R1_001.fastq.gz→ eDNA sample 1, subsample 1 from site 1, forward readneg-9D-pl3_S194_L001_R1_001.fastq.gz→ PCR negative control (plate 3), forward read
All reverse reads follow the same structure with R2.
All FASTQ files contained in the .tar archives follow this naming convention.
| site | longitude | latitude | habitat | site_name |
|---|---|---|---|---|
| S1 | -14.84952 | 16.58457 | river_2 | Ndiawara |
| S2 | -14.93601 | 16.59875 | river_2 | Ouali Diala |
| S3 | -14.92521 | 16.59762 | canal | Guia |
| S4 | -14.88948 | 16.59714 | river_2 | Dioundou |
| S5 | -14.96063 | 16.60565 | river_2 | Fonde Ass |
| S6 | -14.94503 | 16.59575 | river_2 | Khodit |
| S7 | -15.80198 | 16.27096 | lake | Mbane |
| S8 | -15.80171 | 16.24226 | lake | Saneinte |
| S9 | -16.34956 | 16.10936 | river_1 | Lampsar |
File: Suppl_table_2.csv
Description: file containing the individual snail shell size (length and width).
Variables:
sample_id: unique snail identifiershell_length: shell length (mm)shell_width: shell width (mm)
Delimiter: ;
Decimal notation: ,
File: suppl_table_4.csv
Description: asv table containing all samples, prior to any cleaning step.
Variables:
asv: ASV identifiersample_id: sample namesample_type: sample type (eDNA control, negative control, eDNA, snail)count: read count
Delimiter: ;
File: suppl_table_5.csv
Description: table showing all ASVs present in the negative controls, with taxonomy details.
Variables:
asv_idsample_idcounttaxonomy
Delimiter: ;
File: suppl_table_6.csv
Description: table showing the read counts per sample during cleaning steps.
Variables:
sample_idRead count after preprocessingRead count after removing mitochondria readsRead count after removing chloroplast readsRead count after removing kingdom = NA readsRead count after removing phylum = NA readsRead count after global abundance filter (0.005%)Read count after individual abundance filter (0.1%)Read counts after rarefaction to 5000 reads
Delimiter: ;
File: Suppl._table_2-_Microsatellite_primers_sequences_W1_W2.xlsx
Description: Microsatellite loci used in the study, including locus name, motif, number of repeats, amplicon size, primer sequences, sequencing coverage, and genotyping metrics.
Column description
- Multiplex — multiplex panel used for amplification (e.g., W1, W2).
- Locus name — unique identifier of the microsatellite locus.
- ReadID+ — reference sequence ID used for locus mapping.
- Microsatellite motif — repeated motif (e.g., ATAG).
- Number of repeats — number of motif repetitions.
- Amplicon size — expected PCR product length (bp).
- Forward primer sequence* — forward primer (5’→3’).
- Reverse primer sequence — reverse primer (5’→3’).
- Mean coverage/loci/sample — average sequencing depth per locus per sample.
- SD coverage/loci/sample — standard deviation of sequencing depth.
- Analysis strategy — genotyping strategy (e.g., FullLength).
- Parameter set* — parameter set used for genotyping.
- NallelesSequence — number of alleles detected by sequence.
- NalleleSize — number of alleles detected by size.
- Raw missing genotype $ — proportion of missing genotypes.
- Allelic error — allelic error rate.
The samples were collected during February 2022 in Senegal. At each site, we first filtered water along the water column from the surface to the water-sediment interface. We used filtration capsules of 0.45 µM mesh size (Waterra USA Inc.) connected to an electric water pump as previously described in (Douchet et al. 2022). Filtration was performed until the filtration capsule was clogged and the final filtration volume was retrieved. Once the filtration completed, the filtration capsule was depleted from its water content, filled with 50 mL of Longmire solution, vigorously shaken, and preserved at ambient temperature until used for environmental DNA (eDNA) extractions. Following eDNA sampling, Bulinus truncatus were sampled at each site. Total eDNA from water filtrations were extracted following (Douchet et al. 2022) using DNeasy PowerSoil Pro kit (Qiagen) according to manufacturer’s protocol. Total genomic DNA from each of the snails removed from their shell was extracted using DNeasy 96 Blood and Tissue Kit (Qiagen), according to the manufacturer’s protocol.
We targeted the variable V3-V4 loops region of the 16S sDNA gene using the 341F (5’-CCTACGGGNGGCWGCAG-3’) and 805R (5’-GACTACHVGGGTATCTAATCC-3’) primers (Klindworth et al. 2013) combined with universal Illumina adapters. The first PCRs were performed using the Q5® High-Fidelity 2X Master Mix (New England BioLabs), in a 25µL final volume containing 2 µL of template DNA and using a PCR program consisting in an initial denaturation step of 30s at 98°C followed by 32 cycles containing a denaturation step of 6 sec at 98°C, an annealing step of 30 sec at 55°C, and an elongation step of 8 sec at 72°C and ending with a final elongation step of 60 sec at 65°C. We used 5 µL of the PCR products to check the PCR products quality and integrity through electrophoresis on agarose gels. The 20µL left were sent to the Bio-Environment platform (University of Perpignan Via Domitia, France). Libraries were performed using Illumina Nextera index kit and Q5 high fidelity DNA polymerase (New England Biolabs). Indexed PCR products were then normalized with SequalPrep plates (ThermoFisher) and paired-end sequenced on a MiSeq instrument using v2 chemistry (2 x 250 bp).
Changes after Jun 12, 2026:
Uploading of microsatellite primer sequences that were missing from the initial submission (Suppl._table_2-_Microsatellite_primers_sequences_W1_W2.xlsx).
