Data from: Limitations of common molecular markers in fungal biodiversity analysis and the benefits of their synergistic use
Data files
Apr 28, 2026 version files 3.89 GB
-
big_mock_processed.zip
15.52 MB
-
info_processed.txt
115 B
-
info_raw.csv
75.27 KB
-
mock_raw.zip
3.86 GB
-
mock_references.zip
147.79 KB
-
README.md
7.34 KB
-
subsets_mock_preprocessed.zip
11.50 MB
Abstract
High-throughput sequencing of the Internal Transcribed Spacer (ITS) regions is the primary method for estimating fungal diversity from environmental DNA. However, reliance solely on ITS markers is complicated by its high variability in sequence length and the presence of multiple variants within a single genome, which can bias diversity estimates. The core objective of the dataset is to provide a comparative analysis of the ITS regions (ITS1 and ITS2) against five alternative markers: Mitochondrial Large Subunit ribosomal DNA (mtLSU), the second largest subunit of RNA polymerase II (rpb2a and rpb2b regions), the translation elongation factor 1-alpha (ef1-α), and the minichromosome maintenance complex component 7 (mcm7). The data comprises raw Illumina MiSeq amplicon sequences (FASTQ format) derived from a mock community of 413 fungal species belonging to Ascomycota, Basidiomycota, Mucoromycota, and Zoopagomycota. It also includes manually curated reference sequences (FASTA format) and supplementary tabular data mapping primers and tags to samples. These data have high reuse potential for bioinformatics benchmarking and for testing or validating fungal barcode identification pipelines. These data are provided as open-access resources and are fully open for reusage.
This dataset supports the study titled "Limitations of common molecular markers in fungal biodiversity analysis and the benefits of their synergistic use", which evaluates the performance of seven molecular markers for fungal metabarcoding (https://doi.org/10.1111/1755-0998.70123).
The mock community consisted of an equimolar mixture of genomic DNA extracted from 676 fungal specimens obtained either from pure cultures (613) or field collections (63). All these fungi belonged to the group Eumycota: 103 Ascomycota strains, 554 Basidiomycota strains, 18 Mucoromycota strains and 1 Zoopagomycota strain, corresponding to 413 different species, 202 genera, 96 families, 32 orders and 11 classes. Altogether, 71 genera were represented by several species, and 125 species were represented by more than one specimen.
All fungal specimens were identified based on morphology and full-length ITS region sequences. We received biological materials from the Culture Collection of Fungi, Department of Botany, Faculty of Science of Charles University in Prague; the Culture Collection of Basidiomycetes, Institute of Microbiology CAS; collections of the Laboratory of Fungal Genetics and Metabolism, Institute of Microbiology CAS; collections from the Department of Forest Protection and Wildlife Management at the Faculty of Forestry and Wood Technology, Mendel University in Brno; collections from the Department of Environmental Geology and Geochemistry, Institute of Geology CAS; and from the herbarium of the National Museum in Prague. The origin of each fungal strain is documented in Supplementary Table S1 in the manuscript.
File describtions
We have submitted raw amplicon data (mock_raw.zip), curated reference sequences (mock_references.zip), processed datasets (big_mock_processed.zip, subsets_mock_preprocessed.zip), bioinformatic processing details (info_processed.txt), and a comprehensive mapping file (info_raw.csv).
1. info_raw.csv
Variable Metadata for info_raw.csv
The following table describes the variables included in the primary metadata mapping file:
| Fieldname | Field content and format | Values from vocabulary | Note |
|---|---|---|---|
| Sample code | Unique identifier of the sample | YES (bim41, SM1-63, SMBLANK) |
Internal laboratory code for the mock community or subset. |
| Description | Categorization of the sample type | YES (big_mock, ref_subset, ref_blank) |
Identifies if the sample is the primary community, a reference subset, or a negative control. |
| Adapter Sequence | Oligonucleotide sequences | NO | Illumina adapters used in library preparation kits. |
| Gene code | Barcoding region targeted by sequencing | YES (ITS1, ITS2, EF-1α, MCM7, mtLSU, RPB2_A, RPB2_B) |
Specifies the molecular marker amplified in the corresponding row. |
| Full code | Concatenated unique identifier | NO | Combination of Sample code and Gene code (e.g., bim41_its1). |
| Forward Sequence | Oligonucleotide sequences | NO | Forward PCR primer without adapters or tags. |
| fwd TAG+2nt of primer | Oligonucleotide sequences | NO | Forward primers with sample-specific tags used for sample demultiplexing |
| Adapter+Forward | Oligonucleotide sequences | NO | The full sequence string of the adapter and the tagged forward primer. |
| Reverse Sequence | Oligonucleotide sequences | NO | Reverse PCR primer sequence without adpters or tags. |
| Adapter+Reverse | Oligonucleotide sequences | NO | The full sequence string of the adapter and the reverse primer. |
2. info_processed.txt
Briefly describes the bioinformatic processing steps applied in SEED2 pipeline to generate the "processed" datasets, including filtering reads with PHRED quality score (Q-score) mean below 30 (Q30), ambiguous bases or a mismatch in the tag (noAMB), trimming adapters and tags (noTAG), filtering ITS1 and ITS2 reads shorter than 30 bp, and mtLSU, ef1-α, rpb2a, rpb2b, mcm7 reads shorter than 200 bp (min sizes), filtering putative chimeric reads using de novo strategy in the UCHIME (noCHIM), truncating ITS1 and ITS2 reads to contain only the highly variable region using the utility ITSx, and filtering reads unassigned to the proper marker and non-fungal sequences, i.e. those with a best hit in the NCBI database to the non-fungal taxa (Fungi+Gene-verified).
3. big_mock_processed.zip
Contains quality-filtered amplicon sequences for the mock community.
4. mock_raw.zip
The complete set of raw, unprocessed Illumina MiSeq FASTQ files for the mock community.
5. mock_references.zip
A curated collection of FASTA reference sequences for all seven markers, generated by Illumina MiSeq sequencing of 64 small subsets of fungal specimens included in the mock community.
6. subsets_mock_preprocessed.zip
Contains quality-filtered amplicon sequences for the 64 subsets (10–11 species each) used to build the reference datasets.
Code/Software
We used SEED 2 pipeline (https://www.biomed.cas.cz/mbu/lbwrf/seed/) for all bioinformatics processing. R (version 4.2.0) was used for statistical analysis and visualization, specifically utilizing the following packages:
- dplyr (v1.1.4): For data manipulation.
- rstatix (v0.7.2): For Spearman's rank correlation.
- ggplot2 (v3.4.4): For bar plot visualization.
- UpSetR (v1.4.0): For matrix layout visualization of detected species.
Statistical scripts and bioinformatic parameters are available on GitHub: http://www.github.com/vassishap/fungal_multimarker_metabarcoding.
Benefit-sharing statement
Specimens originating from the Czech Republic were collected in accordance with national legislation. For specimens originating from other countries, we have exercised due to diligence to confirm that these genetic resources were acquired in full compliance with the Nagoya Protocol and the national legislation of the provider countries.
These data are shared on a non-monetary basis and are open for reusage.
