Integrated reanalysis of global riverine fish eDNA datasets
Data files
Apr 21, 2026 version files 157.91 GB
-
1_Info.zip
12.48 KB
-
3_Data.zip
2.54 MB
-
README.md
25.95 KB
-
RS001.zip
161.09 MB
-
RS002.zip
98.62 MB
-
RS003.zip
998.56 MB
-
RS004.zip
43.90 MB
-
RS005.zip
550.50 MB
-
RS006.zip
540.58 MB
-
RS007.zip
474.76 MB
-
RS008.zip
8.50 GB
-
RS009.zip
84.36 MB
-
RS010.zip
2.19 GB
-
RS011.zip
182.29 MB
-
RS012.zip
28.91 MB
-
RS013.zip
300.38 MB
-
RS014.zip
1.07 GB
-
RS016.zip
214.38 MB
-
RS017.zip
6.92 GB
-
RS018.zip
361.30 MB
-
RS019.zip
3.68 GB
-
RS020.zip
1.02 GB
-
RS021.zip
54.93 GB
-
RS022.zip
44.66 GB
-
RS023.zip
35.72 MB
-
RS024.zip
33.48 MB
-
RS025.zip
154.74 MB
-
RS026.zip
301.30 MB
-
RS027.zip
63.22 MB
-
RS028.zip
1.01 GB
-
RS029.zip
26.25 MB
-
RS030.zip
998.56 MB
-
RS031.zip
317.09 MB
-
RS032.zip
1.03 GB
-
RS033.zip
25.86 MB
-
RS034.zip
8.96 GB
-
RS035.zip
366.49 MB
-
RS036.zip
1.03 GB
-
RS037.zip
8.96 GB
-
RS038.zip
78.95 MB
-
RS039.zip
34.95 KB
-
RS040.zip
44.35 MB
-
RS041.zip
25.86 MB
-
RS042.zip
1.27 GB
-
RS043.zip
25.86 MB
-
RS044.zip
25.86 MB
-
RS045.zip
25.86 MB
-
RS046.zip
2.29 GB
-
RS047.zip
398.53 MB
-
RS048.zip
887.60 MB
-
RS049.zip
331.05 MB
-
RS050.zip
289.94 MB
-
RS051.zip
86.27 MB
-
RS052.zip
118.79 MB
-
RS053.zip
296.24 MB
-
RS054.zip
274.22 MB
-
RS055.zip
324.61 MB
-
RS056.zip
442.17 MB
-
RS057.zip
360.52 MB
Abstract
Environmental DNA (eDNA) has revolutionized biodiversity monitoring in aquatic ecosystems. One of its possible hallmarks is the potential global comparability of datasets, yet this opportunity has been largely left unexplored due to diverse sampling protocols and bioinformatic workflows. This dataset compiles and harmonizes riverine fish community data derived from environmental DNA (eDNA) metabarcoding studies conducted worldwide. In total, this dataset comprises 58 riverine fish eDNA metabarcoding datasets from 1,818 sampling sites, comprising both published and novel datasets with the goal of assessing and comparing biodiversity patterns from the individual original studies with a unified analysis under a common bioinformatic pipeline. We thereby demonstrate the power of integrating individual eDNA metabarcoding datasets for biodiversity assessment and monitoring.
Project leaders:
- Florian Altermatt (University of Zurich, Switzerland)
- Xiaowei Zhang (Nanjing University, China)
- Loïc Pellissier (ETH Zurich, Switzerland)
- Yan Zhang (Nanjing University, China)
README version date:
20 April 2026
1. General description of the archive
This archive contains metadata, raw sequencing datasets, reformatted eDNA datasets, and unified downstream analysis outputs used in the project “Integrated reanalysis of global riverine fish eDNA datasets”.
The archive includes:
- archive-level metadata and dataset crosswalk files
- raw sequencing datasets
- standardized reformatted datasets
- outputs from unified downstream analyses of raw sequencing data
2. Archive contents
This archive includes the following main components:
1_Info.zip
Archive-level metadata and dataset crosswalk files.
Contents:
AllDataTableCheckList.csv: metadata checklist for all reformatted datasets, identified byDataID(D0XX)00PrimerList.csv: metadata for all raw sequencing datasets archived asRS0XX00PrimerList_db_use.csv: parameters used to generate the barcode reference database for all primer sets
Raw sequencing datasets
The following compressed folders contain raw sequencing data for individual sequencing datasets:
RS001.zipRS002.zipRS003.zipRS004.zipRS005.zipRS006.zipRS007.zipRS008.zipRS009.zipRS010.zipRS011.zipRS012.zipRS013.zipRS014.zipRS016.zipRS017.zipRS018.zipRS019.zipRS020.zipRS021.zipRS022.zipRS023.zipRS024.zipRS025.zipRS026.zipRS027.zipRS028.zipRS029.zipRS030.zipRS031.zipRS032.zipRS033.zipRS034.zipRS035.zipRS036.zipRS037.zipRS038.zipRS039.zipRS040.zipRS041.zipRS042.zipRS043.zipRS044.zipRS045.zipRS046.zipRS047.zipRS048.zipRS049.zipRS050.zipRS051.zipRS052.zipRS053.zipRS054.zipRS055.zipRS056.zipRS057.zip
Each RS###.zip file corresponds to one raw sequencing dataset and contains:
- raw sequencing data files in FASTQ format
metadata.csv, which provides sample-level metadata for the sequencing samples
Please refer to 1_Info/00PrimerList.csv for metadata associated with each raw sequencing dataset.
3_Data.zip
Processed eDNA data products and outputs from unified analyses.
This folder contains the following subfolders:
3_Data/1_Reformat
Standardized and renumbered datasets.
This folder contains multiple dataset folders (D0XX; 62 folders in total). Each dataset folder corresponds to one reformatted dataset identified by DataID in AllDataTableCheckList.csv.
Each D0XX folder contains:
SiteSpeciesMatrix_cor.csv: species-by-sample matrix with standardized/corrected species namesSiteMetadata.csv: sample-level metadata describing the samples included in the paired species matrix
The file descriptions and variable definitions given below for SiteSpeciesMatrix_cor.csv and SiteMetadata.csv apply to all D0XX folders in 3_Data/1_Reformat.
3_Data/2_Reclassify
Outputs from unified analysis of raw sequencing data.
Files include:
Format_SeperateDataset_ASVs_Stan.RDS: ASV and taxonomic tables for each raw sequencing dataset after standardized processingDataset_Fish_Recored_C.RDS: historical fish species records for each studied river basinFormat_SeperateDataset_Spetb_LocalAssign_Global.RDS: species-by-site table without correction based on local historical species recordsFormat_SeperateDataset_Spetb_LocalAssign_Ecoregion_Country_Basin.RDS: corrected species-by-site tables based on local historical species records at the ecoregion, country, or basin level
3. Relationship between files
RawSeqIDidentifies raw sequencing datasets and corresponds toRS0XXDataIDidentifies reformatted datasets and corresponds toD0XX00PrimerList.csvprovides metadata linkingRawSeqID,DataID, primer information, and major raw-sequencing processing steps00PrimerList_db_use.csvprovides primer-specific parameters used to generate the barcode reference databaseAllDataTableCheckList.csvprovides cross-dataset metadata for the reformatted datasets in3_Data/1_ReformatRS0XX/metadata.csvprovides sequencing-sample metadata for the raw sequencing datasets3_Data/1_Reformat/D0XX/SiteMetadata.csvprovides sample-level metadata for the reformatted datasets3_Data/1_Reformat/D0XX/SiteSpeciesMatrix_cor.csvcontains the paired species-by-sample matrix for the same reformatted datasets
4. Missing data and abbreviations
Missing data
Unless otherwise specified, NA indicates one of the following:
- The information was not available or not reported in the original source; or
- The field was not applicable to that dataset, sample, or record.
Abbreviations
ASV: amplicon sequence variantOTU: operational taxonomic uniteDNA: environmental DNASE: single-end sequencingDE: paired-end sequencing
5. Data dictionary
5.1 File: 1_Info/00PrimerList.csv
Purpose:
Metadata for all raw sequencing datasets archived as RS0XX.
Applies to:
All raw sequencing datasets are included as RS0XX.zip.
| Column | Definition | Unit / format | Notes |
|---|---|---|---|
RawSeqID |
Unique identifier for the raw sequencing dataset | text; format RS0XX |
Links this table to the corresponding raw sequencing archive |
N_PN |
Internal short code for the primer set used in the raw sequencing dataset | text | Primer-related shorthand code used in the archive |
DataID |
Unique identifier for the corresponding reformatted dataset | text; format D0XX |
Links the raw sequencing dataset to the reformatted dataset |
PrimerName |
Full name of the primer set used | text | May repeat across multiple raw sequencing datasets |
PrimerF |
Forward primer sequence | DNA sequence text | Primer sequence as used for the dataset |
PrimerR |
Reverse primer sequence | DNA sequence text | Primer sequence as used for the dataset |
Target region |
Target genetic marker region amplified by the primer set | text | For example, 12S, 16S, or COI |
Length of product (bp) |
Expected amplicon length | base pairs (bp) |
Numeric value |
SeqType |
Sequencing read type | text | For example, SE for single-end and DE for paired-end sequencing |
Demultiplex |
Demultiplexing status | text | Indicates whether and/or how reads were separated by sample/barcode |
Trim&Filter |
Trimming and filtering status | text | Indicates whether and/or how trimming and filtering were performed |
Assignment |
Taxonomic assignment status | text | Indicates whether and/or how taxonomic assignment was performed |
Note |
Additional notes on the dataset | free text | May include comments on processing, quality, or special cases |
5.2 File: 1_Info/00PrimerList_db_use.csv
Purpose:
Parameters used to generate the barcode reference database for each primer set.
Applies to:
Primer-level reference-database construction information.
| Column | Definition | Unit / format | Notes |
|---|---|---|---|
PrimerName |
Name of the primer set | text | Links this file to primer information in other tables |
Forward |
Forward primer sequence | DNA sequence text | Sequence used in reference-database generation |
Reverse |
Reverse primer sequence | DNA sequence text | Sequence used in reference-database generation |
Region |
Target marker region for the primer set | text | For example, 12S, 16S, or COI |
Len_min |
Minimum sequence length retained or targeted in the database-generation workflow | base pairs (bp) |
Primer-specific parameter |
Len_max |
Maximum sequence length retained or targeted in the database-generation workflow | base pairs (bp) |
Primer-specific parameter |
error |
Error or mismatch parameter used in the database-generation workflow | numeric | Primer-specific processing parameter |
min_value |
Minimum threshold value used in the database-generation workflow | numeric | Primer-specific processing parameter |
nData |
Number of datasets associated with the primer set | count | Integer value |
Status |
Processing status of reference-database generation for the primer set | text | For example, completion status |
5.3 File: 1_Info/AllDataTableCheckList.csv
Purpose:
Cross-dataset metadata checklist for all reformatted datasets identified by DataID (D0XX).
Applies to:
All reformatted datasets are included in 3_Data/1_Reformat.
| Section | Column | Definition | Unit / format | Notes |
|---|---|---|---|---|
| Reference info. | DataID |
Unique dataset identifier for the reformatted dataset | text; format D0XX |
Dataset ID |
| Reference info. | RawSeqID |
Raw sequencing dataset identifier linked to the dataset | text; format RS0XX |
Raw sequencing dataset ID |
| Reference info. | Author |
Author list associated with the dataset or study | text | Author names as reported by the source |
| Reference info. | Publication (if published) |
Publication or DOI information, if available | text | Citation information when available |
| Process | SiteMetadata |
Status of formatting for site/sample metadata | text | Indicates whether formatted metadata are available |
| Process | SiteSpeMatrix |
Status of formatting for the species matrix | text | Indicates whether formatted species matrix is available |
| Process | RawData Reanalysed |
Whether raw sequencing data were reanalysed | text | Indicates whether raw reads were reprocessed in this project |
| Studied River info. | RiverName |
Name of the studied river | text | River name associated with the dataset |
| Studied River info. | EcosystemType (river/stream) |
Ecosystem type of the study system | text | For example, river or stream |
| Studied River info. | Continent |
Continent where the study was conducted | text | Geographic region |
| Studied River info. | Country |
Country where the study was conducted | text | Geographic location |
| Sampling info. | Number of sampling events |
Number of distinct sampling events in the dataset | count | Integer value |
| Sampling info. | Sampling year |
Year or years in which sampling was conducted | year or text | May be a single year or a range/list |
| Sampling info. | Mesh size (um) |
Mesh size of the filter or sampling material | micrometers (um) |
Numeric value where reported |
| Sampling info. | Volume extracted (liters) |
Water volume extracted or filtered, where reported | liters (L) |
Numeric value where reported |
| Sampling info. | Filter replicates per site |
Number of filter replicates collected per site | count | Integer value where reported |
| Sampling info. | Number of sites |
Number of sampling sites in the dataset | count | Integer value |
| Metabarcoding experiment | Extraction kit |
DNA extraction kit or extraction method used | text | Free-text description |
| Metabarcoding experiment | Primer name |
Name of the primer set used in metabarcoding | text | Amplification primer name |
| Metabarcoding experiment | Primer F |
Forward primer sequence used in metabarcoding | DNA sequence text | Sequence as reported or standardized |
| Metabarcoding experiment | Primer R |
Reverse primer sequence used in metabarcoding | DNA sequence text | Sequence as reported or standardized |
| Metabarcoding experiment | Target region |
Target marker region amplified in metabarcoding | text | For example, 12S, 16S, or COI |
| Metabarcoding experiment | Length of product (bp) |
Expected amplicon length | base pairs (bp) |
Numeric value |
| Metabarcoding experiment | Sequencing platform |
Sequencing platform used for the dataset | text | Platform name as reported |
| Metabarcoding experiment | Sequencing depth (k reads) |
Approximate sequencing depth | thousand reads (k reads) |
Numeric value where reported |
| Bioinformatics | Merge pairs method |
Method used to merge paired reads | text | Free-text description of the pipeline step |
| Bioinformatics | Trimming method |
Method used for read trimming and filtering | text | Free-text description |
| Bioinformatics | Quality score |
Quality score threshold or criterion used in processing | numeric or text | As reported in the source or workflow |
| Bioinformatics | ASV/OTUs |
Type of sequence unit used in the analysis | text | Indicates whether the dataset uses ASVs or OTUs |
| Bioinformatics | Clustering algorithm |
Algorithm used to cluster reads or define OTUs/ASVs | text | Free-text description |
| Bioinformatics | Reference database |
Reference database used for taxonomic assignment | text | Database name or source |
| Bioinformatics | Assignment method |
Method used for taxonomic assignment | text | Free-text description |
| Bioinformatics | Assignment threshold |
Threshold used for taxonomic assignment | numeric or text | For example, similarity or confidence threshold |
| Bioinformatics | Abundance cutoff |
Minimum abundance threshold applied during filtering | numeric or text | As reported in the source or workflow |
| Bioinformatics | Occurrence cutoff |
Minimum occurrence threshold applied during filtering | numeric or text | As reported in the source or workflow |
| Bioinformatics | If rarefied |
Whether rarefaction was applied | text | Indicates whether reads were rarefied before downstream analysis |
5.4 File: RS0XX/metadata.csv
Purpose:
Sample-level metadata for raw sequencing samples contained in the corresponding RS0XX raw sequencing dataset.
Applies to:
All RS0XX folders containing raw sequencing data.
| Column | Definition | Unit / format | Notes |
|---|---|---|---|
SeqID |
Original unique identifier for the sequencing record or sequencing sample | text | Sequence-level or sequencing-sample identifier within the raw sequencing dataset |
N_SamID |
Standardized sample identifier used in the archive | text | Harmonized sample identifier |
DataID |
Unique identifier for the corresponding reformatted dataset | text; format D0XX |
Links the raw sequencing sample to the reformatted dataset |
N_SiteID |
Standardized site identifier used in the archive | text | Harmonized site-level identifier |
SampleID |
Original sample identifier | text | Sample identifier as reported in the original source or raw metadata |
RawSeqID |
Unique identifier for the raw sequencing dataset | text; format RS0XX |
Links the sample metadata to the corresponding raw sequencing archive |
Notes:
- Each row represents one sequencing sample.
NAindicates that the value was not available, not reported, or not applicable.
5.5 File: 3_Data/1_Reformat/D0XX/SiteSpeciesMatrix_cor.csv
Purpose:
Species-by-sample matrix for one reformatted dataset.
Applies to:
All D0XX folders in 3_Data/1_Reformat.
Structure:
- rows = species names
- columns = sample names
- cell entries = recorded matrix values for each species in each sample in the harmonized dataset
Notes:
- Species names in this file are standardized/corrected.
- Sample names correspond to the paired metadata file
SiteMetadata.csv.
5.6 File: 3_Data/1_Reformat/D0XX/SiteMetadata.csv
Purpose:
Sample-level metadata for the samples included in SiteSpeciesMatrix_cor.csv.
Applies to:
All D0XX folders in 3_Data/1_Reformat.
Each row describes one sample.
| Column | Definition | Unit / format | Notes |
|---|---|---|---|
ItemID |
Unique identifier for the metadata record | text | One row per sample record |
DataID |
Unique identifier for the reformatted dataset | text; format D0XX |
Links the sample to the dataset |
N_RivID |
Standardized river identifier used in the archive | text | Harmonized river-level identifier |
N_SiteID |
Standardized site identifier used in the archive | text | Harmonized site-level identifier |
samID |
Sample identifier | text | Sample-level identifier used within the dataset |
Continent |
Continent where the sample was collected | text | Geographic location |
Country |
Country where the sample was collected | text | Geographic location |
RefID |
Reference/source identifier associated with the sample or dataset | text | Links to the original source or internal reference system |
RiverName |
Name of the river from which the sample was collected | text | River name as standardized in the archive |
Primer |
Primer set used for the sample | text | Primer information associated with amplification |
Year |
Sampling year | year | Year in which the sample was collected |
Month |
Sampling month | month or integer | Month in which the sample was collected |
Lon |
Longitude of the sampling location | decimal degrees | Geographic coordinate |
Lat |
Latitude of the sampling location | decimal degrees | Geographic coordinate |
SR |
Original individually reported species richness | count | Species richness reported in the original dataset/source for the corresponding sample, site, or record, as applicable |
SR_re |
Unified analyzed species richness | count | Species richness obtained after unified analysis/harmonization in this project |
Notes:
- Each row represents one sample.
- Coordinates are reported in decimal degrees where available.
NAindicates that the value was not available, not reported, or not applicable for that sample.
