Data from: Applied paleoecology for macrophyte restoration: Defining Anthropocene baselines in Lake Liangzi
Data files
Apr 15, 2026 version files 68.13 KB
-
Fig._3a_3b_3c.csv
7.14 KB
-
Fig._4a_4l.csv
14.64 KB
-
Fig._7a_7b_7c_7d.csv
5.88 KB
-
Fig.2.csv
7.15 KB
-
Fig.5_Fig.S6.csv
1.48 KB
-
Fig.6.csv
111 B
-
Fig.S2.csv
431 B
-
Fig.S3.csv
462 B
-
Fig.S4_Fig.S5a.csv
2.22 KB
-
Fig.S5b_5c.csv
6.19 KB
-
Fig.S8.csv
1.78 KB
-
README.md
20.64 KB
Abstract
The global degradation of freshwater lakes threatens biodiversity and critical ecosystem services. Restoring aquatic macrophytes is essential for reversing this decline, yet a fundamental challenge persists: the debate between aiming for historical baseline conditions or accepting novel ecosystems in the Anthropocene. This dilemma is exacerbated by a lack of long-term data about pre-degradation states. Current assessments are based mainly on short-term in situ observations, providing limited insights into historical reference conditions and hindering the development of effective restoration targets. Here, we develop an evolutionary restoration methodological approach that integrates sedimentary ancient DNA (sedaDNA), pollen, macrofossils, satellite remote sensing, and contemporary surveys to determine the centennial-scale trajectory of macrophyte communities and inform their restoration targets. Taking Lake Liangzi (eastern China) as a case study, our results show that submerged taxa (e.g., Potamogeton crispus, Najas minor, and Chara spp.) dominated under oligotrophic conditions until the 1960, after which nutrient pollution drove a shift to floating-leaved and emergent taxa (e.g., Nelumbo nucifera, Nymphaea macrosperma, and Pontederia crassipes), and eventually to an algal-dominated regime. Ecological Quality Ratio (EQR) assessments reveal a clear degradation trajectory in Lake Liangzi; macrophytes ecological quality declined from Good (pre-2010) to Poor (present). The results indicate that restoration to pre-1960 baseline conditions is unlikely to be directly achievable. Instead, we propose a phased restoration pathway that first targets re-establishment of the ecologically functional Anthropocene baseline represented by the 1960–2010 coexistence state, with the longer-term aspiration of recovering submerged macrophyte-dominated communities. Synthesis and applications: Findings move beyond the polarized historical vs. novel ecosystem debate by advancing a dynamic restoration pathway and targets for macrophytes. This pathway presents the Anthropocene baseline as a pragmatic stepping stone toward long-term recovery and offers managers phased, realistic targets that sustain ecosystem services while keeping the operational historical reference as an aspirational goal. The method established offers a practical basis for developing effective restoration strategies and can be applied to Lake Liangzi and similar shallow lakes. It directly informs global freshwater restoration under the UN Decade on Ecosystem Restoration.
Dataset DOI: 10.5061/dryad.w6m905r3k
Description of the data and file structure
This dataset includes 11 CSV files, corresponding to Figures 2–7 and Supplementary Figures S2–S8 in the associated manuscript. Each file is described below with its structure, column definitions, units, and organization. Missing data are represented by blank cells (empty strings between commas).
1. Fig.2.csv – Data for Figure 2
File purpose
Time‑series data of multiple proxies in lake sediments, including aquatic macrophyte abundance, nutrient levels, and catchment erosion indicators. Different proxies may have different temporal resolutions.
File structure
- Format: CSV (comma‑separated values)
- Rows: 55 (header + 54 data rows)
- Columns: 22
- Encoding: UTF‑8 with BOM
- Missing values: blank cells (no NA)
Column definitions
Year(year AD) – common age axis forOUT readsand some other proxies.OUT reads(unitless) – high‑throughput sequencing output read counts.Year1(year AD) – age forPollendata.Pollen(grains/g or grains/cm³) – pollen concentration.Year2(year AD) – age forMacrofossilsdata.Macrofossils(remains/cm³ or /g) – plant macrofossil concentration.Year3(year AD) – age forTN(%)andTOC(%).TN(%)(%) – total nitrogen content.TOC(%)(%) – total organic carbon content.Year4(year AD) – age forChl-a(ug/L).Chl-a(ug/L)(μg/L) – chlorophyll‑a concentration.Year5(year AD) – age forEF-TP.EF-TP(unitless) – enrichment factor corrected total phosphorus.Year6(year AD) – age forDI-TP(ug/L).DI-TP(ug/L)(μg/L) – diatom‑inferred total phosphorus concentration.Year7(year AD) – age forTP(ug/L).TP(ug/L)(μg/L) – total phosphorus concentration.Year8(year AD) – age forPb(mg/kg).Pb(mg/kg)(mg/kg) – lead concentration.Year9(year AD) – age forFe/Mn.Fe/Mn(unitless) – iron‑to‑manganese mass ratio.Year10(year AD) – age forχlf(10-8 m3/kg-1).χlf(10-8 m3/kg-1)(10⁻⁸ m³ kg⁻¹) – low‑frequency mass magnetic susceptibility.Median grain size(um)(μm) – median grain size; its age is taken from theYearcolumn of the same row.
Data organization
Each row represents one sediment sample (depth/age). Different proxies have different age resolutions; blanks indicate no data for that proxy at that depth. Rows are ordered from youngest (top, ~2021) to oldest (bottom, ~1900).
Correspondence to manuscript figures
This file is used to generate Figure 2.
OUT reads,Pollen,Macrofossils→ panel (a)TN(%),TOC(%),Chl-a,EF-TP,DI-TP,TP→ panel (b)Pb,Fe/Mn,χlf,Median grain size→ panel (c)
Original figure caption from the manuscript (for reference)
(a) Total abundance of aquatic macrophytes reconstructed from three proxies, as indicated by sedaDNA reads, plant macrofossil remains, and pollen concentrations.
(b) Lake nutrient levels, indicated by TOC, TN, DI-TP and EF-TP (derived from sedimentary reconstruction), along with monitored Chl-a and TP in water (obtained from instrumental water quality data).
(c) Human activities and catchment erosion, represented by heavy metal indicators Pb, Fe/Mn, magnetic susceptibility, and median grain size.
2. Fig._3a_3b_3c.csv – Data for Figure 3
File purpose
Species‑level read counts (absolute abundance) and relative abundances of aquatic macrophytes (submerged, floating‑leaved, emergent) over time.
File structure
- Format: CSV
- Rows: ~25 (header + 24 data rows)
- Columns: 40
- Encoding: UTF‑8 with BOM
- Missing values: none (all cells filled)
Column definitions
Year(year AD) – age axis for all species data.- Columns without suffix (e.g.,
Ceratophyllum demersum) – absolute abundance (counts or concentration) of each taxon. - Columns with suffix
1(e.g.,Ceratophyllum demersum1) – relative abundance (proportion or percentage) of the same taxon.
All taxon names are as follows (species level):
Ceratophyllum demersum, Hydrilla verticillata, Myriophyllu spicatum, Nitellopsis obtusa, Stuckenia pectinata, Potamogeton perfoliatus, Potamogeton natans, Potamogeton maackianus, Potamogeton lucens, Potamogeton gramineus, Potamogeton distinctus, Potamogeton crispus, Potamogeton berchtoldii, Nelumbo nucifera, Nymphaea macrosperma, Phragmites australis, Zizania latifolia, Nymphoides peltata, Pontederia crassipes, Trapa natans.
Data organization
Each row represents one sediment layer (age), ordered from youngest (~2021) to oldest (~1902). Absolute and relative abundances of all species are given for the same Year.
Correspondence to manuscript figures
- Absolute abundance columns → Figure 3a
- Relative abundance columns → Figures 3b and 3c
Original figure caption from the manuscript (for reference)
(a) Time series variation of OTUs reads in macrophyte community (Blue represents submerged macrophytes, red represents floating-leaved plants, and green represents emergent plants).
(b) Macrophyte community composition of LZHZ23-1 at the species level, abundance (log+1) is expressed as a percentage over total abundance.
(c) Macrophyte community composition of LZHZ23-1 at the species level according to functional groups (blue represents submerged macrophytes, purple represents floating-leaved plants, and green represents emergent plants).
3. Fig._4a_4l.csv – Data for Figure 4
File purpose
α‑diversity indices, β‑diversity (TBI), NMDS coordinates, and three sets of co‑occurrence network edge lists with full taxonomic annotation for each node.
File structure
- Format: CSV
- Rows: ~70 (header + 69 data rows)
- Columns: 64
- Encoding: UTF‑8 with BOM
- Missing values: many blanks in later rows (only network columns are filled after row 26)
Column definitions (first part)
Year(year AD) – age (only first 26 rows have data).Richness,Shannon,Simpson,Chao1,ACE,Pielou(all unitless) – α‑diversity indices.Height– mean macrophyte height (units as in original study).TBI(unitless) – temporal beta diversity index.NMDS1,NMDS2(unitless) – NMDS ordination scores.
Network columns – three identical sets (suffixes , 1, 2):
source,target– OTU IDs (nodes).weight– edge weight (0–1).correlation– Spearman’s correlation coefficient (positive/negative).Id,label– OTU identifier.Domain,Phylum,Class,Order,Family,Genus,Species– full taxonomy of the node.
Data organization
Rows 1–26 contain diversity indices, NMDS scores, and all three network sets. From row 27 onward, only network columns (edges) are filled; diversity and NMDS columns are blank. The three network sets correspond to three time intervals.
Correspondence to manuscript figures
- Diversity indices and NMDS → Figures 4a–4i
- Network set 1 (no suffix) → Figure 4j (1900–1960)
- Network set 2 (suffix
1) → Figure 4k (1960–2010) - Network set 3 (suffix
2) → Figure 4l (2010–2021)
Original figure caption from the manuscript (for reference)
Temporal variations in α and β-diversity indices of macrophyte community and differences in the co-occurrence network pattern (1900–1960; 1960–2010; 2010–2021) in Lake Liangzi based on sedimentary ancient DNA (sedaDNA) analysis.
(a) Richness; (b) Shannon’s index; (c) Simpson’s index; (d) Chao1; (e) abundance based coverage estimator (ACE); (f) Pielou’s evenness; (g) The weighted mean macrophyte height of macrophytes functional traits; (h) Temporal beta diversity (TBI); (i) NMDS plot of the macrophyte community. Comparisons of the co-occurrence network patterns of macrophyte taxa (at the species level) in the three intervals: (j) 1900–1960, (k) 1960–2010, and (l) 2010–2021. (Note: The size of a circle represents the relative abundance of the taxon, and the color indicates the importance of the taxon; the thickness of each connecting line represents the strength of the significant Spearman’s correlations (|r| > 0.6, p < 0.05) between the nodes; green edges represent negative correlations, and red edges represent positive correlations.)
4. Fig.5_Fig.S6.csv – Data for Figure 5 and Figure S6
File purpose
Presence/absence of different proxy types (sedaDNA, pollen, macrofossils) for each taxon per decade.
File structure
- Format: CSV
- Rows: 12 (header + 11 data rows for 11 decades)
- Columns: 14
- Encoding: UTF‑8 with BOM
- Missing values: empty strings (no data for that taxon/decade)
Column definitions
Year range– decade interval (e.g.,1900_1910).- Columns with
s_(species level) org_(genus level):
s_Ceratophyllum demersum,s_Hydrilla verticillata,s_Myriophyllum spicatum,s_Nelumbo nucifera,s_Potamogeton crispus,g_Nymphoides,g_Nymphaea,g_Najas,g_Vallisneria,g_Trapa,g_Myriophyllum,g_Potamogeton,g_Ceratophyllum.
Cell values:SeDNA_Macrophyte,SeDNA_Pollen,Macrophyte_Pollen,SeDNA_Macrophyte_Pollen, or empty (no record).
Data organization
Each row is a decade (1900–1910 to 2000–2010). Each column is a taxon. The cell indicates which proxy recorded that taxon in that decade.
Correspondence to manuscript figures
Used for both Figure 5 and Figure S6.
Original figure caption from the manuscript (for reference)
Comparison of macrophyte communities based on sedimentary ancient DNA, pollen, plant macrofossils (The prefixes "g" and "s" preceding species names indicate the taxonomic identification level: "g" denotes that the macrophyte was identified at the genus level, while "s" denotes identification at the species level. The "macrophyte" data presented here are mainly derived from sedimentary records, including aquatic vegetation reconstructed from sedaDNA, pollen, and plant macrofossils.)
5. Fig.6.csv – Data for Figure 6
File purpose
Ecological Quality Ratio (EQR) scores for macrophyte status over different periods.
File structure
- Format: CSV
- Rows: 7 (header + 6 data rows)
- Columns: 2
- Encoding: UTF‑8
- Missing values: none
Column definitions
Period– time interval as string (e.g.,1960-1970).EQR(unitless, 0–1) – Ecological Quality Ratio (higher = better ecological status).
Data organization
Rows ordered from the earliest period (1960–1970) to the latest (2010–2020). EQR decreases over time.
Correspondence to manuscript figures
Directly used for Figure 6.
Original figure caption from the manuscript (for reference)
EQR scoring for macrophyte status over different periods.
6. Fig._7a_7b_7c_7d.csv – Data for Figure 7
File purpose
Two time‑series datasets of standardized environmental/biological variables (for Mantel tests) and variable importance from random forest models.
File structure
- Format: CSV
- Rows: 13 (header + 12 data rows)
- Columns: 32
- Encoding: UTF‑8 with BOM
- Missing values: present (later rows have blanks in some sections)
Column definitions
First set (columns Year to Fish) – standardized values for period ~1959–1902:
Year, OUT, TBI, Height, Shannon, Richness, TOC, EF-TP, Chla, Pb, Cd, Mediangrainsize, Pre, Temp, Magneticsusceptibility, Fish.
Second set (columns Year1 to Fish1) – same variables for period ~2021–1961 (different time window).
Third set (random forest model 1):
Variable1– predictor names.%IncMSE1– increase in mean squared error (%).IncNodePurity1– increase in node purity.
Fourth set (random forest model 2):
Variable– predictor names.%IncMSE– increase in MSE (%).IncNodePurity– increase in node purity.
Data organization
Rows 1–11 contain all four sets. Row 12 onward: first set blank, second set partially filled, third/fourth sets may be blank. Age order is from young to old within each set.
Correspondence to manuscript figures
- First two sets → Figures 7a and 7b (Mantel test correlation matrices).
- Third and fourth sets → Figures 7c and 7d (random forest importance plots).
Original figure caption from the manuscript (for reference)
Relationships between the macrophyte community structure, diversity, and environmental factors.
(a) 1900–1960 and (b) 1960–2021, biotic environmental drivers of macrophyte communities at Lake Liangzi over the past 120 yr. Pairwise comparisons of environmental factors are expressed as a color gradient using Spearman’s correlation coefficients. Macrophyte OTUs, temporal beta diversity (TBI), functional traits (the weighted mean plant height), and diversity (richness and Shannon index) are tested for correlations with individual environmental factors using the Mantel test. Line widths represent Mantel’s statistics for the corresponding correlations, and colors indicate significance.
(c) 1900–1960 and (d) 1960–2021, the random forest analysis (ntree = 1000) of multiple factors (Increase in MSE (%) and Increase in Node Purity, larger values indicate greater variable importance; the color gradient from dark to light represents decreasing contribution).
7. Fig.S2.csv – Data for Figure S2
File purpose
Depth profiles of ¹³⁷Cs and ²¹⁰Pbₑₓ activities for sediment core dating.
File structure
- Format: CSV
- Rows: 21 (header + 20 data rows)
- Columns: 3
- Encoding: UTF‑8
- Missing values: none (0 for ¹³⁷Cs below detection limit)
Column definitions
Depth(cm)(cm) – core depth.137Cs(Bq/kg)(Bq/kg) – ¹³⁷Cs specific activity.210Pbex(Bq/kg)(Bq/kg) – excess ²¹⁰Pb activity.
Data organization
Rows ordered by increasing depth (1 cm to 57 cm). Depth intervals are irregular (2 cm, then 4 cm, etc.).
Correspondence to manuscript figures
Used to generate Figure S2 (age‑depth model).
Original figure caption from the manuscript (for reference)
The profiles of ²¹⁰Pbₑₓ, ¹³⁷Cs specific activities, and the CRS age models (the ¹³⁷Cs-corrected CRS model using ¹³⁷Cs time marker of 1963 CE) derived from ²¹⁰Pb and ¹³⁷Cs activities, including the linear fitting model of significant ¹³⁷Cs time markers and the core top (red stars), with a high fitting degree (R² = 0.91, p < 0.001).
8. Fig.S3.csv – Data for Figure S3
File purpose
Remote‑sensing derived coverage of floating‑leaved (FEAV) and submerged (SAV) aquatic vegetation from 2008 to 2023.
File structure
- Format: CSV
- Rows: 17 (header + 16 data rows)
- Columns: 3
- Encoding: UTF‑8
- Missing values: none
Column definitions
Year(year AD) – observation year.FEAV(unitless relative coverage) – floating‑leaved aquatic vegetation.SAV(unitless relative coverage) – submerged aquatic vegetation.
Data organization
Rows in chronological order (2008 to 2023).
Correspondence to manuscript figures
Directly used for Figure S3.
Original figure caption from the manuscript (for reference)
Aquatic vegetation coverage data retrieved by remote sensing from 2008 to 2023.
9. Fig.S4_Fig.S5a.csv – Data for Figures S4 and S5a
File purpose
Absolute abundance of aquatic macrophyte species versus depth (S4) and versus age (S5a) from sedaDNA.
File structure
- Format: CSV
- Rows: 26 (header + 25 data rows)
- Columns: 22
- Encoding: UTF‑8
- Missing values: none (zero indicates absence)
Column definitions
Depth(cm) – core depth.Year(year AD) – age.- Columns 3–22 – same species as in
Fig.3a_3b_3c.csv(absolute abundance).
Data organization
Rows ordered by increasing depth (young to old). Depth and age are inversely related.
Correspondence to manuscript figures
- Depth vs. abundance → Figure S4 (CONISS cluster analysis).
- Age vs. abundance → Figure S5a (sedaDNA time series).
Original figure caption from the manuscript (for reference)
Fig. S4: CONISS cluster analysis results of sedaDNA from Lake Liangzi (LZHZ23-1) over the past century.
Fig. S5a: sedaDNA time series of aquatic macrophytes (Blue represents submerged macrophytes, red represents floating-leaved plants, and green represents emergent plants).
10. Fig.S5b_5c.csv – Data for Figures S5b and S5c
File purpose
Pollen‑based (S5b) and macrofossil‑based (S5c) time series of aquatic macrophytes.
File structure
- Format: CSV
- Rows: 47 (header + 46 data rows)
- Columns: 48
- Encoding: UTF‑8
- Missing values: many blanks (second group empty after ~row 20)
Column definitions
First group (columns 1–17) – pollen‑based data:
Year, Najas_minor, Potamogeton_crispus, Ceratophyllum_demersum, Gloeotrichia_echinulata_colonies, Potamogeton, Myriophyllum_spicatum, Myriophyllum_sp., Najas_marina, Vallisneria_spinulosa, Nitella_sp., Chara_sp., Hydrilla_verticillata, Nymphaeaceae, Trapa_natans, Asteraceae, Euryale_ferox. All abundance units are counts per unit sediment (grains/g or remains/g).
Second group (columns 18–48) – macrofossil‑based data (genus level):
Year1, Vallisnena, Najas, Ceratophyllum, Potamogeton1, Myriophyllum, Salvinia, Nymphaea, Nelumbo, Acorus, Sagittaria, Alternanthera, Lobelia, Actinostemma, Polygonum, Lemna, Azolla, Nuphar, Eichhornia, Nymphoides, Callitriche, Hydrocharis, Trapa, Ceratopteris, Typha.
Data organization
Rows 1–20 contain both groups. From row 21 onward, only the first group (pollen) is present. The first group spans ~2008–1901, the second group only ~2008–1950.
Correspondence to manuscript figures
- First group → Figure S5b (pollen).
- Second group → Figure S5c (macrofossils).
Original figure caption from the manuscript (for reference)
(b) pollen-based time series of aquatic macrophytes; (c) macrofossil-based time series of aquatic macrophytes (Blue represents submerged macrophytes, red represents floating-leaved plants, and green represents emergent plants).
11. Fig.S8.csv – Data for Figure S8
File purpose
Standardized environmental and biological variables (including water level) from 1993 to 2021 for time‑series comparison.
File structure
- Format: CSV
- Rows: 9 (header + 8 data rows)
- Columns: 18
- Encoding: UTF‑8
- Missing values: none
Column definitions
All variables are standardized (mean ≈0, SD ≈1).
Year,Year1(identical, redundant) – year AD.OUT,TBI,Height,Shannon,Richness,TOC,EF-TP,Chla,Pb,Cd,Mediangrainsize,Pre,Temp,Magneticsusceptibility,Fish,Wate level(water level).
Data organization
Rows ordered from youngest (2021) to oldest (1993). Year and Year1 are identical.
Correspondence to manuscript figures
Directly used for Figure S8.
Original figure caption from the manuscript (for reference)
Comparison of macrophytes communities based on sedimentary ancient DNA (sedaDNA), pollen, and plant macrofossils.
Sharing/Access information
Data are publicly available at Dryad:
https://doi.org/10.5061/dryad.w6m905r3k
The dataset is associated with the following publication:
Jin, Y., Zhang, K., Lin, Q., Liu, S., Huang, S., Han, Y., Luo, J., Meadows, M. E. (2026). Applied paleoecology for macrophyte restoration: Defining Anthropocene baselines in Lake Liangzi. Journal of Applied Ecology.
Code/Software
No custom code or software was used to generate the primary data. The data are presented as raw values collected from laboratory analyses (sedaDNA sequencing, pigment analysis, radiometric dating, etc.) and remote sensing. Statistical analyses (Mantel test, random forest, NMDS, co‑occurrence networks) were performed using R (version 4.2.1) with packages vegan, randomForest, igraph, and ggplot2. The scripts are available from the corresponding author upon reasonable request.
Note on figure captions: The original figure captions from the manuscript are provided above for reference only. The dataset is fully interpretable without them, as all column definitions, units, and data organization are explicitly described in this README.
