Contrasting associations of urbanization and lateral hydrological gradients with functional structure and rarity of spontaneous plants in river corridors: Evidence from Changde, China
Data files
Sep 25, 2026 version files 358.56 KB
-
Code.zip
18.68 KB
-
processed.zip
314.25 KB
-
README.md
11.35 KB
-
results.zip
14.28 KB
Abstract
This dataset supports analyses of spontaneous plant communities in urban river corridors in Changde, China. It contains data from 1,144 sampling plots and 248 plant species, including species importance values, relative cover, functional traits, species diversity, functional diversity, community weighted mean traits, functional rarity, urbanization gradients, lateral hydrological gradients, statistical outputs, R scripts, and supplementary tables. The data were used to examine plant community assembly and functional responses to impervious surface area (ISA) and water distance. The dataset can be reused for studies of urban vegetation, riparian ecology, functional diversity, and functional rarity.
Description:
This package contains the processed data, R scripts, statistical outputs, and supplementary tables used to support the analyses reported in the manuscript.
The study examines spontaneous plant communities in urban river corridors in Changde, China. The analyses evaluate species diversity, functional diversity, functional trait composition, functional rarity, and their associations with impervious surface area and water distance.
Directory structure:
Code.zip/
R scripts used for the main analyses.
processed.zip/
Processed datasets used as inputs for the analyses.
results.zip/
Statistical outputs generated for the main analyses reported in the manuscript.
Supplementary.zip/
Supplementary tables and supporting analyses.
The scripts use relative file paths and are intended to be run with the Code folder as the working directory. Required R packages are listed in the individual scripts.
Data dictionary and conventions:
- Level.of.urbanization and urban_level: L = low urbanization (ISA < 0.20); M = medium urbanization (0.20 <= ISA <= 0.50); H = high urbanization (ISA > 0.50). ISA is the proportion of impervious surface within a 500 m radius buffer, stored as a fraction from 0 to 1 (unitless).
- Habitat: W = water; WS = waterside; R = revetment; FR = riverfront. These categories follow the lateral position from water toward land.
- WD = water distance in meters (m). Woody_coverage = woody plant cover in percent (%).
- ecotype: Aquatic = aquatic; Hydrophytic = wet-site adapted; Mesophytic = intermediate-moisture adapted; Xerophytic = dry-site adapted.
- leaf_texture: Fleshy, Herbaceous, Leathery, Membranous, and Papery denote the named leaf-texture categories; no_leaf_texture denotes no assigned leaf-texture category.
- Pollination_mode: Anemophilous = wind pollination; Entomogamous = insect pollination; Selfed = self-pollination; no_Pollination_mode denotes no assigned pollination category.
- dispersal_mode: Anemochory = wind dispersal; Hydrochory = water dispersal; Autochory = self-dispersal; Zoochory = animal dispersal.
- A plus sign in a pollination or dispersal value indicates a combination of the named modes (for example, Anemochory+Hydrochory).
- rarity_status: rare = classified as functionally rare; non_rare = classified as not functionally rare.
- NA denotes a missing or undefined value. Plot and site identifiers are labels, not measured quantities.
Abbreviations used below: CSV = comma-separated values; CWM = community-weighted mean; FDis = functional dispersion; FDiv = functional divergence; FEve = functional evenness; FRic = functional richness; RaoQ = Rao's quadratic entropy; SEM = structural equation modeling; PCoA = principal coordinates analysis; PERMANOVA = permutational multivariate analysis of variance; FDR = false discovery rate; FAMD = factor analysis of mixed data. R is the statistical programming language in the scripts section (the separate Habitat value R denotes revetment). Statistical indices, proportions, standardized values, and regression coefficients are unitless unless otherwise specified.
PROCESSED/
Species_sample_matrix.csv
Species by plot matrix used to calculate alpha diversity.
Rows:
Plant species
Columns:
Sampling plots
Values:
Species importance values.
comm.csv
Plot by species community matrix used to calculate functional diversity and community weighted mean traits.
Rows:
Sampling plots
Columns:
Plant species
Values:
Relative species cover.
Group.csv
Grouping information for the 1,144 sampling plots.
Variables:
- GroupID
- Level.of.urbanization: L, M, and H (see data dictionary)
- Habitat: W, WS, R, and FR (see data dictionary)
Trait_matrix.csv
Functional trait matrix for the plant species included in the analyses.
Rows:
Plant species
Columns:
- ecotype
- leaf_texture
- Pollination_mode
- dispersal_mode
- plant_height: original species height in meters (m)
The functional-diversity and functional-structure analyses apply log(plant_height + 1) transformation and standardization where specified in the corresponding scripts. Categorical trait values and combinations are defined in the data dictionary above.
α_diversity.csv
Alpha-diversity indices calculated for all sampling plots.
Variables:
- shannon: Shannon-Wiener index
- simpson: Simpson index
- patrick: Species richness
- pielou: Pielou evenness
- Level.of.urbanization
- Habitat
Pielou evenness is recorded as NA when a plot contains only one species, because evenness cannot be calculated for a single-species plot.
functional_diversity.csv
Functional-diversity indices calculated for the sampling plots.
Variables include:
- FDis: Functional Dispersion
- FDiv: Functional Divergence
- FEve: Functional Evenness
- RaoQ: Rao's quadratic entropy
- FRic: Functional Richness
- Level.of.urbanization
- Habitat
dummy_CWM_matrix.csv
Site-level community weighted mean trait matrix used for functional-structure analyses.
Rows:
Sampling plots
Columns:
Site-level CWM values for plant height and indicator-coded categories of the categorical traits.
Categorical trait columns represent relative-cover-weighted proportions (unitless, 0 to 1) of species belonging to each trait category. CWM_plant_height is based on transformed and standardized height values and is unitless, not meters.
SEM.csv
Dataset used for the structural equation models.
Variables include:
- ISA: impervious surface proportion (unitless, 0 to 1)
- Level.of.urbanization
- CWM_plant_height: community-weighted mean of transformed, standardized plant height (unitless)
- CWM_leaf_texture
- CWM_Pollination_mode
- CWM_ecotype
- CWM_dispersal_mode
- FDis
- Shannon
- Rarity
- Habitat
- WD: water distance in meters (m)
- Woody_coverage: woody plant cover in percent (%)
ISA represents impervious surface proportion within a 500 m buffer. The SEM script applies the transformations and standardization required for model fitting. Other CWM traits, FDis, Shannon, and Rarity are unitless indices or trait-category proportions.
non-rare+rare_LMH.csv
Species-level trait and rarity-status data summarized across urbanization levels.
Variables include:
- species
- urban_level
- rarity_status
- ecotype
- leaf_texture
- Pollination_mode
- dispersal_mode
- plant_height: transformed and standardized height value (unitless; not the original height in meters)
non-rare+rare_Habitat.csv
Species-level trait and rarity-status data summarized across habitat types.
Variables include:
- species
- Habitat
- rarity_status
- ecotype
- leaf_texture
- Pollination_mode
- dispersal_mode
- plant_height: transformed and standardized height value (unitless; not the original height in meters)
RESULTS/
01_Shannon_FDis_group_comparisons.csv
Kruskal-Wallis tests and pairwise Wilcoxon tests for Shannon-Wiener diversity and FDis across urbanization levels and habitat types. Pairwise P values are adjusted using the Benjamini-Hochberg method.
02_Shannon_FDis_group_descriptives.csv
Group-wise sample sizes and descriptive statistics for Shannon-Wiener diversity and FDis, including the mean, standard deviation, median, interquartile range, minimum, and maximum.
03_Shannon_FDis_Spearman_by_group.csv
Spearman rank correlations between Shannon-Wiener diversity and FDis within each urbanization level and habitat type.
04_Rarity_ISA_WD_regression.csv
Linear regression results describing the associations of functional rarity with standardized ISA and log transformed standardized WD. The file includes sample size, slope, standard error, test statistic, P value, R squared, adjusted R squared, and residual degrees of freedom.
05_CWM_ISA_WD_results.csv
Regression results for community weighted mean traits along the ISA and WD gradients. The file includes the slope, standard error, P value, R squared, adjusted R squared, sample size, and FDR-adjusted P value.
06_PCoA_axis_explained_variance.csv
Eigenvalues and explained variance for the first two PCoA axes based on the Gower distance among functional traits.
07_PCoA_PERMANOVA_results.csv
PERMANOVA results for differences in community functional composition among urbanization levels and habitat types. The tests use Euclidean distances among site-level scores calculated from the PCoA functional trait space and 999 permutations.
08_PCoA_betadisper_tests.csv
Tests of homogeneity of multivariate dispersion for urbanization levels and habitat types. The dispersion tests use the same site-level Euclidean distance matrix as the PERMANOVA.
09_SEM_main_results.csv
Fit indices, standardized path associations, and R squared values for the ISA and WD overall structural equation models.
10_SEM_total_associations.csv
Bootstrap summaries of standardized total associations with functional rarity for the ISA and WD models. The file includes direct associations, indirect associations, total associations, model R squared, relative association magnitude, and percentile-based 95% bootstrap confidence intervals based on 5,000 case-resampling replicates.
R SCRIPTS/
01_α_diversity.R
Calculates alpha-diversity indices from the species importance-value matrix and produces comparisons across urbanization levels and habitat types.
02_functional_diversity.R
Calculates functional-diversity indices from relative-cover data and the functional trait matrix, and produces comparisons across urbanization levels and habitat types.
03_spearman.R
Calculates Spearman rank correlations between Shannon-Wiener diversity and FDis within urbanization levels and habitat types.
04_Functional_structure_and_trait.R
Examines functional trait-space patterns and community weighted mean trait associations along ISA and WD gradients. The script includes PCoA, PERMANOVA, multivariate dispersion diagnostics, CWM calculations, and trait regression analyses.
05_Rarity_ISA_WD.R
Examines the linear associations between functional rarity and standardized ISA and WD.
06_Non-rare and rare species stacking.R
Summarizes and visualizes the functional trait composition of non-rare and rare species across urbanization levels and habitat types.
07_SEM.R
Fits the overall ISA and WD structural equation models, calculates standardized total associations, estimates uncertainty using 5,000 bootstrap resamples, and produces the SEM and total-association figures.
SUPPLEMENTARY/ (Zenodo)
Supplementary_1_Sampling_and_trait_proportion.docx
Supporting information on sampling structure and trait composition.
Supplementary_2_SEM_screening_and_diagnostics.docx
Environmental-variable screening, model diagnostics, and SEM statistical checks.
Supplementary_3_SEM_sensitivity_and_robustness.docx
Sensitivity and robustness analyses for spatial scale, spatial dependence, model specification, environmental-variable selection, and related SEM assessments.
Supplementary_4_Trait_selection_and_FAMD.docx
Functional-trait screening results and FAMD outputs supporting the selection of traits used in the main analyses.
SOFTWARE
R version 4.5.2
Vegetation data were collected across urbanization levels and lateral habitat types. Species importance values were used for alpha diversity, and relative cover was used for functional diversity and community weighted mean trait calculations. Functional traits included ecotype, leaf texture, pollination mode, dispersal mode, and plant height. Trait distances, PCoA, PERMANOVA, regression models, and structural equation models were used to evaluate community patterns and environmental associations. SEM total associations were estimated using 5,000 bootstrap resamples with percentile based 95% confidence intervals.
