Data from: Established short-rotation willow (Salix spp.) as an effective facilitator of native hardwood tree regeneration in New York State, USA
Data files
Jul 22, 2026 version files 2.81 MB
-
1_Submited_Data_and_Code.zip
2.79 MB
-
README.md
14.95 KB
Abstract
These are the compiled datasets for the field and lab work associated with two consecutive studies that took place between September 2024 and December 2025. Data consist of point and plot data for estimating seedling density in willow stands. Experiment one used the point-quarter method, which requires the distance to the nearest seedling in each quadrat; height and species were also recorded. Experiment one includes data to estimate seedling density based on distance along transects extending from a seed source. Experiment two utilized 2m-diameter (generally) fixed plots where the species and height of seedlings within the plot were measured. Complementary plot data were also collected concerning the willow crop, understory vegetation, soils, and seed sources within 300 m of the plot.
The files are contained in two folders. Each folder contains files associated with statistical and data analyses using SAS to produce the tables and figures in the manuscript: input data, SAS programs, and output data. The SAS programs are annotated so the tables and figures with which they are associated are identified. The SAS programs are designed to run standalone (the required datasets are embedded within the code). A duplicate of the input data is provided in a comma-delimited format (.CSV) for easier viewing. The SAS output has been saved as complete website (*.HTML) files with their associated figures stored in sub-directories of the same name. These can be opened by any browser to see the output as it was displayed by SAS. Ancillary output files created by the SAS code are also included, either as *.CSV or *.RTF file formats are also provided.
Description of data and files
1_Submited_Data_and_Code.zip: Zip file containing relevant files, folders, and subfolders
Folder Content and File Descriptions
FOLDER: “Study 1_Analyses”
a. SAS PROGRAM FILE: “Study 1 - Automated summary Final.SAS”
i. This code was used to produce results for Table 2, Figure 2, and Figure 3 in the manuscript
ii. Associated Files
- FILE: “Study 1 - Automated summary Final_inputdata.CSV”
a. This file contains the raw data collected for point-quarter method plots described in the manuscript.
b. Variable descriptions are defined in the Data Dictionary section below
-
FILE: “Study 1 - Automated summary Final_SAS_output.HTML”
-
FILE: “ht_hist_all.CSV” – output by program
-
FILE: “ht_hist_by_site.CSV” – output by program, but not used in manuscript
-
FILE: “spp_all.CSV” – output by program
-
FILE: “spp_by_site.CSV” – output by program, but not used in manuscript
-
FILE: “tpha_hist_all.CSV” – output by program
-
FILE: “tpha_hist_by_site.CSV” – output by program, but not used in manuscript
b. SAS PROGRAM FILE: “Study 1 - Regression final v6.SAS”
i. This code is used to produce Equation 4, and Figure 4 in the manuscript
ii. Associated Files
- FILE: “Study 1- Regression final v6_inputdata.CSV”
a. This data describes the point-quarter plots that were aligned as transects relative to potential seed sources on the FOX and HOM sites.
b. Variable descriptions are defined in the Data Dictionary section below
- FILE: “Study 1 - Regression final v6_SAS_output.HTML”
FOLDER: “Study 1 - Regression final v6_SAS_output_files”
a. Figures created for Study 1 - Regression final v6_SAS_output.HTML by SAS
FOLDER: “Study 2_Analyses”
a. SAS PROGRAM FILE: “Study 2_fixed_master_cluster_final.SAS”
i. This code is used to produce results for Tables 3, 4, and 5; also to produce Figures 5 in the manuscript
ii. Associated Files
- FILE: “Input_Dataset.CSV”
a. All three SAS programs for study 2 use the same input file
b. Variable descriptions are defined in Data Dictionary section below
- FILE: “ArithmeticMeans.RTF”
a. SAS-generated table that was used for initial drafts of Tables 3, 4, and 5
- FILE: “OrdinalSlopes.RTF”
a. SAS-generated table that was used to screen variables for regression analysis in [SAS PROGRAM FILE] “Study 2_fixed_master_logistic_3.3_mindist.SAS”
- FILE: “Study 2_fixed_master_cluster_final_SAS_output.HTML”
FOLDER: “Study 2_fixed_master_cluster_final_SAS_output_files
a. Figures created for Study 2_fixed_master_cluster_final_SAS_output.HTML by SAS
b. SAS PROGRAM FILE: “Study 2_fixed_master_logistic_3.3_mindist.SAS”
i. This code is used to produce results for Table 6 and Figure 6 in the manuscript
ii. Associated Files
- FILE: “Input_Dataset.CSV”
a. All three SAS programs for study 2 use the same input file
b. Variable descriptions are described in Data Dictionary section below
- FILE: “Study 2_fixed_master_logistic_3.3_mindist_SAS_output.HTML”
FOLDER: “Study 2_fixed_master_logistic_3.3_mindist_SAS_output_files”
a. Figures created for “Study 2_fixed_master_logistic_3.3_mindist_SAS_output_files.HTML” by SAS
b. SAS PROGRAM FILE: “Study 2_Distance_NonlinearCompare.SAS”
i. This code is used to produce results for Figures 7, and Equations 5 and 6 in the manuscript
ii. Associated Files
- FILE: “Input_Dataset.CSV”
a. All three SAS programs for study 2 use the same input file
b. Variable descriptions are described in Data Dictionary section below
- FILE: “Study 2_Distance_NonlinearCompare_SAS_output.HTML”
a. SAS output for this program
Data/Variable Dictionaries for CSV and RTF files
FOLDER: “Study 1_Analyses”
FILE: “Study 1 - Automated summary Final_inputdata.CSV”
- site: Site abbreviations: AVA, BEL, BUF, FOX, HOM, LEV, MUH, MAS
- plot: Plot number within site
- quar: Quartile within plot
- dist: Distance (ft) from seed source (FOX only in this dataset, not used)
- dis: Distance (cm) of seedling from point center to estimate seedling density using the point quarter method
- ht: Seedling height (cm)
- spp: Seedling species (general): ash (ash), asp (aspen), box (box elder), ced (cedar), dog (dogwood), elm (elm), fir (fir), hic (hickory), map (maple), rok (red oak), wok (white oak), wpi (white pine)
- count (counting variable)
FILE: “Study 1- Regression final v6_inputdata.CSV”
- site: Site abbreviations: AVA, BEL, BUF, FOX, HOM, LEV, MUH, MAS
- partial: 0/1 Variable to indicate if the transect was under canopy or open sky for its entire length.
- type: Canopy (plots that had willow >1 growing season in age overtopping the seedlings); Open (plots that had been harvested that season; willow was in its first growing season)
- line: transect number
- plot: plot number
- dist: Distance (ft) from seed source. Distance along transects
- dis: Distance (cm) of seedling from point center to estimate seedling density using the point quarter method
- ht: Seedling height (cm)
- spp: Seedling species (general): ash (ash), box (box elder), dog (dogwood), hic (hickory), map (maple), rok (red oak)
- count (counting variable)
- cover: dummy variable where 1 (type=Canopy), 0 (type=Open)
- home: dummy variable where 1 (site=HOM), 0 (site=FOX)
- canopy: dummy variable, same as cover
FILE: “ht_hist_all.CSV” (SAS generated output)
- ht_bin_cm: Bin for height classes 30=0-30cm, 60=30-60cm, 90=60-90cm, etc.
- n: count
- pct: percent of total count
FILE: “ht_hist_by_site.CSV” (SAS generated output)
- Site: Site abbreviation
- ht_bin_cm: Bin for height classes: 30=0-30cm, 60=30-60cm, 90=60-90cm, etc.
- n: count
- pct: percent of total count
FILE: "spp_all.CSV” (SAS generated output)
- Spp: Species code
- n: count
- pct: percent of total count
FILE: “spp_by_site.CSV” (SAS generated output)
- Site: Site abbreviation
- Spp: Species code
- n: count
- pct: percent of total count
FILE: “tpha_hist_all.CSV” (SAS generated output)
- tpha_bin: Bin for seedling density classes: 500 (0-500 seedlings per hectare), 1000 (500-1000 seedlings per hectare), --- 5000 (more than 5000 seedlings per hectare).
- n: count
- pct: percent of total count
FILE: “tpha_hist_by_site.CSV” (SAS generated output)
- Site: Site abbreviation
- tpha_bin: Bin for seedling density classes: 500 (0-500 seedlings per hectare), 1000 (500-1000 seedlings per hectare), etc., 5000 (more than 5000 seedlings per hectare).
- n: count
- pct: percent of total count
FOLDER: “Study 2_Analyses”
FILE: “Study 2_Input_Dataset.CSV”
- Date: Date of site visit
- Sn: plot serial number (100’s place for site, one’s place for plot number)
- Site: Site code
- Tug: 1=Tug Hill Plateau; 0=Lake Plain
- Plot: Plot number
- Lat: North DD.MM.mmmm (Degrees, and decimal minutes)
- Long: West DD.MM.mmmm (Degrees, and decimal minutes)
- Size: (not used)
- TPha_pqm: Seedlings per hectare as determined by the point quarter method
- TPha_pqm30: Seedlings per hectare as determined by the point quarter method for seedlings greater than 30 cm height
- pdia: diameter (m) of plot (two plots were reduced size due to the abundance of seedlings)
- height: height of willow (m)
- TPha_fix: Seedlings per hectare as determined fixed plot area
- TPha_fix30: Seedlings per hectare as determined fixed plot area for seedlings greater than 30 cm
- TPha_fix60: Seedlings per hectare as determined fixed plot area for seedlings greater than 30 cm
- TPha_fix90: Seedlings per hectare as determined fixed plot area for seedlings greater than 30 cm
- lnTPha_fix: ln(lnTPha_fix)
- lnTPha_fix30: ln(lnTPha_fix30)
- lnTPha_fix60: ln(lnTPha_fix60)
- lnTPha_fix90: ln(lnTPha_fix90)
- mindist: distance (m) to closest seed source (Figure 1 in manuscript)
- dist_360: mean distance (m) to closest seed source for 16 sectors (Figure 1 in manuscript)
- dist_W180: mean distance (m) to closest seed source for 9 sectors (Figure 1 in manuscript)
- dist_W90: mean distance (m) to closest seed source for 5 sectors (Figure 1 in manuscript)
- dist_top3: mean distance (m) to closest seed source for top 3 sectors (Figure 1 in manuscript)
- dist_top5: mean distance (m) to closest seed source for top 5 sectors (Figure 1 in manuscript)
- arc_360_50m: degrees of horizon arc where a seed source is present within 50 m for 16 sectors (Figure 1 in manuscript)
- arc_W180_50m: degrees of horizon arc where a seed source is present within 50 m for 9 sectors (Figure 1 in manuscript)
- arc_W90_50m: degrees of horizon arc where a seed source is present within 50 m for 5 sectors (Figure 1 in manuscript)
- arc_top3_50m: degrees of horizon arc where a seed source is present within 50 m for top 3 sectors (Figure 1 in manuscript)
- arc_top5_50m: degrees of horizon arc where a seed source is present within 50 m for top 5 sectors (Figure 1 in manuscript)
- arc_360_100m: degrees of horizon arc where a seed source is present within 100 m for 16 sectors (Figure 1 in manuscript)
- arc_W180_100m: degrees of horizon arc where a seed source is present within 100 m for 9 sectors (Figure 1 in manuscript)
- arc_W90_100m: degrees of horizon arc where a seed source is present within 100 m for 5 sectors (Figure 1 in manuscript)
- arc_top3_100m: degrees of horizon arc where a seed source is present within 100 m for top 3 sectors (Figure 1 in manuscript)
- arc_top5_100m: degrees of horizon arc where a seed source is present within 100 m for top 5 sectors (Figure 1 in manuscript)
- arc_360_165m: degrees of horizon arc where a seed source is present within 150 m for 16 sectors (Figure 1 in manuscript)
- arc_W180_165m: degrees of horizon arc where a seed source is present within 150 m for 9 sectors (Figure 1 in manuscript)
- arc_W90_165m: degrees of horizon arc where a seed source is present within 150 m for 5 sectors (Figure 1 in manuscript)
- arc_top3_165m: degrees of horizon arc where a seed source is present within 150 m for top 3 sectors (Figure 1 in manuscript)
- arc_top5_165m: degrees of horizon arc where a seed source is present within 150 m for top 5 sectors (Figure 1 in manuscript)
- forcov_360_100m: percent of sector occupied by seed source within 100 m for 16 sectors (Figure 1 in manuscript)
- forcov_W180_100m: percent of sector occupied by seed source within 100 m for 9 sectors (Figure 1 in manuscript)
- forcov_W90_100m: percent of sector occupied by seed source within 100 m for 5 sectors (Figure 1 in manuscript)
- forcov_top3_100m: percent of sector occupied by seed source within 100 m for top 3 sectors (Figure 1 in manuscript)
- forcov_top5_100m: percent of sector occupied by seed source within 100 m for top 5 sectors (Figure 1 in manuscript)
- forcov_360_165m: percent of sector occupied by seed source within 150 m for 16 sectors (Figure 1 in manuscript)
- forcov_W180_165m: percent of sector occupied by seed source within 150 m for 9 sectors (Figure 1 in manuscript)
- forcov_W90_165m: percent of sector occupied by seed source within 150 m for 5 sectors (Figure 1 in manuscript)
- forcov_top3_165m: percent of sector occupied by seed source within 100 m for top 3 sectors (Figure 1 in manuscript)
- forcov_top5_165m: percent of sector occupied by seed source within 100 m for top 5 sectors (Figure 1 in manuscript)
- wil_ht: height of willow canopy (1-ft height classes)
- wil_sur: percent of planted willow surviving in 10-tree row sections (average of 4)
- wil_cov: estimated percent canopy closure from densiometer
- bare_pct: percent bare soil (visual estimate per methods in manuscript)
- under_pct: percent ground cover (visual estimate per methods in manuscript)
- under_ht: predominant height (cm) of understory vegetation within the plot
- n_pct: percent nitrogen in bulk soil
- c_pct: percent carbon in bulk soil
- bd_b: bulk density grams per cubic cm of soil Ap horizon (top 20 cm)
- bd_f: bulk density of fine fraction in grams per cubic cm of soil Ap horizon (top 20 cm)
- crs_p: proportion of course materials by volume
- bd_parent: bulk density of course materials in grams per cubic cm
- por_b: ratio of voids:solids for soil Ap horizon (top 20 cm)
- por_f: ratio of voids:solids for soil Ap horizon excluding fines (top 20 cm)
- theta_g : gravimetric water content at time of sampling (grams water:grams fresh core)
- theta_v: volumetric water content at time of sampling (volume water:volume fresh core)
- water_mm: mm of water per hectare present in upper 20 cm of soil at time of collection
- N_mgha: nitrogen content in upper 20 cm of soil scaled to Mg/ha
- C_mgha: carbon content in upper 20 cm of soil scaled to Mg/ha
- lit_mgha: litter (dry weight 60C) present on soil surface scaled to Mg/ha
FILE: “ArithmeticMeans.RTF” (Draft of Tables 3, 4, and 5 in manuscript)
- Variable: Matches variable names in [FILE] “Study 2_Input_Dataset.CSV”
- Ord1: Corresponds to the “Low” seedling density plots in manuscript
- Ord2: Corresponds to the “Medium Low” seedling density plots in manuscript
- Ord3: Corresponds to the “Medium High” seedling density plots in manuscript
- Ord4: Corresponds to the “High” seedling density plots in manuscript
- Global: Average of all plots
FILE: “OrdinalSlopes.RTF” (Results of a variable screening step. In the SAS program “Study 2_fixed_master_cluster_final.SAS”, the first screening step was to use PROC GLIMMIX to determine if a slope exists for single-component models for variable across ordinal levels (1-4). This was a variable elimination step with the advantage that it only used 1 degree of freedom. Other screening steps followed using mean separations in the SAS program “Study 2_fixed_master_cluster_final.SAS”)
- DF: Degrees of freedom in ordinal model
- Response: Variable matching names in [FILE] “Study 2_Input_Dataset.CSV”
- Order: An iteration number
- TrendSig: Was a significant slope detected at the alpha = 0.05 level (see p-value in final column); if “Yes”, means separations were conducted, if “No” variable was not considered in logistic models.
- Slope: Slope of ordinal model
- SE: standard error of ordinal model
- T: t-score of ordinal model
- P: p-value of ordinal model
