Data and code from: Fire-survival strategies of first-year acacia seedlings
Data files
Apr 20, 2026 version files 1.49 MB
-
Acacia_Garden_revised_for_archiving.zip
1.45 MB
-
README.md
33.91 KB
Abstract
Savannas—defined by discontinuous tree cover and continuous grass cover—are maintained by factors such as drought, fire, and herbivory which impose bottlenecks on tree populations. Despite evidence that young plants are particularly vulnerable to fire, surprisingly little is known about the traits that enable first-year seedlings to survive fire. This limits our understanding of the mechanisms underlying tree establishment and savanna dynamics more generally.
We conducted a full factorial common garden experiment to measure the separate and combined impacts of fire, drought, and herbivory on rates on three- to six-month-old seedlings of 12 ‘Acacia’ (sensu lato) species in the genera Vachellia and Senegalia.
Results demonstrate an overwhelming impact of fire on mortality, enhanced by drought, but the impact of herbivory was not detectable. Stem thickness was the best predictor of topkill and survival, with threshold diameters of 9.56 mm (95% CI 9.51–9.59) for well-watered and 11.52 mm (95% CI 10.88–12.69) for droughted plants, above which seedlings had favorable odds of avoiding topkill by fire and surviving.
The ability of seedlings to survive after being topkilled by fire varied by species and was linked to the drought treatment and the existence of stem bases belowground. Thus, fire represents a major demographic bottleneck for first-year acacia seedlings, even when seedlings are established and have sufficient access to water.
Synthesis: Acacia seedlings (eleven of twelve species) acquire some ability to survive grass fires in their first year, either by resisting topkill with thickened stems and/or by regrowing from below-ground stems. Our results suggest that small differences in these fire-survival strategies at this crucial life stage translate to large differences in savanna physiognomy and species composition.
Dataset DOI: 10.5061/dryad.15dv41p9z
Description of the data and file structure
Prepared April 14, 2026
Data derive from the common garden experiment on Acacia trees at Nelson Mandela - African Institution of Science and Technology in Arusha, Tanzania. Seeds were collected in 2022 and the data were collected in 2023 and 2024. All field data derive from the first (“main”) experiment.
Users should consult the published article for methodological and other details. Data are available without restrictions; as a courtesy, we ask that users alert the first and last authors and cite the published article in any works arising from reuse of these data.
For more information, contact both A.B. Potter (abp47@cornell.edu) and T. Michael Anderson (anderstm@wfu.edu).
Files and variables
File: Acacia_Garden_revised_for_archiving.zip
Description:
The root project folder “Acacia_Garden_revised_for_archiving” contains all raw data and code to reproduce the analyses. It is a complete and real project folder, aimed at reproducing research as performed. As such, the numerous redundancies, naming and formatting discrepancies, and inefficiencies are not corrected or deleted, but rather presented as is.
The root folder contains three broad folders:
- data
- R
- working_copy_data
In brief,
- "data" contains all raw data. These are extensively documented in this README.
- The folder "R" contains all R code, which contains extensive documentation.
- The folder "working_copy_data" contains cleaned data files derived from the raw data. These data are redundant. Its sole purpose is to enable the analyses to be run without having to clean the data each time.
Contents of all three broad folders follows.
Contents of each broad folder
Folder "data"
The folder “data” contains uncleaned data files. These represent all raw data used. These were generally typed from paper datasheets. These are organized into the following five subfolders:
- Herlocker_table
- JacobTypedCopySep17
- misc_data
- TMA_environmental_data
- WinnieTypedCopyAug26
Subfolder 'Herlocker_table'
This contains one file, 'Habitat_affiliations_Herlocker.csv' created by ABP.
File 'Habitat_affiliations_Herlocker.csv'
These are data generated from the following publication:
Herlocker, D. J. (1974). Woody Vegetation of the Serengeti National Park. Caesar Kleberg Research Program in Wildlife Ecology and Department of Range Science, Texas A & M University System.
Variables:
Species: The species of Acacia tree, as per the taxonomy used in the Potter et al. manuscript.
Taxonomy used in Herlocker (1974): The name of Acacia tree, as per Herlocker.
Topography and soil (Herlocker 1974): Herlocker's description of where each Acacia grows.
Vegetation formations (Herlocker 1974): Herlocker's classification of the major plant communities of the Serengeti.
Subfolder 'JacobTypedCopySep17'
This is a folder containing the single datasheet typed up by Jacob Kamana.
File 'datasheet_dig_measure_entering (MIGUNGA).csv'
This datasheet contains data on Acacia seedling in the main experiment collected while these seedlings were being excavated at the end of the experimental period. Empty cells represent no data, due to (1) the data not being applicable (e.g. mud, Zea mays) (2) the data being correctly missing due to the acacia seedling having decomposed or burned (3) the data being incorrectly missing due to errors in data collection, a possibility that is unlikely but cannot be discounted. Existing R code is equipped to handle these empty cells.
Variables:
Block: The Block within the main experiment.
Subplot: The subplot within each block the main experiment
Position: The position within each subplot in the main experiment
Species: Species of AcaciaDig
date: The date this seedling was excavated dd-mm-yyyy
Stem green?: Whether (yes=1, no=0) a seedling had a visibly green or living stem.
Roots alive?: Whether (yes=1, no=0) a seedling had a visibly living root.
Root into another subplot?: Whether (yes=1, no=0) a seedling had a root that went into another subplot from the one in which it was planted or if it went out of the experiment into a drainage ditch.
Nodules present?: Whether (yes=1, no=0) a seedling had root nodules, suggested N-fixing ability.
Root exposed?Ê: Whether (yes=1, no=0) a seedling had the top-most portion of its root exposed.
Depth_RootCollar_mm: The depth of the junction between the root and the shoot in millimeters.
Root Collar_diam_mm: The diameter of the junction between the root and the shoot in millimeters.
Diam_at_10cmdepth_mm: The diameter of the root at 10 centimeters below the soil surface.
Diam_at_30cmdepth_mm: The diameter of the root at 30 centimeters below the soil surface.
WholeRoot_max_depth_cm: The depth of the deepest part of the taproot, given that the entire taproot was excavated.
WholeRoot_end_Diam_mm: The diameter of the deepest part of the taproot, given that the entire taproot was excavated.
BrokenRoot_dug_depth_cm: The depth of the deepest part of the taproot excavated, given that the taproot was broken and thus the true depth was deeper.
BrokenRoot_end_Diam_mm: The diameter of the deepest part of the taproot excavated, given that the taproot was broken.
Notes: empty field
Subfolder 'misc_data'
The subfolder contains four miscellaneous sources of data.
Seedlings were germinated in three cohorts and the seedlings in the main experiment represent a mix of cohorts. Unfortunately, germination data are unavailable. Fortunately, notes were taken on which seedlings were transplanted where and when in the main experiment, and these notes record information on plastic tags that were used to track seedlings. These notes were incomplete and there were many minor changes of plan (e.g. seedling mortality after transplanting), which sometimes forced us to replace seedlings of one cohort with seedlings of another cohort. After the fact, we examined photos taken during the transplanting process that corroborated these notes to determine the cohort of each seedling that was ultimately retained in the experiment. The high redundancy of our notes and photos makes up for the incompleteness of our records. As such, we can confidently piece together the missing data on germination cohorts.
File 'Cohort_block_column_row_indivdual_Douglas.xlsx'
Three tabs in this uncleaned spreadsheet describe the cohort (1, 2, or 3) that each seedling in the experiment belonged to, using images and notes.
Empty cells represent no data, due to (1) the cohort not being applicable (species = NA, Zea mays) (2) the data being correctly missing due to a plant tag having been lost or rendered illegible and no notes of photographic evidence that would suggest which cohort a seedling belonged to. (3) the data being incorrectly missing due to errors in data collection, a possibility that is unlikely but cannot be discounted. Existing R code is equipped to handle these empty cells.
Tab: Cohort_information_bydate
Block: Block in the main experiment.
Subplot: Subplot in the block.
Position: Position within the subplot
Species: Species of Acacia seedling.
Remaining columns: The cohort to which a seedling belonged, as determined by (1) Images with dates (2) Notes taken at the time of transplant (3) Notes taken by Ellen W.
Tab: Cohort_with_Notes
As in the previous tab "Cohort_information_bydate" except that this tab contains stray notes, such when tags were missing and the seed stock used (e.g. VACSEY5 is the 5th seed collection of Vachellia seyal var. fistula).
Tab: Ellen_7_8_23_Chrt_Blk_Clmn_Row
Ellen collected these data by walking through the main experiment in order and noting down any information on plastic tags. Each of the rows corresponds to a position in the experiment (Code), following an identical order to all other data sheets (1.1.1 to 10.9.6).
Ellen 7/28/23: A determination made by Ellen W. on July 28 2023 as to which cohort a seedling belonged, based on any tags visible.
Block: The Block of germination tubes in the germination shelter.
Column: The column of tubes in a block.
Row: The Row of tubes in a column.
Individual: The seed source from which a seedling was collected.
Notes: Notes on the tag.
As an example, the first line of data in this spreadsheet tells us that the seedling at 1.1.1 (block 1, subplot 1, position 1) belonged to cohort 1 where it was planted in block 11 in the shade house in column 6 and row 3 and that this seedling (vacsie, I remember) derived from the first set of seeds of this species collected (individual 1).
File 'Compiled_cohort_information_Douglas_18Dec2024.csv'
This compiled spreadsheet is a cleaned version of 'Cohort_block_column_row_indivdual_Douglas.xlsx' . Each column describes redundant information on the cohort (1, 2, or 3) that each seedling in the experiment belonged to, using images and notes. This information was used to simply identify the cohort (1, 2, or 3) to which each seedlings belonged, which is part of basic experimental design.
Variables:
As in tab "Cohort_information_bydate" in file "Cohort_block_column_row_indivdual_Douglas.xlsx"
The first three variable identify the position or seedling within the main experiment, and the remaining tabs identify the source of information and cohort determination.
File 'datasheet_template_treatments.csv'
This file contains the intended experimental design, which was generated in silico.
Variables:
(First column which is empty): serial number 1-540 representing a unique spot in the experiment.
number: serial number 1-540 representing a unique spot in the experiment.
Block: The Block within the main experiment.
Subplot: The subplot within each block the main experiment
PlantPosition: The position within each subplot in the main experiment
Species: The species of Acacia that was intended to be planted.
Fire: The intended treatment with fire (Fire or No Fire)
Drought: The intended treatment with drought (Drought or NoDrought)
Herbivory: The intended herbivory treatment (Herbivory or NoHerbivory)
identity: The intention of whether a spot in the experiment is to be plant
GrassCompetition: The intended level of grass competition, which is "NoGrass" for the entire experiment (that is, grass was prevented from establishing). This column is included as a redundancy to distinguish the main experiment from a simultaneous experiment studying the effects of grass competition.
File 'Postfire_score2'
This spreadsheet records observations on seedlings in the weeks following the application of the fire treatment.
Variables:
Block: The Block within the main experiment.
Subplot: The subplot within each block the main experiment
Position: The position within each subplot in the main experiment
Species: The species of Acacia.
Green_Leaves: Whether or not (yes=1, no=0) the seedling had living leaf tissue.
Green_stem: Whether or not (yes=1, no=0) the seedling had living stem tissue.
Appears Topkilled by Fire in Photo: Whether or not (Topkilled, NotTopkilled, Unclear) the seedling looked like it was topkilled in photos taken after the fire or whether the photo evidence was unclear. Empty cells are not applicable (not subjected to the fire treatment). Existing R code is equipped to handle these empty cells.
Postfire_Regrowth: Whether or not (yes=1, no=0) the seedling appeared to be regrowing after the fire. Empty cells are not applicable (not subjected to the fire treatment). Existing R code is equipped to handle these empty cells.
Notes: Any observations of plants exposed to fire. Empty cells indicate that no notes were taken for this line of data (missing data, because not noteworthy).
Time: An arbitrary time estimate (17:40) associated with the determinations of green leaves and green stems.
Date_for_GS_GL: The date associated with the determinations of green leaves and green stems DD-MM-YY
(untitled column): Additional observations of plants exposed to fire.
14_Feb_Judgement: A single topkill determination made of Feb 14.
Topkill_from_photo_and_data: A combined judgement of whether or not a plant was topkilled, given both evidence from photos and from data.
(untitled column): Notes on how the judgement of whether or not a plant was topkilled was reached. Empty cells indicate that no notes were taken for this line of data (missing data, because not noteworthy).
Subfolder 'TMA_environmental_data'
File 'Env.data.seed.collections.csv'
Environmental conditions at the GPS locations were seeds used to germinate seedlings were collected, as determined from existing spatial datasets from the following manuscript:
Anderson, T. M., G.P. Hempson, J.E. Donaldson, C.M. Beale, M. te Beest, C. Courtney-Mustaphi, J.P.G.M. Cromsigt, C. Foy, R. Fynn, N. P. Hanan, C.L. Parr, J. Probert, L. le Roux, K. Sianga, I.P.J. Smit, A.C. Staver, & S. Archibald. (2025). Identifying ecological knowledge and research gaps via the African Database on Savanna Protected Areas (ADSPA). Diversity and Distributions, Revision Invited.
Variables:
True.identification: The correct scientific name of the Acacia species (and variety, if applicable), as consistent with the manuscript.
Labeled.Species: The scientific name recorded at the time of seed collection. This name is sometimes taxonomically ambiguous or incorrect, given that these were preliminary field identifications.
X_Coord: The latitude at which seeds were collected. Presumably in WGS84.
Y_Coord: The longitude at which seeds were collected. Presumably in WGS84.
MAT: The mean annual temperature in Celsius.
MAP: The mean annual precipitation, in millimeters.
Soil.N Soil nitrogen content in ppm.
Soil.P: Soil phosphorus content in ppm.
DEM: Elevation in meters
ETo: Reference Evapotranspiration in millimeters
SAND: Percentage sand
Subfolder 'WinnieTypedCopyAug26'
This folder contains the most important data files of the experiment. There are 41 data files, representing raw data. Of these 41, 34 files are typed up versions of weekly measurement datasheets and 7 files are miscellaneous data files. The 34 weekly measurement datasheets are so similar that they are treated as a group of files.
File: Cohort_block_column_row_indivdual.csv
A version of a tab in File 'Cohort_block_column_row_indivdual_Douglas.xlsx' .
Because of record keeping errors, we had to piece together which seedling in the experiment came from which exact spot in the shade house where germination was conducted, which in turn came from a particular seed collection event. This was done when Ellen W wrote down information on plastic tags that accompanied each seedling from the shade house to the main experiment field. Collecting these data was not always possible. As such, these data are incomplete.
Empty cells in variables "Cohort", "Block", "Column", and "Individual" represent no data, due to (1) the data not being applicable (e.g. Species =NA, Zea mays) (2) the data being missing because no information was readily available due to lack of notes or damage to the tags during transplanting, loss of the tags ("missing tag" in notes), and photodegradation of the plastics rendering them illegible. Existing R code is equipped to handle these empty cells.
As an example, the first Vachellia sieberiana from which seeds were collected was Vachellia sieberiana 1, and seeds from this collection were planted in the 11th block in the 6th column and 3rd row of the shade house as part of the 1st cohort of germination. The vacsie seedling growing in this tube was eventually transplanted into the main experiment in block 1, subplot 1, position 1 by Arjun Potter.
Variables:
Block: The block in the experiment.
Subplot: The subplot in the experiment.
Position: The position (1-6) in the experiment.
Species: The species.
Cohort: Which germination cohort each seedling belongs to (1, 2, or 3).
Block (the second one): The "block" of plastic tubes in the shade house where seeds were germinated.
Column: The column of plastic tubes, arranged as a grid, in a block in the shade house.
Row: The row of plastic tubes, arranged as a grid, in a block in the shade house.
Individual: The seed source from which each seedling was collected. E.g. 1 means the "Vacsie 1" in the seed collection datasheet. Seedlings with the same individual number were collected at the same time and place (usually from the same tree).
Notes: Notes on the tag or plant. Empty cells indicate that no notes were taken for this line of data (missing data, because not noteworthy).
File: Fire Datasheet.csv
This datasheet describes the fire treatment conducted in the main experiment.
Empty values represent missing data due to errors in data collection. Specifically, the empty cells in "T1.Avgtemp.C" and "T2.Avgtemp.C" were due to a failure of a thermocouple.
Variables:
Block: The Block in the main experiment. "Pilot" is a test burn conducted away from the experiment.
Subplot: The Block in the main experiment. "Pilot" is a test burn conducted away from the experiment.
Wt.AirDryGrass.kg: The air-dry weight of grass fuel added, in kilograms.
GrassSampleTaken: Whether or not a sample of this grass was taken.
Time: The time of the burn.
Date: The date of the burn dd/mm/yy.
PhotoBefore: Whether a photo was taken before or not.
H2OTempBef.C: The temperature of water in a small metal container before the fire was lit.
AirTemp.C: The air temperature.
T1.Maxtemp.C: The peak temperature of the fire measured by thermocouple 1, at 5 cm above the soil surface.
T2.Maxtemp.C :The peak temperature of the fire measured by thermocouple 2, at 15 cm above the soil surface.
T1.Avgtemp.C: The average temperature of the fire measured by thermocouple 1, at 5 cm above the soil surface.
T2.Avgtemp.C :The average temperature of the fire measured by thermocouple 2, at 15 cm above the soil surface.
H2OTempAft.C: The temperature of water in a small metal container after the fire finished burning.
VideoFire : Whether a video was taken of the fire.
PhotoAfter : Whether a photo was taken after the fire.
PhotoWithColorCard: Whether a photo of the ash was taken with a color correction card.
Collect.ResidualBiomass: Whether unburnt biomass was collected.
Flags: An internal reminder to remove any flagging tape or flags prior to ignition.
Notes: Qualitative notes on the fire.
File: Grass Sample Weights.xlsx
This raw excel file was used to calculate the moisture content of the air-dried grass fuel used in the fire treatment. The corresponding cleaned csv version is "Grass_fuel_samples_main_experiment.csv" and contains the same data.
Variables
S/N: Serial number of bag.
Date: Date at which grass fuel sample was harvested dd/mm/yy.
Block: Block in the main experiment from which grass fuel was sampled.
Wet weights for grass samples with medium bags: The weight of the sample of the air-dried grass fuel, in grams, including the weight of one medium-sized brown paper bag.
Dry weights for grass samples with medium bags: The weight of the oven-dried grass fuel, in grams, including the weight of one medium-sized brown paper bag.
File: Grass_fuel_samples_main_experiment.csv
This cleaned csv file was used to calculate the moisture content of the air-dried grass fuel used in the fire treatment. The corresponding raw excel file is "Grass Sample Weights.xlsx" and has the same data.
Variables
S/N: Serial number of bag.
Date: Date at which grass fuel sample was harvested yyyy-mm-dd.
Block: Block in the main experiment from which grass fuel was sampled.
Wet weights for grass samples with medium bags: The weight of the sample of the air-dried grass fuel, in grams, including the weight of one medium-sized brown paper bag.
Dry weights for grass samples with medium bags: The weight of the oven-dried grass fuel, in grams, including the weight of one medium-sized brown paper bag.
File: Herbivory.sheet_data_entry.csv
This datasheet documents the intended and actual implementation of the herbivory treatment, in terms of the which plants had how many stems cut and at what lengths and how much mass was removed. An attempt was made to separate leaves from stems, but this was imperfectly done. An attempt was also made to weigh removed stem material both while fresh and also after drying in the oven. Due to the variable sizes of acacia seedlings, the amounts removed varied widely and as such there sometimes multiple paper bags of tissue removed. Similarly, it was sometimes easiest to weigh tissue in the paper bag and sometimes easiest to remove it from the bag when weighing. These complications result in several, near-identical variables.
Empty cells have several meanings. Empty cells represent not applicable data in the cases where Herbivory = NoHerbivory or where Species is NA or Zea mays. Exceptions would indicate a deviation in the experimental protocol. In other cases, empty cells mean that there is no data for that variable for that line of data, which could be due to (1) there actually being no data for this value (not applicable) or (2) an error in data collection and thus a missing value (not available). Because these possibilities are difficult to distinguish or discount, we did not infill NA but kept these values empty. Existing R code is equipped to handle these empty cells.
Variables:
(Blank column): This is the Block in the experiment.
Subplot: This is the subplot in the experiment.
PlantPosition: The is the position of the plant (1:6) in its subplot.
Species: The species of plant.
Herbivory: Whether the plant was assigned to the Herbivory treatment or not.
Stem.LengthW4: The stem length in Week 4, in cm.
Stump.Length.Goal: The length of the stem, in cm, intended to remain after clipping for the herbivory treatment.
Actual.Stump: The actual length of the stem, in cm, remaining following the herbivory treatment. This factors in deviations from the ideal.
N.stems.cut: The number of stems that had to be pruned to follow the herbivory treatment.
Dead.included: Whether or not dead stem material was removed.
N.bags: The number of paper bags of stem material.
Wet.mass.leaf.and.stem.with.bag: Fresh mass in grams removed. Paper bag included.
Wet.mass.leaf.and.stem.no.bag: Fresh mass in grams removed.
Wet.mass.leaves.with.bag: Fresh mass in grams removed. Paper bag included.
Wet.mass.stems.with.bag: Fresh mass in grams removed. Paper bag included.
Extra.wet.bag: Fresh mass in grams removed. Paper bag presumably included.
Dry.mass.leaf.and.stem.with.bag: Oven-dry mass in grams removed. Paper bag included.
Dry.mass.leaf.and.stem.no.bag: Oven-dry mass in grams removed.
Dry.mass.leaves.with.bag: Oven-dry mass in grams removed. Paper bag included.
Dry.mass.stems.with.bag: Oven-dry mass in grams removed. Paper bag included.
Extra.dry.bag: Oven-dry mass in grams removed. Paper bag presumably included.
Comments: Comments on this sample.
File: Objects_Weighing_NM-AIST.xlsx.csv
This datasheet describes the weights and sizes of certain objects that were frequently used. In particular this table has information on paper bags used for weighing purposes, which simplifies the process of subtracting bag weight from data, allowing samples to be weighed within paper bags.
Variables
Object_Name: The name given to each object.
Wet_Weight_grams: The air-dry or normal weight in grams.
Dry_Weight_grams: The oven-dry weight in grass of the same object.
Number_of_items_in_Object_Name: The number of items in each object. Generally, a bunch of 10 paper bags was weighed at one time, and that bunch was treated as an object. This was done to average over variations in paper bag size.
Length_cm: The length of a single item of an object, in centimeters.
Width_cm: The width of a single item of an object, in centimeters.
File: Residual biomass weights.csv
This datasheet was used for quantifying the unburnt fraction in each burn in the fire treatment.
Variables:
Block: The block in the experiment.
Subplot: The subplot within the block.
Residual.bag.number: The bag number, used to distinguish multiple bags of unburnt material from a single fire treatment.
Oven-dry.with.medium.size.bag: The oven-dry weight in grams of all unburnt grass, charcoal, and ashes remaining after the grass fuel was ignited within a subplot.
Files: 34 similar weekly measurement datasheets beginning with either the word "Acacia" or the word "Week"
These 34 weekly measurement datasheet have file names beginning with either the word "Acacia" or the word "Week". These files contain regular weekly measurements of parameters in the main experiment. These files contain 31 weeks' worth of data, represent every week from "Week0" to "Week30" inclusive. Week 1 represents the first week of the drought treatment. Week 0 represents measurements prior to the start of the drought. The first three weekly datasheets (Weeks 0, 1, and 2) were originally typed up as xlsx files. The remaining weekly datasheets (Weeks 3 until 30) were typed up as csv files. Our R code reads in the three xlsx files and retroactively writes them into csv files, but we decided to archive the original xlsx files in additon. For this reason, there are 34 files to represent 31 weekly datasheets.
Variables:
These weekly measurement all have a very similar, but not identical, set of variables. The naming of variables likewise differs from datasheet to datasheet. Rather than listing the variables for each of the 34 files individually, which would be very redundant, it is more efficient to describe all of them in one place. The file "working_copy_weekly.csv" is a single file that merges all weekly measurement sheets. This file is found in the "working_copy_data" folder. Please see that section of this README document for a concise explanation.
In addition, the variable names, their units, and their meanings are documented in the two accompanying R scripts, particularly in "01_clean-raw-data.R", as well as in the main text of the manuscript.
Folder "R"
The folder “R” contains two R scripts, which constitute all code to clean and analyze data. File paths and directories may need to be modified as needed.
- The R script “01_clean-raw-data.R” inputs uncleaned data from the “data” folder and writes clean data files into “working_copy_data” . This R script is primarily focused on data cleaning, particularly detecting and fixing errors and standardizing formats.
- The R script “02_weekly_analyses_improved.R” inputs both uncleaned data from the “data” folder and cleaned data from “working_copy_data” folder. Code will output figures and is set up to put these into a “figs” folder in the root directory (a folder with this name can be created). This R script contains all analyses.
Folder "working_copy_data"
The folder “working_copy_data” contains cleaned intermediate versions of the same raw data in the "data" folder. Its purpose is to enable running “02_weekly_analyses_improved.R” without having to run “01_clean-raw-data.R” . That is, it exists to save time.
Since the data contained herein are wholly derived from the raw data, the files in "working_copy_data" can be deleted and re-created from scratch using the R code. Existing R code that uses the same directory structure will overwrite (and thus refresh) these cleaned data files each time the code is run.
Variable names generally follow those in the corresponding raw data files in the "data" folder. Minor discrepancies are documented in the accompanying R code.
File: fire_datasheet_cleaned.csv
This is simply a cleaned version of the raw data file "Fire Datasheet.csv".
Variables:
Variable names are the same as in "Fire Datasheet.csv" (in the "data" folder)
File: postfire_topkill_score_cleaned.csv
This file indicates which seedlings were topkilled by the fire treatment.
Variables:
Block: The block in the experiment (1:10)
Subplot: The subplot in each block (1:9)
Position: The position within each subplot (1:6)
Topkill_from_photo_and_data: A categorical variable derived from several sources of data that indicates whether all aboveground portions of a seedling were killed by the fire treatment ("topkilled") or whether some aboveground portions survived ("not topkilled).
Code: A unique identifier for each seedling or physical location in the experiment that combines the block, subplot, and position.
File: working_copy_roots_for_James.csv
This is simply a cleaned version of the raw data file "datasheet_dig_measure_entering (MIGUNGA).csv" (in the "data" folder).
Variables
Variable names and data follow the corresponding raw data file.
File: working_copy_treatment_timeline_for_James
This is a table of dates for events associated with seedlings.
Variables:
Cohort: The cohort (1, 2, or 3) to which a seedling belonged.
Code: A unique identifier for each seedling or physical location in the experiment that combines the block, subplot, and position.
Block: The block in the experiment (1:10)
Subplot: The subplot in each block (1:9)
Position: The position within each subplot (1:6)
Fire: Whether or not a physical location, and thus the seedling present, was subject to the fire treatment.
Drought: Whether or not a physical location, and thus the seedling present, was subject to the drought treatment.
Herbivory: Whether or not a physical location, and thus the seedling present, was subject to the herbivory treatment.
identity: Whether the physical location (code) represents a regular Acacia seedling ("plant") or a low-density control seedling ("lonecontrolplant" or "seyalcontrol") or a maize plant ("maizecontrol") or an empty bare spot ("mud").
density: Whether a code was in a high-density subplot, with 6 seedlings, or a low-density subplot, with 2 seedlings and 1 maize.
Acacia: If the code represents an Acacia seedling.
Date_Time_for_Excavation: The date and time at which the Acacia seedling was excavated YYYY-MM-DD.
Ultimate_Fate: Whether or not an Acacia seedling was dug up ("Exhumed") or not ("Decomposed").
burn_unit: The unit at which the fire treatment was conducted, specifically subplots within blocks.
Date_Time_for_Fire_Treatment: The date and time at which the fire treatment was implemented YYYY-MM-DD.
Date_Time_for_Herbivory_Treatment: The date and time at which the herbivory treatment was implemented YYYY-MM-DD..
Date_Time_for_Drought_Treatment: The date and time at which the herbivory treatment was first implemented (and continuously implemented thereafter through the end of the experiment) YYYY-MM-DD..
File: working_copy_weekly.csv
The contains the most important data in the main experiment: cleaned and compiled weekly observations of seedling growth and mortality. This csv is produced by the R script “01_clean-raw-data.R". These data are a cleaned and combined version of the separate raw weekly measurement datasheets (34 similar weekly measurement datasheets beginning with either the word "Acacia" or the word "Week") , which are in the "WinnieTypedCopyAug26" subfolder of the "data" folder.
Variables:
Block: The block in the experiment (1:10)
Subplot: The subplot in each block (1:9)
Position: The position within each subplot (1:6)
Species: The species of Acacia
Stem_Diameter_mm: The diameter of the seedling's main stem at 1 cm above the ground, in millimeters.
Green_Leaves: The preliminary, field-based determination as to whether an Acacia had green leaves.
Notes: Any notes on that plant at that week.
Datasheet: Which raw weekly datasheet produced this line of data.
WeekNo: The week number in the experiment (week 1 is the first week of drought).
Height_cm: The height of the seedling in centimeters.
Soil_VWC_Percent: The volumetric water content of the soil near the seedling, probing down into the surface, expressed as a percentage.
Stem_Length_cm: The length of the main stem in centimeters.
Post_Fire_Regrowth: The presence of new growing sprouts following the fire treatment.
Date_For_Regrowth: The date at which post-fire regrowth was noticed.
Code: A unique identifier for each seedling or physical location in the experiment that combines the block, subplot, and position.
Acacia: Whether the line in the data represents an Acacia seedling or not.
Genus: The genus of Acacia tree (two genera, either Vachellia or Senegalia)
Subgeneric: The clade at a rank between genus and species, grouping similar species but splitting the genera.
Date_Time_for_VWC: The date and time at which VWC data was collected.
Date_Time_for_PH: The date and time at which plant height data was collected.
Date_Time_for_SD: The date and time at which stem diameter data was collected.
Date_Time_for_SL: The date and time at which stem length data was collected.
GL: Whether or not an Acacia seedling had green leaves. This is a cleaned data variable that fixes errors in the original "Green_Leaves" variable, using a protocol described within the R code.
Green_Stems: Whether or not an Acacia seedling had green stems. This is a cleaned data variable that is derived from other variables and notes.
Still_Alive: Whether or not an Acacia seedling is likely to be alive. This is a cleaned data variable that integrates several lines of information using a protocol described in the R code.
Flag: A binary variable indicating whether a line in the data was worth scrutinizing (1 = flagged, 0 = no obvious problem). This subjective variable was created and used during data cleaning only.
Code/software
All analyses were conducted in R version 4.4.2 (2024-10-31) -- "Pile of Leaves"
