Data from: Cybianthus anthuriophyllus (Primulaceae), a new record for the flora of Colombia: Distribution modeling and conservation assessment
Data files
May 04, 2026 version files 45.22 MB
-
Attached_files.zip
45.21 MB
-
README.md
9.74 KB
-
S1.csv
548 B
Abstract
Cybianthus anthuriophyllus (Primulaceae) was previously known from eastern Ecuador and northern Peru. Here we document the first confirmed occurrences of this species in Colombia and develop a distribution model based on collections from the Andean–Amazonian foothill of Caquetá, Cauca, and Putumayo departments, in order to clarify its conservation. Specifically, to refine its potential distribution and inform conservation efforts, we developed a species distribution model (SDM) using MaxEnt and assessed its extinction risk following IUCN Red List Criteria. The model identified high suitability in the northwestern Amazon, particularly along the Andean foothill, and response curves indicated that annual precipitation and isothermality were the primary predictors of habitat suitability. The species qualifies as Endangered (EN B2ab(iii)) due to its limited area of occupancy, few known locations, and ongoing habitat loss. These findings extend the distribution known range of Cybianthus anthuriophyllus, close the floristic gap in the subgenus Comomyrsine, highlight the conservation importance of the Andean–Amazon foothill corridor, and provide evidence for the continued need to prioritize biological inventories in this area.
Dataset DOI: 10.5061/dryad.cz8w9gjk6
Description of the data and file structure
This dataset includes all input data used for the Species Distribution Model (SDM) analyses. It contains occurrence records in CSV format compatible with MaxEnt, a shapefile of Napo Province sensu Morrone (2014) biogeography classification, and the climatic variables obtained from the WorldClim database (www.worldclim.org), which were clipped using the shapefile of area M as a spatial mask.
Additionally, we provide MaxEnt output files from model runs used during the training phase, as well as all output files from the final model.
Finally, we include a folder containing the input and output files for each iteration of the jackknife leave-one-out method of Pearson et al. (2007) with their respective executable file pValueCompute.exe used for this validation.
Files and variables
File: Attached_files.zip
Description: This dataset contains all input data, intermediate files, and outputs used for Species Distribution Model (SDM) analyses conducted with MaxEnt. The dataset is organized to ensure full reproducibility of the modeling workflow, including occurrence data, environmental predictors, spatial masks, and validation outputs.
All files are organized into three main folders:
1) Inputs for SDM
a) Occurrence data: The file “TOTALPOINTS.csv” contains species occurrence records formatted for use in MaxEnt.
b) Study area shapefile: The folder “shape_file_napo” contains the shapefile used as a spatial mask, corresponding to the Napo Province sensu Morrone 2014 biogeography classification, based on Löwenberg-Neto 2014 biogeography dataset.
c) Environmental variables: The folder “variables_shp_m” contains climatic predictors obtained from the WorldClim database (www.worldclim.org) at 30 arc-second resolution (~1 km²). These layers were clipped using the study area shapefile described above and represent the environmental inputs used in MaxEnt.
2) Final_SDM
This folder contains all output files generated by MaxEnt for the best-performing model. The files include spatial predictions, model summaries, performance metrics, and tabular outputs (CSV/Excel) that describe model behavior and results. The different types of files can be classified as follows:
a) Raster files (.asc)
- Format: ASCII raster grid
- Description: Spatial predictions of habitat suitability
- Values:
- Continuous values ranging from 0 to 1 (logistic output)
- Represent relative habitat suitability (dimensionless)
b) HTML files (.html)
- Description: Automatically generated MaxEnt reports
- Include:
- Model performance (AUC)
- Variable contributions
- Response curves
- Jackknife plots
c) Lambdas files (.lambdas)
- Description: Model parameter files
- Contain:
- Feature coefficients
- Regularization parameters
- Use: Required to replicate model predictions
d) Tabular outputs (.csv / Excel files)
- MaxEnt results table (e.g., maxentResults.csv or equivalent Excel file): Summary of model performance and variable importance metrics.
| Column | Description | Units |
|---|---|---|
species |
Name of the modeled species | — |
AUC.training |
Area Under the Curve for training data | Dimensionless (0–1) |
AUC.test |
Area Under the Curve for test data (if applicable) | Dimensionless (0–1) |
AUC.diff |
Difference between training and test AUC | Dimensionless |
regularizationMultiplier |
Regularization parameter used in the model | Dimensionless |
features |
Feature classes used (e.g., L, Q, H) | — |
contribution_[variable] |
Percent contribution of each environmental variable | % |
permutation_importance_[variable] |
Permutation importance of each variable | % |
- Samples predictions (e.g., samplePredictions.csv): Predicted suitability values for each occurrence (and background) point.
| Column | Description | Units |
|---|---|---|
species |
Species name | — |
longitude |
Longitude of the point | Decimal degrees |
latitude |
Latitude of the point | Decimal degrees |
prediction |
Predicted habitat suitability | 0–1 (dimensionless) |
presence |
Presence (1) or background (0) | Binary |
- Environmental data per sample (e.g., samplesWithData.csv): Environmental variable values extracted for each point used in the model.
| Column | Description | Units |
|---|---|---|
species |
Species name | — |
longitude |
Longitude | Decimal degrees |
latitude |
Latitude | Decimal degrees |
bio1, bio2, ... |
Bioclimatic variables (WorldClim) | See below |
Common WorldClim variables
| Variable | Description | Units |
|---|---|---|
bio1 |
Annual Mean Temperature | °C |
bio2 |
Mean Diurnal Range | °C |
bio3 |
Isothermality | % |
bio12 |
Annual Precipitation | mm |
(Note: Temperature variables from WorldClim may be scaled; verify if values are multiplied by 10.)
3) Test_Jacknifee
This folder contains input files generated during the validation phase using the jackknife approach described by Pearson et al. 2007.
-
The folder “filesCSV_for_test_Jackknife” contains CSV files used for iterative runs in which each occurrence point was excluded in turn.
Column name Description Interpretation Units iteration_without_ith_pointIteration number corresponding to the jackknife procedure, where each occurrence point is removed once Identifies which point was excluded in each model run — With_all_pointsPredicted suitability value at the focal point when the model is trained using all occurrence records Baseline prediction for comparison Dimensionless (0–1) Without_ith_point_deleted (Pi)Predicted suitability value at the same point when that specific occurrence is excluded from model training Measures the influence of that point on model prediction Dimensionless (0–1) success/failure (Xi)Binary indicator of prediction success based on a predefined threshold 1 = successful prediction (model correctly predicts the excluded point); 0 = failure Binary (0/1) -
The file “Input_data_for_Jackknife_test” contains the summarized data used for the jackknife validation.
Literature.
Löwenberg-Neto, P. 2014. Neotropical region: a shapefile of Morrone’s (2014) biogeographical regionalisation. – Zootaxa 3802: 300–300.
Morrone, J. J. 2014. Biogeographical regionalisation of the Neotropical region. – Zootaxa 3782: 001–110.
Pearson, R. G., Raxworthy, C. J., Nakamura, M., and Peterson, A. T. 2007. Predicting species distributions from small numbers of occurrence records: a test case using cryptic geckos in Madagascar. – J. Biogeogr. 34: 102–117.
File: S1.csv
Description: CSV file with the occurrences used for the SDM analysis.
Variables
- Name: species name
- Long: longitude
- Lat: latitude
Code/software
The data can be accessed and processed using the following free and open-source software:
- QGIS (version 3.34) was used for spatial data processing, including visualization, clipping of environmental layers, and preparation of the study area mask (area M). Raster layers (ASCII format, .asc) and vector data (ESRI shapefiles) can be opened and edited directly in QGIS.
- MaxEnt (version 3.3.3) was used to run the Species Distribution Models (SDMs). Occurrence data in CSV format and environmental layers in ASCII format were used as inputs. Model outputs (e.g., .asc, .html, .lambdas, .csv) can be viewed using MaxEnt or standard GIS software.
Access information
Data was derived from the following sources:
- GBIF. 2026. GBIF occurrence download for Cybianthus anthuriophyllus. https://doi.org/10.15468/dl.z3tufm
Occurrence data collection
During botanical explorations in the municipalities of Florencia (Caquetá department), Piamonte (Cauca department), and Villagarzón (Putumayo department) from Colombia, we recorded the presence of Cybianthus anthuriophyllus. Phenological observations were noted in the field with the support of a local collaborator (see Acknowledgements). To compile the full known distribution of the species, we reviewed the original protologue (Pipoly 1998), consulted specimens housed at COL, CUVC, HEAA, HUA, HUAZ, LAMUA, and UIS (following Thiers 2023), and gathered georeferenced records from the Tropicos database (www.tropicos.org) and GBIF (2026). After removing duplicate entries, records without specimens’ images available before 2023, and records with geographic uncertainty, we retained a set of 14 presence points used for subsequent analyses (S1).
Species distribution modeling (SDM)
We estimated the potential distribution of Cybianthus anthuriophyllus using the MaxEnt algorithm (version 3.3.3k), which is particularly suitable for presence-only data (Phillips et al. 2006). Climatic predictors were sourced from the WorldClim database (www.worldclim.org) at 30 arc-second resolutions (~1 km²). To avoid overfitting under a small sample size (n = 14), we implemented a structured variable selection procedure. Initially, Pairwise Pearson Correlations were calculated, and highly correlated variables (r ≥ 0.8) were excluded to reduce multicollinearity and inflation of variable importance. Consequently, four bioclimatic predictors with known ecological relevance to premontane humid forest were evaluated: isothermality (Bio3), temperature seasonality (Bio4), annual temperature range (Bio7), and annual precipitation (Bio12).
Model calibration used the following settings: 20 replicate runs, 10,000 background points (after testing lower background values proportional to presence points according to Rausell-Moreno et al. 2025), and a regularization multiplier of 1, which were selected after testing alternative values (0.5, 1.5, and 2) based on response curve evaluation. The accessible area M of the species (Barve et al. 2011) was defined as the Napo Province sensu Morrone (2014), and a polygon shapefile provided by Löwenberg-Neto (2014) was used as a mask layer to constrain the model geographically.
Given the small number of occurrence records, we evaluated model performance using the jackknife validation approach described by Pearson et al. (2007). This leave-one-out method assesses predictive ability by iteratively excluding each point and calculating whether it falls within the predicted suitable area of the model. We used the strictest omission threshold—Lowest Presence Threshold (LPT)—and calculated Pi (proportion of predicted area) and Xi (model success/failure) accordingly. Models including additional correlated predictors produced unstable response curves and did not improve predictive performance under jackknife validation. Therefore, the final model retained two ecologically interpretable and statistically stable variables: isothermality (Bio3) and annual precipitation (Bio12). Bio 3 captures the proportional oscillation between diurnal and annual temperature variation, which is particularly relevant in premontane environments where thermal stability may influence physiological tolerance and niche specialization (O’Donnell and Ignizio 2012). Whereas Bio 12 reflects overall water availability, a key driver in any terrestrial ecosystem (Gutiérrez-Hernández and García 2021). The combination of these two variables captures both thermal buffering and moisture availability, fundamental axes structuring in plant distribution.
Conservation assessment
We assessed the preliminary conservation status of Cybianthus anthuriophyllus based on the IUCN Red List Categories and Criteria proposed by the IUCN Standards and Petitions Committee (2024). We focused on Criterion B, which includes the Extent of Occurrence (EOO, subcriterion B1) and the Area of Occupancy (AOO, subcriterion B2). These metrics were calculated using the GeoCAT tool (www.geocat.kew.org), applying a 2 km² grid size for AOO, as recommended by the IUCN guidelines.
To evaluate the additional subcriteria (a), (b), and (c) required under Criterion B, we used spatial data layers on forest cover and land use obtained from national and global sources of DIVA-GIS (https://diva-gis.org), SIAC (https://www.siac.gov.co), MapBiomas project with the initiatives of Colombia (https://colombia.mapbiomas.org), Ecuador (https://ecuador.mapbiomas.org), and Peru (https://peru.mapbiomas.org). Fragmentation was inferred by overlaying known localities and the modeled distribution on recent forest/non-forest maps to assess habitat continuity and the degree of isolation between populations. Habitat degradation and ongoing decline were assessed based on recent deforestation trends and land use change across the Andean–Amazonian foothills. The number of locations was estimated using the precautionary principle, considering the spatial clustering of records and potential shared threats.
