Shared and unique patterns of trait expression in a sulfide spring fish population in Guanacaste, Costa Rica
Data files
Jul 24, 2026 version files 29.63 MB
-
Costa_morph_dv.Rmd
15.56 KB
-
counts_matrix_cr_countreadpairs.txt
25.75 MB
-
CR_2025_LOE.xlsx
13.96 KB
-
cr_env_data.xlsx
9.68 KB
-
CR_RNA_2015_final.Rmd
37.48 KB
-
Final_Total_Geomorph.xlsx
670.90 KB
-
links.xlsx
9.14 KB
-
LOE_cr.Rmd
11.26 KB
-
README.md
5.89 KB
-
samples.xlsx
11.19 KB
-
SL_CR_2015.xlsx
14.58 KB
-
Table_S2-Poecilia_mexicana_Annotations_1.csv
3.09 MB
Abstract
Questions about the repeatability and predictability of evolution have long fascinated scientists. Organisms exposed to similar habitats are generally expected to exhibit similar traits due to natural selection but may instead show unexpected traits due to other forces. Organisms in sulfide spring habitats provide clear examples of convergent evolution. These springs test the limits of survival as they are rich in toxic hydrogen sulfide (H2S) and severely hypoxic. Despite these highly challenging conditions, members of the family Poeciliidae have repeatedly colonized toxic sulfide springs, convergently evolving similar life history, physiological, and morphological traits. Here, we quantify these same traits in a population of Poecilia gillii from a sulfide spring in Guanacaste, Costa Rica and compare this species’ morphological and physiological traits related to sulfide tolerance with a nearby non-sulfidic population and with well-studied populations of P. mexicana from southern Mexico. Like sulfidic P. mexicana, the Costa Rican sulfide spring fish had larger heads, a more elongated abdomen, reduced standard length, and increased expression of genes associated with H2S detoxification, sulfur transport, and the cellular response to hypoxia. Previous studies indicate that standing genetic variation is an important driver of convergent evolution of traits in the sulfide spring fishes, but our study highlights how similar selection gradients can lead to predictable adaptive evolution even when ancestral standing genetic variation is unlikely.
Dataset DOI: 10.5061/dryad.zw3r228ng
Description of the data and file structure
This dataset includes measures of environmental data, body size, body shape, hypoxia tolerance, hydrogen sulfide tolerance, and gene expression of Poecilia gillii from a sulfidic and non-sulfidic habitat.
Files and variables
File: Costa_morph_dv.Rmd
Description: This R script was used to analyze body size and body shape.
File: CR_2025_LOE.xlsx
Description: This dataset includes information on the hypoxia and hydrogen sulfide tolerance of P. gillii from sulfidic and non-sulfidic habitats. Missing values are indicated using blank cells.
Variables
- Date: Date of Trial
- Trial: Trial type (Hypoxia or Sulfide)
- Fish: Fish ID number
- Habitat: Habitat of origin (S= sulfidic, NS= non-sulfidic)
- LOE: time until loss of equilibrium (s)
- temp: Water temperature of trial measured in degrees Celsius
- SL_cm: Standard length (measured from snout to caudal peduncle in cm)
- Mass: Mass measured in g
- Survival: Whether fish lasted the entire time of the trial without losing equilibrium 0= lost equilibrium, 1= did not lose equilibrium
- Sex: Sex of fish, Male or Female
File: counts_matrix_cr_countreadpairs.txt
Description: Counts matrix used for analysis of gene expression. Columns 2-6 are removed before analysis. Column 1 refers to the gene ID, the remaining columns are all individual fish. Each row in a column refers to the counts for that gene.
File: cr_env_data.xlsx
Description: Environmental data collected in 2015 and 2024. Missing values are indicated using blank cells.
Variables
- Site: Site name (Middle Earth or Tribute to Temp 1)
- DO: dissolved oxygen measured in mg/L
- Cond: conductivity measured in µS/cm
- pH: pH
- Temp: Water temperature measured in °C
- Date: Sampling date
- lat: site latitude
- long: site longitude
- Sulfide: Sulfide measurements
File: CR_RNA_2015_final.Rmd
Description: This R script was used to analyze gene expression of sulfidic and non-sulfidic fish.
File: Final_Total_Geomorph.xlsx
Description: This Excel file includes body shape data. This datasheet is split into two sheets (Core Male and Core Female) and the description of the variables is the same for each sheet.
Variables
- Specimen: Specimen ID
- Genus: Genus of the fish
- Species: Species of the fish
- Lineage: Lineage of the fish (Species+ whether it came from a sulfidic or non-sulfidic environment)
- H2S: Habitat type, 1= sulfidic, 0= nonsulfidic
- Site: Site name
- Sex: Sex, F= female, M=Male
- Centroid: centroid size, the square root of the summed squared distances of each landmark from the centroid of the landmark configuration. This is a proxy for body size.
- Resized X#: The X landmark coordinate location. Each fish had sixteen landmarks so there are 16 X variables: Resized X1, Resized X2, Resized X3, etc.
- Resized Y#: The Y landmark coordinate location. Each fish had sixteen landmarks so there are 16 Y variables: Resized Y1, Resized Y2, Resized Y3, etc.
File: links.xlsx
Description: This file is used to provide information on how body shape landmarks are linked together
Variables
- [,1]: landmark that is linked with the landmark in [,2]. For example, if [,1]=1, and [,2]=10, then landmarks 1 and 10 are linked.
- [,2] landmark that is linked with the landmark in [,1].
File: LOE_cr.Rmd
Description: This R script is used to analyze hypoxia and hydrogen sulfide tolerance data.
File: SL_CR_2015.xlsx
Description: This dataset includes information on fish body size.
Variables
- Fish: Fish ID
- Sex: Fish sex, F= female, M= male
- SL: Standard length (measured from snout to caudal peduncle in mm)
- Pop: habitat of origin (sulfidic or Nonsulfidic)
- Site: Site ID
File: Table_S2-Poecilia_mexicana_Annotations_1.csv
Description: This dataset includes protein annotations for genes in the gene expression analysis.
Variables
- gene ID: internal gene ID
- gene name: internal gene name
- Subject sequence ID: subject or target (reference genome) sequence ID
- E-value: The Expect value (E) is a parameter that describes the number of hits one can “expect” to see by chance when searching a database of a particular size. It decreases exponentially as the Score (S) of the match increases. Essentially, the E value describes the random background noise. For example, an E value of 1 assigned to an alignment means that, in a database of the same size, one expects to see 1 match with a similar score or higher simply by chance.
- Bit score: The bit-score is a rescaled alignment score to indicate the alignment quality, which is independent of the size of the search database. The higher the bit score, the better the alignment is.
- Protein annotations: Protein annotation of the gene
File: samples.xlsx
Description: Information on fish used in the gene expression analyses. Missing values are indicated using blank cells.
Variables
- Sample: Sample ID
- Species: species (P.gillii, L. Perugiae, P.mex, or P.sulph)
- Site: site name
- Sex: sex of fish, F= female, M=male
Code/software
This data was analyzed using R version 4.5.1.
Environmental data (cr_env_data.xlsx), body size (SL_CR_2015.xlsx), and body shape (Final_Total_Geomorph.xlsx and links.xlsx) data were analyzed using the R script Costa_morph_dv.
Hypoxia and hydrogen sulfide tolerance (CR_2025_LOE.xlsx) were analyzed using the R script LOE_cr.
Gene expression data (counts_matrix_cr_countreadpairs.txt, samples.xlsx, and Table_S2-Poecilia_mexicana_Annotations 1.xlsx) were analyzed using the script CR_RNA_2015_final.
