Taphonomic megabiases constrain phylogenetic information in the squamate fossil record
Data files
Aug 29, 2025 version files 1.59 MB
-
Data_S1a_Squamate_Fossil_record.csv
408.98 KB
-
Data_S1b_CCM2_Character_List.csv
226.01 KB
-
Data_S1c_Taxon_scoring.csv
118.27 KB
-
Data_S1d_Squamate_collections_PBDB.csv
335.14 KB
-
Data_S1e_Squamate_Formations_PBDB.csv
379.96 KB
-
Data_S2_Statistical_Tests.xlsx
84.44 KB
-
Data_S3_R_Code_examples.R
9.85 KB
-
Data_S3a_timescale_age_ranges.txt
809 B
-
Data_S3b_Squamate_Species_Thru_Time.txt
400 B
-
Data_S3c_Lizards_CCM2.csv
3.77 KB
-
Data_S3d_Species_Specimens_CCM2.csv
2.02 KB
-
Data_S3e_All_Lithologies.csv
8.78 KB
-
README.md
8.33 KB
Abstract
Fossil data is subject to inherent biological, geological, and anthropogenic filters that can distort our interpretations of ancient life and environments. The inevitable presence of incomplete fossils thus requires a holistic assessment of how to navigate the downstream effects of bias on our ability to accurately reconstruct aspects of biology in deep time. In particular, we must assess how biases affect our capacity to infer evolutionary relationships, which are essential to analyses of diversification, paleobiogeography, and biostratigraphy in Earth history. In this study, we use an established completeness metric to quantify the effects of taphonomic filters on the amount of phylogenetic information available in the fossil record of 795 extinct squamate (e.g., lizards, snakes, amphisbaenians, and mosasaurs) species spanning 242 million years of geologic time. This study found no meaningful relationship between spatiotemporal sampling intensity and fossil record completeness. Instead, major differences in squamate fossil record completeness stem from a combination of anatomy/body-size and affinities of different squamate groups to specific lithologies and depositional environments. These results reveal that naturally occurring processes create structural megabiases that filter anatomical and phylogenetic data in the squamate fossil record, while anthropogenic processes play a secondary role.
Dataset DOI: 10.5061/dryad.hhmgqnktp
Description of the data and file structure
Datasets, statistical analyses, and R Code used in this study
- File Supplemental Information: PDF file containing supplementary methods, discussion, figures, and tables used for this study.
- Files that begin with the title "Data_S1..." include detailed occurrence data for all squamate taxa examined in this study, including links to online versions of references where applicable, as well as datasets used for assessing the Character Completeness Metric (CCM2), and the corresponding scores for each squamate species.
- The first sheet is entitled "Data_S1a_Squamate_Fossil_record.csv". It contains all occurrence data for the 795 squamate species sampled in this study. The sheet contains 27 columns that include taxonomic information (Column A-D); Specimen information (E-G), Age (H-J); Geologic information (K-N); Locality Information (O-S); Completeness score (T); and citations/references (U-AA).
- The second sheet is entitled "Data_S1b_CCM2_Character_List.csv" which includes all the characters in our combined dataset we used to assess the completeness of each squamate taxon. Each character is organized by region of the skeleton.
- The third sheet is titled "Data_S1c_Taxon_scoring.csv" and includes breakdowns of the number of characters scored for each of the 795 squamate taxa and final percentages.
- The fourth sheet is titled "Data_S1d_Squamate_collections_PBDB.csv" which includes data from the Paleobiology Database (PBDB) on every published squamate collection in the database for comparison to our CCM2 data.
- The fifth sheet is titled "Data_S1e_Squamate_Formations_PBDB.csv", which includes all squamate-bearing geologic formations on the PBDB.
- File Data_S2_Statistical_Tests.xlsx: Excel file containing results of all statistical tests in this study, formatted for ease of reading.
- There are 10 sheets within this .xlsx file:
- The first sheet is entitled "Simple Linear Regression", which compares geologic stage-level values among the four time series variables: Median CCM2 per stage, Squamate-bearing formations per stage, squamate collections per stage, squamate species per stage. These tests were carried out in R and compiled in Microsoft Excel.
- The second sheet is entitled "GLS (log-transformed", which includes results for the Generalized Least-Squares regression tests performed on log-transformed values for the same time-series variables as above: Median CCM2 per stage, Squamate-bearing formations per stage, squamate collections per stage, squamate species per stage. These tests were carried out in R and compiled in Microsoft Excel.
- The third sheet is entitled "Anatomical Gap Taxon Tests" and includes results from the non-parametric statistical comparisons (Mann-Whitney U-Test [comparison of median CCM2 scores]; Kolmogorov-Smirnov Test [comparison of distribution shapes]) between the four main anatomical groups of squamates: lizards, snakes, amphisbaenians, and mosasaurs. These tests were carried out in R and compiled in Microsoft Excel.
- The fourth sheet is entitled "Anatomical Gap Taxon Tests (LC)" which carries out the same non-parametric statistical tests as above, but using corrected CCM2 scores for limb-reduced or limbless taxa (e.g., snakes and amphisbaenians). The LC stands for "Limbless-Corrected", and is used for all statistical tests. These tests were carried out in R and compiled in Microsoft Excel.
- The fifth sheet is entitled "Clade Taxon Tests (LC)", and includes non-parametric statistical comparisons of completeness score distribution of all major squamate clades. These tests include the Mann-Whitney U-Test [comparison of median CCM2 scores] and Kolmogorov-Smirnov Test [comparison of distribution shapes]. These tests were carried out in R and compiled in Microsoft Excel.
- The sixth sheet is entitled "Landmass Tests", and includes non-parametric statistical comparisons of completeness score distributions across sampled landmasses. These tests include the Mann-Whitney U-Test [comparison of median CCM2 scores] and Kolmogorov-Smirnov Test [comparison of distribution shapes]. These tests were carried out in R and compiled in Microsoft Excel.
- The seventh sheet is entitled "Landmass Tests (LC)" and includes non-parametric statistical comparisons of limbless-corrected completeness score distributions across sampled landmasses. These tests include the Mann-Whitney U-Test [comparison of median CCM2 scores] and Kolmogorov-Smirnov Test [comparison of distribution shapes]. These tests were carried out in R and compiled in Microsoft Excel.
- The eighth sheet is entitled "Depositional Setting Tests" and includes non-parametric statistical comparisons of completeness score distributions across sampled depositional settings containing squamate fossils. These tests include the Mann-Whitney U-Test [comparison of median CCM2 scores] and Kolmogorov-Smirnov Test [comparison of distribution shapes]. These tests were carried out in R and compiled in Microsoft Excel.
- The ninth sheet is entitled "Depositional Setting Tests (LC)" and includes non-parametric statistical comparisons of limbless-corrected completeness score distributions across sampled depositional settings containing squamate fossils. These tests include the Mann-Whitney U-Test [comparison of median CCM2 scores] and Kolmogorov-Smirnov Test [comparison of distribution shapes]. These tests were carried out in R and compiled in Microsoft Excel.
- The tenth and final sheet is entitled" Lithology Tests (LC)" and includes non-parametric statistical comparisons of limbless-corrected completeness score distributions across sampled lithologies preserving squamate fossils. These tests include the Mann-Whitney U-Test [comparison of median CCM2 scores] and Kolmogorov-Smirnov Test [comparison of distribution shapes]. These tests were carried out in R and compiled in Microsoft Excel.
- There are 10 sheets within this .xlsx file:
- File Data_S3_R_Code_examples.R: R file including generalized code used to carry out graphical displays and statistical analyses in this study. The following files, derived from Data_S1_Squamate_fossil_record.xlsx, can be used to run the example analyses:
- Data_S3a_timescale_age_ranges.txt. This is a file containing abbreviated geologic stage names and their corresponding age ranges to upload in the geoscale functions in Data_S3_R_Code_examples.R for replicating the time series plots in this study.
- Data_S3b_Squamate_Species_Thru_Time.txt. This is a text file containing the number of sampled squamate species for each geologic stage in Data S3a. This can be uploaded in Data_S3_R_Code_examples.R for replicating the time series plots in this study.
- Data_S3c_Lizards_CCM2.csv. This file is a single-column CSV file that includes the Character Completeness Metric 2 (CCM2) scorings for all sampled lizard taxa. This can be uploaded in Data_S3_R_Code_examples.R for replicating the violin plots in this study.
- Data_S3d_Species_Specimens_CCM2.csv. This file is a multicolumn CSV file that includes all time series data in this study. This can be uploaded in Data_S3_R_Code_examples.R for replicating the Generalized Least Squares (GLS) regression analyses in this study.
- Data_S3e_All_Lithologies.csv. This file is a multicolumn CSV file that includes the CCM2 scores of all squamates preserved in specific rock lithologies around the globe. This can be uploaded in Data_S3_R_Code_examples.R for replicating the non-parametric statistical tests we used in this study.
Supplemental_Information.pdf: It contains additional figures, statistical test tables, datasets, and extensive references on fossil squamates, providing detailed evidence and analyses of how preservation biases affect our understanding of lizard and snake evolutionary history.
Code/software
All analyses can be carried out in the latest version of R.
Citation: R Core Team (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL https://www.R-project.org/.
