Serum CST4 and routine laboratory indicators data for gastrointestinal tumor SVM diagnostic model
Data files
May 04, 2026 version files 103.82 KB
-
Processed_Data_SVM_Model.csv
17.99 KB
-
RAW_Data_Gastrointestinal_Tumor_2022-2023.csv
77.72 KB
-
README.md
8.11 KB
Abstract
This dataset supports the research titled "Serum Cystatin 4 Combined with Routine Clinical Laboratory Indicators: A Support Vector Machine Diagnostic Model for Early Screening of Gastrointestinal Tumors". It includes clinical laboratory data from 344 subjects, consisting of 214 patients with pathologically confirmed gastrointestinal tumors (91 gastric cancer, 80 colorectal cancer, 43 esophageal cancer) and 130 non-tumor individuals who underwent physical examinations at the same institution between January 2022 and June 2025.
The dataset contains 38 laboratory indicators per subject, including 14 blood routine parameters (e.g., white blood cell count [WBC], hematocrit [HCT], platelet count [PLT]), 16 biochemical indicators (e.g., total protein [TP], albumin [ALB]), 8 traditional tumor markers (e.g., carcinoembryonic antigen [CEA], carbohydrate antigen 50 [CA50]), and serum cystatin 4 (CST4) detected by enzyme-linked immunosorbent assay (ELISA). All data have undergone preprocessing, including mean imputation for missing values and Z-score method (|Z|>3) for outlier handling to ensure data quality.
This dataset serves as the foundational data for constructing and validating machine learning-based diagnostic models for early gastrointestinal tumor screening. Researchers can use it to reproduce the support vector machine (SVM) model developed in the study, compare the performance of different algorithms, or explore additional predictive biomarkers. Detailed variable definitions, preprocessing protocols, and usage guidelines are provided in the accompanying README file.
Dataset DOI: 10.5061/dryad.3ffbg79zx
Description of the data and file structure
Summary of Redman Gastrointestinal Tumor-Related Detection Dataset README
This document (Version v1.0) is compiled by Bi Huijuan from the Laboratory Department of Anhui Provincial Public Health Clinical Center. It focuses on the gastrointestinal tumor-related detection dataset from 2022 to 2023, detailing data-related information and usage requirements.
I. Basic Information of the Dataset
- Includes two CSV-format files (UTF-8 encoding): a processed dataset (core indicators) and a raw dataset (complete indicators). Missing values are marked as empty values, with commas as delimiters.
- Total sample size exceeds 344 cases, divided into the Gastrointestinal Tumor Group (214 cases, pathologically confirmed, Group=1) and the Control Group (130 cases, healthy physical examinees/non-tumor population, Group=0), using consecutive inclusion sampling.
II. Core Research Information
- Background & Objectives: Addressing the high incidence of gastrointestinal tumors, the study explores the diagnostic value of tumor markers (e.g., CST4), blood routine, and biochemical indicators, analyzes the impact of factors like age and gender, and provides support for constructing a multi-indicator combined diagnostic model.
- Study Design: A cross-sectional study. Eligible participants are aged ≥20 years with complete tests and clinical data; those with missing key data, combined other malignant tumors, etc., are excluded. Ethical approval (No. SL-XJ2023-014) has been obtained.
- Sample & Detection: Serum samples are collected via venous blood draw and standardized centrifugation. Detection is performed using instruments (e.g., Siemens, Mindray) and supporting reagents, covering 8 tumor markers, 14 blood routine indicators, and 16 biochemical indicators, with strict internal and external quality control.
III. Data Processing and Coding
- Data Cleaning: Involves format standardization, missing value handling (mean/median imputation for missing rate <5%), outlier correction (rectifying entry errors, retaining annotation of detection limit-related values), duplicate record removal (via SubjectID), logical consistency checks, and post-cleaning result validation.
- Coding Rules: Adopts a three-level system "Project Code-Group Identifier-Sample Serial Number" (format: CST4-XX-000). "CST4" is the project code; "G" represents the Gastrointestinal Tumor Group, "C" the Control Group; "000" is a 3-digit serial number starting from 001.
Files and variables
File: Processed_Data_SVM_Model.csv
Description: This preprocessed dataset (named "Processed Data SVM Model.csv") is derived from the Redman Gastrointestinal Tumor-Related Detection Dataset, which integrates clinical detection data collected from January 2022 to May 2023. It is a refined subset of the raw dataset ("RAW Data Gastrointestinal Tumor 2022-2023.csv"), specifically optimized for the construction and validation of SVM (Support Vector Machine) diagnostic models. The dataset undergoes strict data cleaning and standardization to ensure high quality, with all samples anonymized for privacy protection while retaining core diagnostic information. It serves as a direct and efficient data foundation for exploring the diagnostic value of key indicators in gastrointestinal tumors and developing multi-indicator combined diagnostic tools.
Variables
- SubjectID: The format is CST4-XX-000, where CST4 represents the project code, XX is a letter-based group identifier (G for tumor group, C for control group), and 000 is a 3-digit sample serial number incrementing from 001.
- label:1=Gastrointestinal Tumor Group; 0=Control Group
- Age:
- CST4:Cystatin S Cystatin S
- WBC: leucocyte count
- HCT: hematocrit
- PLT: platelet count
- TP: total protein
- ALB: albumin
File: RAW_Data_Gastrointestinal_Tumor_2022-2023.csv
Description: The raw dataset (named "RAW Data Gastrointestinal Tumor 2022-2023.csv") is the foundational source of the Redman Gastrointestinal Tumor-Related Detection Dataset, integrating comprehensive clinical and laboratory data collected from January 2022 to May 2023. It is derived from the Laboratory Information System (LIS) and Hospital Information System (HIS) of Anhui Provincial Public Health Clinical Center, covering consecutive eligible subjects including gastrointestinal tumor patients (pathologically confirmed) and control populations (healthy physical examinees/non-tumor individuals). The dataset retains all original detection records without excessive variable screening, ensuring the completeness and authenticity of information. After standardized formatting and preliminary quality control (e.g., removing severely invalid records), it serves as the core raw material for subsequent data cleaning, variable derivation, and multi-dimensional research (such as exploratory analysis of detection indicators, construction of combined diagnostic models, and analysis of influencing factors).
Variables
- CollectionDate: CollectionDate refers to the date when venous blood specimens were collected from research subjects. It is formatted as "YYYY-MM-DD" to ensure standardized time recording.
- TestDate:estDate denotes the specific date when the collected serum specimens were analyzed and tested in the clinical laboratory, formatted uniformly as "YYYY-MM-DD" for standardized time documentation.
- SubjectID: SubjectID refers to a unique code used to anonymously identify individual participants in a study, ensuring data confidentiality."
- group:1=Gastrointestinal Tumor Group; 0=Control Group
- Gender:1=male; 0=female
- Age:
- AFP: Alpha-Fetoprotein Alpha-Fetoprotein
- CEA: Carcinoembryonic Antigen Carcinoembryonic Antigen
- CA19-9: Glycan Antigen 199
- CA-125: Glycan Antigen 125
- CA72-4: Glycan Antigen 724
- CA50: Glycan Antigen 50
- CA242: Glycan Antigen 242
- CST4: Cystatin S Cystatin S
- WBC: leucocyte count
- NEU: Neutrophil count
- LYM: lymphocyte count
- MON: Monocyte count
- EOS: Eosinophil count
- BAS: Basophil count
- RBC: red-cell count
- HGB: hemoglobin
- HCT: hematocrit
- MCV: mean corpuscular volume
- MCH: Mean hemoglobin content
- MCHC: mean hemoglobin concentration
- PLT: platelet count
- RET: reticulocyte count
- TP: total protein
- ALB: albumin
- GLOB: globin
- TBA: total bile acid
- TBIL: total bilirubin
- DBIL: bilirubin direct
- IBIL: indirect bilirubin
- ALT: glutamic-pyruvic transaminase
- AST: glutamic-oxalacetic transaminase
- LDH: lactate dehydrogenase
- ALP: alkaline phosphatase
- BUN: usea nitrogen
- CREA: creatinine
- UA: purine trione
- CHOL: cholesterol
- TG: glycerin trilaurate
Code/software
Data processing, model construction, and visualization were performed using Python 3.10 with standardized libraries including scikit-learn (1.5.2, https://scikit-learn.org), XGBoost (1.7.4), and LGBM (4.0.0, https://lightgbm.readthedocs.io), etc., the models underwent optimization on a 7: 3 training-validation split. SHAP 1.4.6.1. R 4.5.2 (R Foundation for Statistical Computing, Vienna, Austria) was used for calibration curve analysis, decision curve analysis (DCA), and correlation heatmap generation, with the rms package 8.1-0 and ggplot 4.0.0.
Human subjects data
The submitted dataset has been fully de-identified in compliance with Dryad’s requirements for human subjects data. All direct identifiers, including but not limited to names, addresses, birth dates, medical examination and procedure dates, and personal identification numbers, have been permanently removed. Indirect identifiers (age, geographic location, and clinical characteristics) have been generalized or grouped to prevent the re-identification of any individual. No information in the dataset can be used to identify specific research participants.
