Data from: Large language models for automated and audience-tailored labeling of latent classes
Data files
Apr 16, 2026 version files 3.08 KB
Abstract
This study compares multiple LLMs, including ChatGPT, DeepSeek, and Llama 3, to generate meaningful, audience-adapted labels for the existing latent classes among patients with chronic low back pain (cLBP). Phenotypes were derived from baseline data from two cohorts within the NIH HEAL BACPAC consortium: BACKHOME, a large nationwide e-cohort (train set: N=3,025), and COMEBACK, a deep phenotyping cohort (test set: N=450). The analysis included pain characteristics, psychosocial factors, lifestyle habits, and social determinants of health. ChatGPT-4o (OpenAI), DeepSeek-R1, and Llama 3 (Meta) were applied to generate class labels for each combination of audience (clinician, patient, and caregiver), tone (formal, empathetic, and informal), and technicality (high, medium, and low). Latent Class Model (LCM) identified four distinct behavioral phenotypes in patients with cLBP: High Distress and Maladaptive Behaviors, Resilient and Adaptive Coping, Intermediate Maladaptive Patterns, and Emotionally Regulated with High Pain Burden. Previously validated by domain experts, these profiles served as the basis for automated labeling using three LLMs (ChatGPT-4o, DeepSeek-R1, and Llama 3). Using different tones and complexity levels, each model produced class labels specific to clinicians, patients, and caregivers. The generated class names for all LLMs closely matched expert-defined traits like emotional regulation, resilience, and high distress, indicating strong conceptual alignment and the capacity of LLMs to generate precise, audience-specific labels for intricate behavioral and psychological profiles. These results highlight the possibility of integrating LLM-driven labeling into research and clinical practice, helping to achieve more transparent knowledge translation, improved decision-making, and personalized care.
Dataset DOI: 10.5061/dryad.1jwstqk9d
Description of the data and file structure
This dataset contains normalized class profile summaries derived from two independent cohorts used in the study of latent class identification in chronic low back pain. The training dataset (BACKHOME cohort) and the testing dataset (COMEBACK cohort) include aggregated, normalized values for psychosocial, behavioral, and clinical measures across four identified latent classes.
The data provided here represent summary-level (non-individual) values used to characterize latent class profiles. Individual-level patient data are not included due to access restrictions.
Files included in this dataset:
• _Class_Profiles_for_BACKHOME_(Train)and_COMEBACK(Test)_Sets.csv
Contains normalized class profile values for BACKHOME (training) and COMEBACK (testing) cohorts across four latent classes.
The original datasets are available through the Vivli data repository and require approval for access.
Access information
Other publicly accessible locations of the data:
• https://doi.org/10.25934/PR00010820
• https://doi.org/10.25934/PR00010819
Data was derived from the following sources:
• Vivli data repository (https://vivli.org)
Human subjects data
This study was conducted with institutional review board approval (WIRB #20-30368). Access to the original datasets is available through Vivli and requires approval in accordance with their data access policies. The data included in this Dryad submission are summary-level representations and do not contain individual-level patient data. Individual-level patient data are not included due to access restrictions.
