Voppel · Schizophrenia (Heidelberg, Germany) 2025 · computational linguistic analysis with cross-sectional clinical validation · n=98

Analysis of conceptual overlap among formal thought disorder rating scales in psychosis: a systematic semantic synthesis.

Level 4 - case-series / case-control

Cross-sectional scale validation and computational measurement study using an observational clinical sample

PubMed 41398324 · doi:10.1038/s41537-025-00712-z · record verified 2026-08-26

What was done

Natural language processing (Sentence-BERT) was applied to analyze definitions and semantic overlap among formal thought disorder (FTD) rating scales across three validation steps. First, 30 items from the Thought and Language Disorder (TALD) scale were computationally clustered into positive and negative FTD groupings and compared to published factor analyses. Second, a similarity matrix was generated across 103 items from seven FTD scales to generate cross-scale clusters, which were compared against the classifications of six blinded clinical experts. Third, the semantic clusters were tested against cross-scale item correlations (Clinical Language Disorder Scale [CLANG] vs. Thought, Language and Communication [TLC] scale) in a clinical sample of 98 participants (49 healthy controls and 49 patients with schizophrenia or affective psychosis).

What was found

Sentence-BERT item clustering matched prior empirical factor analyses for 73% of TALD items. The cross-scale synthesis identified four coherent FTD clusters: (1) muddled communication & incomprehension, (2) abrupt topic shifts, (3) inconsistent narrative structure, and (4) restricted speech. Blinded expert raters demonstrated moderate-to-high agreement with the automated clusters (Fleiss' kappa = 0.617). In the participant dataset (n = 98), highest-correlating CLANG-TLC item pairs fell within the predicted semantic cluster significantly more often than expected by chance (binomial test, p < 0.001).

Why it matters

Inconsistent terminology across psychopathological rating scales hampers cross-study harmonization in psychosis research. This approach shows that computational linguistics can systematically map overlapping FTD constructs across disparate clinical instruments to improve measurement consistency.

Limits

The human empirical validation dataset was small (98 total participants, 49 with psychosis) and tested cross-scale item mapping on only two specific instruments (CLANG and TLC). Item-level semantic embeddings rely solely on text definitions and may miss subtleties in how clinicians interpret or score items during actual psychiatric evaluations.