检验了语言分析工具LIWC在抑郁分类中的实际作用,发现其预测价值有限。
Disentangling the Interpretive and Predictive Roles of LIWC: Controlled Substitution in Depression-Related Classification

- 用四种替换方式对比LIWC,测试其对抑郁分类的贡献
- 在五组数据中均未发现显著提升,统计上不成立
- 适合用于解释语言特征,但不适合依赖其做预测
语言探究与词频(LIWC)提供可审计的心理语言学类别,广泛用于解读抑郁相关语言,但在现代多模态系统中的增量预测能力仍不明确。我们在五个英文和中文抑郁相关语料库中,采用匹配的参与者级交叉验证评估LIWC。考察其是否提升分类性能,以及性能变化反映什么。将完整LIWC与三种局部替换版本对比:主成分旋转版(移除类别坐标直接访问)、参与者随机打乱版(保留真实LIWC分布但破坏个体对齐)、随机边际版(保留特征分布)。在多个固定表示上下文中,完整LIWC在冻结、参与者级早期融合下未见稳定增益。所有预设数据块对比均未通过多重比较校正。独立的SBERT校准显示,完整版与打乱版、随机版间差异更大,表明更大的参与者对齐信号可在相同流程下产生更明显分离,但无法克服五语料库的统计效力限制。LIWC仍适合作为可审计、语料条件化的解释层。这些结论不可推广至微调、序列感知或端到端架构。
原文摘要 · Abstract (English)
Linguistic Inquiry and Word Count (LIWC) provides auditable psycholinguistic categories that are widely used to interpret depression-related language, but its incremental predictive role in modern multimodal systems remains unclear. We evaluate LIWC across five English and Chinese depression-related corpora under matched participant-level cross-validation. We ask whether LIWC improves classification and what any performance change reflects. Intact LIWC is compared with three fold-local substitutes: a PCA-rotated version that removes direct access to named category coordinates, a participant-shuffled version that preserves real LIWC profiles while breaking participant alignment, and a random-marginal version that preserves feature-wise distributions. Across multiple fixed representation contexts, the results provide limited evidence for stable LIWC gains under frozen, participant-level early fusion. None of the prespecified dataset-blocked contrasts survives multiple-comparison correction. A separate SBERT calibration produces larger observed intact-versus-shuffled and intact-versus-random separations, indicating that larger participant-aligned signals can produce correspondingly larger separations under the same procedure, while not resolving the five-corpus power limitation. LIWC remains useful as an auditable, corpus-conditioned interpretive layer. These conclusions should not be generalized to fine-tuned, sequence-aware, or end-to-end architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。