分析临床访谈数据发现,医生提问方式会误导模型,让其误判抑郁。
When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
- 用医生提问内容训练模型,易学错线索
- 去除医生话语后,模型准确率下降但更依赖真实语言
- 适合关注模型可解释性的研究者
自动从医患对话中检测抑郁症因公开数据集和语言建模进步而日益流行。然而,模型的可解释性仍有限:性能虽高,却常不揭示预测依据。我们分析了三个数据集——ANDROIDS、DAIC-WOZ、E-DAIC,发现半结构化访谈中存在系统性偏差,源于面试官的固定提问模式和位置。基于面试官发言训练的模型,往往通过固定提问句式与位置区分抑郁与正常人群,而非真正理解患者语言。若限制模型仅使用患者话语,决策依据分布更广,反映真实的语言特征。尽管半结构化流程保证一致性,但包含面试官提问会因脚本痕迹虚高模型表现。结果表明该偏差具有跨数据集、架构无关性,强调需按时间与说话人定位决策证据,确保模型真正学习患者语言。
原文摘要 · Abstract (English)
Automatic depression detection from doctor-patient conversations has gained momentum thanks to the availability of public corpora and advances in language modeling. However, interpretability remains limited: strong performance is often reported without revealing what drives predictions. We analyze three datasets: ANDROIDS, DAIC-WOZ, E-DAIC and identify a systematic bias from interviewer prompts in semi-structured interviews. Models trained on interviewer turns exploit fixed prompts and positions to distinguish depressed from control subjects, often achieving high classification scores without using participant language. Restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. While semi-structured protocols ensure consistency, including interviewer prompts inflates performance by leveraging script artifacts. Our results highlight a cross-dataset, architecture-agnostic bias and emphasize the need for analyses that localize decision evidence by time and speaker to ensure models learn from participants' language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。