arXiv:2601.18061cs.AIcs.HC2026-01被引 2

专家对心理健康AI的评估分歧极大,说明集体打分难以作为可靠训练标准。

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

  • 三位精神科医生独立评分,一致性极低(ICC 0.087–0.295)
  • 涉及自残与自杀的回复分歧最大,甚至出现负可靠性(α = -0.203)
  • 分歧源于不同临床理念,非误差,适合安全评估与模型对齐研究者参考

学习人类反馈(LHF)假设经过适当聚合的专家判断可作为训练和评估AI系统的有效真实标签。我们在高安全风险的心理健康领域测试了这一假设。三位持证精神科医生使用校准量表独立评估大语言模型生成的回应。尽管接受相同培训并遵循一致指令,三位医生之间的评分一致性持续偏低(ICC 0.087–0.295),低于重要评估所要求的可接受阈值。分歧在最安全关键的项目上最为显著,尤其是涉及自杀与自伤的回应,且呈系统性而非随机。一项因素甚至导致负可靠性(Krippendorff's α = -0.203),表明分歧程度比随机还严重。定性访谈显示,分歧源于合理但不兼容的临床框架——以安全为先、注重互动、文化敏感等不同取向,并非测量误差。结果表明,专家依赖整体风险直觉而非细粒度因子判别,聚合标签实为算术妥协,抹去了专业哲学基础。该研究将心理健康AI中的专家分歧定义为一种社会技术现象,建议从业者从共识聚合转向保留并学习专家分歧的对齐方法。

原文摘要 · Abstract (English)

Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI systems. We tested this assumption in mental health, where high safety stakes make expert consensus essential. Three certified psychiatrists independently evaluated LLM-generated responses using a calibrated rubric. Despite similar training and shared instructions, inter-rater reliability was consistently poor ($ICC$ $0.087$--$0.295$), falling below thresholds considered acceptable for consequential assessment. Disagreement was highest on the most safety-critical items. Suicide and self-harm responses produced greater divergence than any other category, and was systematic rather than random. One factor yielded negative reliability (Krippendorff's $α= -0.203$), indicating structured disagreement worse than chance. Qualitative interviews revealed that disagreement reflects coherent but incompatible individual clinical frameworks, safety-first, engagement-centered, and culturally-informed orientations, rather than measurement error. By demonstrating that experts rely on holistic risk heuristics rather than granular factor discrimination, these findings suggest that aggregated labels function as arithmetic compromises that effectively erase grounded professional philosophies. Our results characterize expert disagreement in safety-critical AI as a sociotechnical phenomenon where professional experience introduces sophisticated layers of principled divergence. We discuss implications for reward modeling, safety classification, and evaluation benchmarks, recommending that practitioners shift from consensus-based aggregation to alignment methods that preserve and learn from expert disagreement.

AI安全心理健康专家分歧对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。