让心理筛查模型学会区分不同语音采集方式的证据效力,避免胡编症状。
Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

- 按语音采集方式分类证据,限制模型推理范围
- 在抑郁症筛查中达0.8658的AUROC,比最强基线高0.0811
- 适合需要安全可靠临床NLP系统的研究者使用
基于多模态语音与文本的计算心理筛查展现出巨大潜力。然而,现有模型常假设所有临床语音协议具有同等证据效力。事实上,从自由访谈到固定朗读任务等异构协议所支持的证据基础截然不同。强制统一推理会模糊这些边界,导致模型从无关文本中虚构症状或过度宣称支持。即使先进的长链思维大模型也无法解决此问题,因其自由推理可能加剧边界违规。为此,我们重新将多模态筛查建模为证据有界推理问题。提出证据包基准(Evidence Package Benchmark),整合来自六种异构来源的1,870个证据包,包含明确的模态掩码与证据权限。进一步提出EviBound框架:通过用户画像感知规划器限制推理范围,利用五向声学共识协调证据工具,并由边界批判器抑制无支持声明。实证结果显示,EviBound在留出测试集上抑郁症筛查的AUROC达到0.8658,优于最强的全模态直接基线+0.0811 AUROC,且保持零声明违规。本工作推动临床NLP研究从无约束准确率迈向证据一致、协议感知的安全系统。
原文摘要 · Abstract (English)
Computational mental health screening using multimodal speech and text has shown great promise. However, existing models often assume all clinical speech protocols carry equivalent evidentiary validity. In reality, heterogeneous protocols, from free interviews to fixed reading tasks, support fundamentally different evidence. Forcing uniform reasoning flattens these boundaries, causing models to hallucinate symptoms from irrelevant text or overclaim support. Even advanced long chain-of-thought LLMs fail to resolve this issue, as free-form reasoning can exacerbate boundary violations. To address this, we reformulate multimodal screening as an evidence-bounded reasoning problem. We introduce the Evidence Package Benchmark, integrating 1,870 packages across six heterogeneous sources with explicit modality masks and evidence permissions. We further propose EviBound, a protocol-aware evidence control framework. Unlike direct LLM prompting, EviBound uses a profile-aware planner to restrict reasoning scope, orchestrates evidence tools via five-way acoustic consensus, and enforces a boundary critic to suppress unsupported claims. Empirical results show EviBound achieves a held-out test Depression AUROC of 0.8658, exceeding the strongest direct omni-modal baseline by +0.0811 AUROC while maintaining zero claim violations. Our work moves beyond unconstrained accuracy toward evidence-consistent, protocol-aware systems for safer clinical NLP research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。