评测大模型心理推理的可靠性,发现其判断常超出证据支持范围。
MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy
- 构建跨六类任务的统一评测基准MiraMind,评估推理过程与证据的匹配度。
- 20个模型普遍存在判断过度具体或确定,超出证据支持范围的问题。
- 提出新模型Mindora,通过结构化重写和一致性优化,提升推理平衡性。
大语言模型在心理健康推理中面临证据约束判断难题:模型需将有限、主观且常模糊的证据转化为具有适当具体性、确定性、严重性和可操作性的解释、决策或主张。现有基准多聚焦特定临床角色或最终答案,对推理路径的可靠性评估不足。我们提出MiraMind,一个涵盖六类任务与13个数据集的统一基准,覆盖评估、诊断、干预、抽象与验证。该基准不仅评估任务结果,还通过可用性、逻辑结构和信息贡献度评估从证据到判断的推理轨迹。对20个LLMs的评估揭示存在‘克制差距’——模型判断的具体性或确定性超出有限证据支持。我们进一步训练了Mindora(8B参数),通过难例监督、结构化轨迹重写与一致性感知优化,增强证据到判断的过渡。Mindora在MiraMind上取得最佳平均排名,在全部六类任务中均优于基线模型,并生成更均衡的推理轨迹。结果表明,MiraMind能暴露共性克制缺陷,并有效评估针对性后训练对证据约束型心理推理的改进效果。
原文摘要 · Abstract (English)
Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, and often ambiguous evidence into interpretations, decisions, or claims whose specificity, certainty, severity, and actionability remain warranted. Existing benchmarks mainly evaluate specific clinical roles or final answers, leaving the reliability of explicit reasoning trajectories under-specified. We introduce MiraMind, a unified benchmark spanning six task families and 13 datasets across appraisal, diagnosis, intervention, abstraction, and verification. MiraMind evaluates both task outcomes and the trajectories connecting evidence to judgment through usability, logical structure, and informational contribution. Evaluating 20 LLMs reveals a restraint gap, in which the specificity or certainty of model judgments exceeds what limited evidence supports. We further train Mindora, an 8B model that targets evidence-to-judgment transitions through hard-case supervision, structured trajectory rewriting, and consistency-aware optimization. Mindora achieves the best average rank on MiraMind, improves over its backbone across all six task families, and produces more balanced reasoning trajectories. These results show that MiraMind can expose shared restraint failures and evaluate whether targeted post-training improves evidence-constrained mental-health reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。