arXiv:2604.06216cs.CLcs.AI2026-04

融合人类专家与大模型,提升心理健康聊天机器人幻觉与遗漏检测准确率。

Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses

论文配图:Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
图 1 · 摘自论文原文
  • 构建五维分析框架,融合人类专业判断提取可解释特征。
  • 在自建数据集上幻觉检测F1达0.717,公开基准上达0.849。
  • 适合高风险医疗场景下的智能对话系统评估与安全审核。

随着大模型驱动的聊天机器人在心理健康服务中日益普及,检测幻觉与遗漏对用户安全至关重要。然而,现有基于大模型作为评判者的方法在高风险医疗场景中表现不佳,其中部分幻觉检测方法召回率接近零。我们发现根本原因在于大模型难以捕捉领域专家识别出的细微语言与治疗模式。为此,我们提出一种融合人类专家与大模型的框架,从逻辑一致性、实体验证、事实准确性、语言不确定性及专业适切性五个维度提取可解释的领域相关特征。在公开心理健康数据集和新构建的人工标注数据集上的实验表明,基于这些特征训练的传统机器学习模型在自建数据集上幻觉检测的F1为0.717,在公开基准上为0.849;遗漏检测在两个数据集上的F1均在0.59至0.64之间。结果表明,在高风险心理医疗应用中,结合领域知识与自动化方法的评估方式比黑箱大模型评判更可靠、透明。

原文摘要 · Abstract (English)

As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safety. However, state-of-the-art LLM-as-a-judge methods often fail in high-risk healthcare contexts, where subtle errors can have serious consequences. We show that leading LLM judges achieve only 52% accuracy on mental health counseling data, with some hallucination detection approaches exhibiting near-zero recall. We identify the root cause as LLMs' inability to capture nuanced linguistic and therapeutic patterns recognized by domain experts. To address this, we propose a framework that integrates human expertise with LLMs to extract interpretable, domain-informed features across five analytical dimensions: logical consistency, entity verification, factual accuracy, linguistic uncertainty, and professional appropriateness. Experiments on a public mental health dataset and a new human-annotated dataset show that traditional machine learning models trained on these features achieve 0.717 F1 on our custom dataset and 0.849 F1 on a public benchmark for hallucination detection, with 0.59-0.64 F1 for omission detection across both datasets. Our results demonstrate that combining domain expertise with automated methods yields more reliable and transparent evaluation than black-box LLM judging in high-stakes mental health applications.

心理健康幻觉检测人机协作大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。