首个评估工具增强型情绪支持对话的基准,验证工具能有效减少幻觉。
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
- 构建交互式评测框架,含真实情感场景与工具环境
- 强模型用工具更高效,弱模型提升有限
- 适合研究可信对话系统与工具融合的开发者
情绪支持对话不仅需要情感表达,还需基于事实的行动支持以建立信任。现有情绪支持系统和评测大多局限于纯文本情境,忽视了外部工具在多轮对话中实现事实依据、降低幻觉的作用。本文提出TEA-Bench,首个面向工具增强型情绪支持对话的交互式评测基准,包含真实情感场景、类MCP的工具环境以及过程级评估指标,综合衡量支持质量与事实准确性。在九个大模型上的实验表明,工具增强普遍提升支持质量并减少幻觉,但增益高度依赖模型能力:强模型更选择性地使用工具,弱模型仅获微弱改善。我们还发布了TEA-Dialog数据集,发现监督微调虽提升分布内表现,但泛化能力差。结果强调工具使用对构建可靠情绪支持代理的重要性。代码与数据见https://github.com/XingYuSSS/TEA-Bench。
原文摘要 · Abstract (English)
Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance. However, existing ESC systems and benchmarks largely focus on affective support in text-only settings, overlooking how external tools can enable factual grounding and reduce hallucination in multi-turn emotional support. We introduce TEA-Bench, the first interactive benchmark for evaluating tool-augmented agents in ESC, featuring realistic emotional scenarios, an MCP-style tool environment, and process-level metrics that jointly assess the quality and factual grounding of emotional support. Experiments on nine LLMs show that tool augmentation generally improves emotional support quality and reduces hallucination, but the gains are strongly capacity-dependent: stronger models use tools more selectively and effectively, while weaker models benefit only marginally. We further release TEA-Dialog, a dataset of tool-enhanced ESC dialogues, and find that supervised fine-tuning improves in-distribution support but generalizes poorly. Our results underscore the importance of tool use in building reliable emotional support agents. Our code and data can be found in https://github.com/XingYuSSS/TEA-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。