arXiv:2506.12266cs.CLcs.AI2025-06ACL被引 7

对比人类专家与大模型在复杂任务对话中的行为差异,发现差距越大性能越差。

The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs

  • 构建评估框架,量化模型与人类在对话行为、工具使用和知识调用上的差异。
  • 任务越复杂,行为差距越大,最复杂任务中模型对话行为F1仅0.464。
  • 缩小行为差距可提升性能24.3%,适合关注模型可靠性与人机对齐的研究者。

基于大语言模型(LLM)的智能体显著影响任务导向对话系统(TODS),但在零样本场景下仍存在明显性能瓶颈。以往研究虽指出性能差距,但其行为成因尚不明确。本研究提出综合性评估框架,量化AI智能体与人类专家在对话行为、工具使用及知识调用方面的差异。结果表明,行为差距是影响性能的关键因素;随着任务复杂度增加,该差距呈显著扩大趋势(相关性0.963),导致智能体在复杂任务对话中表现下降。在最复杂的任务中,即使采用GPT-4o的智能体,其对话行为F1仅为0.464,工具使用错误率高且不匹配,F1仅0.139,外部知识利用效率低下。减少此类行为差距可带来平均24.3%的性能提升。研究强调了全面行为评估与对齐策略的重要性,以增强LLM驱动的TODS处理复杂任务的能力。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based agents have significantly impacted Task-Oriented Dialog Systems (TODS) but continue to face notable performance challenges, especially in zero-shot scenarios. While prior work has noted this performance gap, the behavioral factors driving the performance gap remain under-explored. This study proposes a comprehensive evaluation framework to quantify the behavior gap between AI agents and human experts, focusing on discrepancies in dialog acts, tool usage, and knowledge utilization. Our findings reveal that this behavior gap is a critical factor negatively impacting the performance of LLM agents. Notably, as task complexity increases, the behavior gap widens (correlation: 0.963), leading to a degradation of agent performance on complex task-oriented dialogs. For the most complex task in our study, even the GPT-4o-based agent exhibits low alignment with human behavior, with low F1 scores for dialog acts (0.464), excessive and often misaligned tool usage with a F1 score of 0.139, and ineffective usage of external knowledge. Reducing such behavior gaps leads to significant performance improvement (24.3% on average). This study highlights the importance of comprehensive behavioral evaluations and improved alignment strategies to enhance the effectiveness of LLM-based TODS in handling complex tasks.

大模型评估对话系统人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。