提出真实用户场景评估框架,发现大模型在复杂交互中表现普遍不佳
Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

- 构建涵盖理想与非理想行为的多轮交互评估基准
- 19个模型平均成功率低于40%,复杂输入下性能显著下降
- 适合关注大模型真实应用能力的研究者与开发者
尽管大型语言模型(LLMs)在工具使用方面取得显著进展,现有评估基准仍难以匹配真实场景。这些基准大多依赖模拟的理想化用户假设,缺乏面向用户体验的评估。该局限性未能涵盖真实用户常见的模糊表达、不合作行为及意图变化。为填补这一空白,我们提出RUT-Bench,一个专用于评估大模型在多样化真实用户工具调用场景下的基准。RUT-Bench支持高保真度仿真,涵盖单轮与多轮对话中的理想理性模式及异构非理想行为。我们在19个广泛使用的开源与专有大模型上进行了全面评估。实验结果表明,所有测试模型的整体成功率均未超过40%,且几乎所有模型在面对更复杂的非理想用户输入时均出现明显性能下降。代码与数据已公开于https://github.com/Miaow-Lab/RUT-Bench。
原文摘要 · Abstract (English)
Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world scenarios. Such benchmarks mostly rely on simulated idealized user assumptions and lacks experience-oriented evaluation. These limitations fail to account for the ambiguity, uncooperative behaviors, and shifting intentions characteristic of real-world users. To fill this gap, we propose RUT-Bench, a dedicated benchmark designed to assess LLMs under diverse Real-world User Tool calling scenarios. RUT-Bench supports high-fidelity simulations covering both ideal rational patterns and heterogeneous non-ideal behaviors across single-turn and multi-turn dialogues. We conduct comprehensive evaluations on 19 widely adopted open-source and proprietary LLMs using our benchmark. Experimental results reveal that no tested LLMs achieve an overall success rate above 40%, and nearly all of them experience noticeable performance drops when facing more complicated non-ideal user inputs. Our code and data is available at https://github.com/Miaow-Lab/RUT-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。