提出多模态代理评估新基准,支持用户人格适配与双控场景评测。
MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings
- 构建双控设置下的多模态评估框架,引入用户人格适应机制。
- 在电信与零售领域验证12项新指标,发现多模态引入增加交互开销。
- 适用于需个性化响应的客服、导购类多模态智能体研发与评估。
当前基于大语言模型的智能体评估框架主要聚焦于文本对话场景,未暴露用户人格信息,处于用户无关环境。在客户体验管理领域,智能体行为会随对用户性格的认知而演变。随着实时语音合成与多模态大模型的发展,基于大模型的智能体正迈向多模态化。为此,我们提出MM-tau-p²基准,用于评估双控场景下多模态智能体在有无用户人格适配时的鲁棒性,并将用户输入纳入规划过程以解决查询。研究发现,即使使用GPT-5、GPT 4.1等前沿模型,引入多模态仍需考虑多模态鲁棒性与回合开销等问题。该基准基于前期工作FOCAL,通过精心设计提示词与评分标准,采用大模型为裁判的方法,在电信与零售领域提供了12项新指标的量化估计。
原文摘要 · Abstract (English)
Current evaluation frameworks and benchmarks for LLM powered agents focus on text chat driven agents, these frameworks do not expose the persona of user to the agent, thus operating in a user agnostic environment. Importantly, in customer experience management domain, the agent's behaviour evolves as the agent learns about user personality. With proliferation of real time TTS and multi-modal language models, LLM based agents are gradually going to become multi-modal. Towards this, we propose the MM-tau-p$^2$ benchmark with metrics for evaluating the robustness of multi-modal agents in dual control setting with and without persona adaption of user, while also taking user inputs in the planning process to resolve a user query. In particular, our work shows that even with state of-the-art frontier LLMs like GPT-5, GPT 4.1, there are additional considerations measured using metrics viz. multi-modal robustness, turn overhead while introducing multi-modality into LLM based agents. Overall, MM-tau-p$^2$ builds on our prior work FOCAL and provides a holistic way of evaluating multi-modal agents in an automated way by introducing 12 novel metrics. We also provide estimates of these metrics on the telecom and retail domains by using the LLM-as-judge approach using carefully crafted prompts with well defined rubrics for evaluating each conversation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。