arXiv:2606.09570cs.CLcs.HC2026-06

首个基于真实用户反馈的AI助手体验评估基准,测对话质量与用户偏好对齐。

UXBench: Benchmarking User Experience in AI Assistants

论文配图:UXBench: Benchmarking User Experience in AI Assistants
图 1 · 摘自论文原文
  • 从7万次交互日志提取7400条数据,构建三任务用户体验评测体系。
  • 发现模型对用户反馈的预测能力可训练,奖励模型准确率达良好校准水平。
  • 揭示大模型评判协议系统性偏差,适合关注用户体验优化的研究者参考。

随着人工智能助手每日服务数百万用户,超越通用模型能力的用户体验(UX)评估日益重要。本文提出UXBench,首个基于真实用户反馈信号的以用户为中心的基准,用于评估偏好对齐与对话生成。该基准包含三个相互关联的任务:UX Judge、UX Eval 和 UX Recovery,共包含7,400个测试实例,源自主流中文AI助手超过7万次交互日志。数据集真实反映用户分布,涵盖8种场景、83个领域及多样化的失败模式,构成严峻挑战。在26个前沿语言模型上的广泛实验揭示了模型感知用户体验的能力及其与对话参与度的关系。通过深入分析模型行为与性能差距,我们证明用户反馈预测是一种可学习能力,基于真实环境反馈训练的奖励模型可实现良好校准的准确性。进一步揭示了大模型作为裁判评价协议的系统性偏差,并对比了直接影响用户体验的典型响应策略。UXBench建立了新的评估范式,呼吁更多关注定制化用户体验优化,推动形成以用户为中心的规模定律,塑造人工智能助手的成功路径。

原文摘要 · Abstract (English)

As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present UXBench, the first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation. The benchmark consists of three interconnected tasks, UX Judge, UX Eval, and UX Recovery, with 7,400 test instances extracted from over 70K interaction logs of a mainstream Chinese AI assistant. The dataset closely reflects real user distributions, covering 8 scenarios, 83 domains, and diverse failure patterns that pose severe challenges. Extensive experiments on 26 frontier language models provide novel insights into how well models perceive user experience and how improvements in model capability contribute to better dialogue engagement. Through comprehensive analysis of model behavior and performance gaps, we show that user feedback prediction is a learnable capability, where a reward model trained from in-the-wild feedback signals can achieve well-calibrated accuracy. We further document the systematic biases of LLM-as-a-judge evaluation protocols and compare typical response strategies that directly affect user experience. UXBench establishes a new evaluation landscape and calls for greater attention to tailored UX optimization, contributing to a user-centric scaling law that shapes the success of AI assistants.

用户体验对话评估基准测试反馈学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。