评测大模型生成的用户体验建议是否可执行,发现不同模型表现差异明显。
UXBench: Measuring the Actionability of LLM-Generated UX Critiques

- 构建可运行的网页样例与浏览器探索机制,强制模型收集操作证据后再评价。
- 8个前沿模型在修复成功率上差异显著,最高提升达32%。
- 适合想评估AIUX建议实用性的产品团队或研究者使用。
大型语言模型(LLMs)正被用于作为用户体验(UX)评判者,检查界面、诊断可用性问题并提出修复建议。然而,目前尚无可控基准来衡量这些批判在多样化产品界面中的可靠性与可操作性。我们提出了 UXBench,一个评估 LLM 作为交互基础的 UX 判官的基准。UXBench 包含十类产品界面的本地可运行网页样例,并结合覆盖约束的浏览器探索,迫使模型在报告前收集交互证据。每个模型生成包含七个维度的结构化 UX 报告;报告质量通过固定下游修复代理能否基于该批判改进界面来衡量。我们在自动化修复提升协议和盲测人类验证研究中评估了八个前沿模型。结果表明,用户体验评判既非饱和也非单一维度:各模型在报告可操作性上存在显著差异,表现出独特的修复特征模式,在不同样例上的可靠性各异,并在各类界面中轮流领先。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs. Yet no controlled benchmark measures whether the resulting critiques are reliable and actionable across heterogeneous product surfaces. We introduce UXBench, a benchmark for evaluating LLMs as interaction-grounded UX judges. UXBench comprises local-first runnable web fixtures spanning ten product-surface families, paired with coverage-gated browser exploration that forces models to collect interaction evidence before reporting. Each judge model produces a structured UX report over seven rubric dimensions; report quality is measured by whether a fixed downstream repair agent can improve the interface based on the critique. We evaluate eight frontier models under both an automated repair-lift protocol and a blind human validation study. Results show that UX judging is neither saturated nor one dimensional: models differ meaningfully in report actionability, exhibit distinct rubric-level repair signatures, vary in fixture-level reliability, and trade leadership across surface categories
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。