arXiv:2606.05697cs.AI2026-06被引 1

用AI模拟真实用户评估界面,让产品迭代更快更便宜。

PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX Evaluation

论文配图:PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX Evaluation
图 1 · 摘自论文原文
  • 基于角色设定的AI代理,模拟特定用户对界面的反应
  • 在多个数据集上达到与真人一致的评估真实度
  • 适合早期产品设计阶段快速验证界面体验

用户界面(UI)和用户体验(UX)评估是产品开发的核心,但可靠反馈仍依赖招募真人参与者或进行线上A/B测试,导致早期迭代缓慢且成本高昂。针对此问题,近期研究探索了多模态大模型作为代理评估者,但现有方法或仅提供表面评价,或反映模型自身偏见而非真实用户反应。本文提出PerceptUI框架,支持基于角色的界面评估,可预测特定用户对界面相关问题的回答,并生成自然语言推理。该框架分两阶段训练:(i) 对比反思微调,从人类决策中提炼教师生成的推理;(ii) 基于模型自身失败记录的反思式提示演化。在多个领域和数据集上,PerceptUI实现人类级真实度,能泛化至未见问题与角色,并生成具有代表性的群体响应分布。

原文摘要 · Abstract (English)

User interface (UI) and user experience (UX) evaluation is central to product development, yet reliable feedback still relies on recruiting human participants or running online A/B tests, making early-stage iteration slow and costly. In light of this, recent work has explored Multimodal Large Language Models as proxy evaluators. However, existing approaches either produce surface-level critiques or a judgment that reflects the model's own biases rather than the genuine response of a particular user. We introduce PerceptUI, a framework for persona-conditioned UI/UX evaluation that predicts how a specific user would answer interface-related questions and produces natural-language rationales. PerceptUI is trained in two stages: (i) contrastive reflection fine-tuning distills teacher-generated rationales by extracting lessons from human decisions, and (ii) a reflective prompt-evolution step from the model's own failure traces. Across multiple domains and datasets, PerceptUI achieves human-level realism, generalizes to unseen questions and personas, and yields population-level response distributions.

UI评估AI代理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。