让AI模型像用户一样看推荐界面,提升模拟真实度。
Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation
- 用眼动数据对齐视觉语言模型的注意力,模仿用户注视习惯。
- 在三个探测器和两个模型结构上均提升注意力匹配度与点击预测准确率。
- 适合做推荐系统评估、个性化模拟的研究者和工程师。
大型语言模型代理正被广泛用于推荐系统的可扩展用户模拟。然而现有模拟器仅通过文本或结构化元数据感知推荐内容,而非真实用户浏览的视觉界面——这一关键差距在于,用户对推荐布局的注意力既受视觉驱动,又高度个性化。我们探究将视觉语言模型(VLM)的视觉注意力对齐于用户特定的眼动模式,能否提升模拟保真度。通过对一个基于轮播式推荐场景的真实眼动数据集分析发现,用户具有稳定且显著预测点击行为的个体化注视模式。基于此,我们提出用于用户模拟的注视对齐微调方法(FixATE)。该方法首先通过可解释性算子探测VLM内部视觉注意力,获得与人类注视可比的槽级相关性分布;随后学习个性化软提示,引导模型注意力朝向每位用户的特征注视模式。在三种基于可解释性的探测算子和两种不同架构的VLM主干网络上的实验均显示,在注意力对齐与点击预测准确率方面取得一致提升。结果表明,让模型‘像用户一样看’是实现更忠实还原用户感知与行为的可行路径。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly deployed as scalable user simulators for recommender system evaluation. Yet existing simulators perceive recommendations through text or structured metadata rather than the visual interfaces real users browse-a critical gap, since attention over recommendation layouts is both visually driven and highly personalized. We investigate whether aligning a vision-language model's (VLM's) visual attention with user-specific gaze patterns can improve simulation fidelity. Analysis of a real-world eye-tracking dataset collected in a carousel-based recommendation setting reveals that users exhibit stable individual gaze patterns strongly predictive of click behavior. Building on this finding, we propose Fixation-Aligned Tuning for user Emulation (FixATE). Our approach first probes the VLM's internal visual attention via interpretability operators to obtain a slot-level relevance distribution comparable with human fixation, and then learns personalized soft prompts to steer the model's attention toward each user's characteristic fixation pattern. Experiments across three interpretability-based probing operators and two architecturally distinct VLM backbones demonstrate consistent improvements in both attention alignment and click prediction accuracy. These results suggest that making the model "see like the user" is a viable path toward simulators that more faithfully reproduce how users perceive and act in recommendation interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。