arXiv:2601.04554cs.IR2026-01KDD被引 6

用多模态用户代理模拟真实行为,低成本替代推荐系统A/B测试

Exploring Recommender System Evaluation: A Multi-Modal User Agent Framework for A/B Testing

  • 构建推荐沙盒环境,支持多模态跨页面交互
  • 生成数据可提升推荐模型性能,验证替代传统A/B测试可行性
  • 适合需要快速评估推荐算法的工业研究者

在推荐系统中,线上A/B测试是评估模型性能的关键方法,但存在成本高、用户体验下降和耗时长等挑战。基于大语言模型的智能体具备替代传统A/B测试的潜力,但现有方法难以模拟用户的感知过程与交互模式,缺乏真实环境和视觉感知能力。为此,我们提出多模态用户代理(A/B Agent),构建用于A/B测试的推荐沙盒环境,支持与真实平台一致的多模态、多页面交互。该代理融合多模态感知、细粒度用户偏好,结合用户画像、行为记忆检索与疲劳机制,模拟复杂人类决策过程。从模型、数据、特征三个维度验证了其作为传统A/B测试替代方案的潜力。进一步发现,由A/B Agent生成的数据能有效增强推荐模型能力。代码已公开于https://github.com/Applied-Machine-Learning-Lab/ABAgent。

原文摘要 · Abstract (English)

In recommender systems, online A/B testing is a crucial method for evaluating the performance of different models. However, conducting online A/B testing often presents significant challenges, including substantial economic costs, user experience degradation, and considerable time requirements. With the Large Language Models' powerful capacity, LLM-based agent shows great potential to replace traditional online A/B testing. Nonetheless, current agents fail to simulate the perception process and interaction patterns, due to the lack of real environments and visual perception capability. To address these challenges, we introduce a multi-modal user agent for A/B testing (A/B Agent). Specifically, we construct a recommendation sandbox environment for A/B testing, enabling multimodal and multi-page interactions that align with real user behavior on online platforms. The designed agent leverages multimodal information perception, fine-grained user preferences, and integrates profiles, action memory retrieval, and a fatigue system to simulate complex human decision-making. We validated the potential of the agent as an alternative to traditional A/B testing from three perspectives: model, data, and features. Furthermore, we found that the data generated by A/B Agent can effectively enhance the capabilities of recommendation models. Our code is publicly available at https://github.com/Applied-Machine-Learning-Lab/ABAgent.

推荐系统A/B测试多模态智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。