arXiv:2606.05621cs.IR2026-06

用智能体主动生成噪声标签,让推荐系统更准地识别用户真实偏好。

ANCHOR: Agentic Noise Creation Framework for Human Simulation and Denoising Recommendation

论文配图:ANCHOR: Agentic Noise Creation Framework for Human Simulation and Denoising Recommendation
图 1 · 摘自论文原文
  • 通过智能体模拟用户行为,主动创建带标签的噪声数据。
  • 在真实数据上测试,显著提升推荐准确率,尤其对边界模糊的噪声效果好。
  • 适合做推荐系统优化的研究者和工程师,尤其关注噪声处理场景。

从嘈杂的隐式反馈中提炼准确的用户偏好仍是推荐系统的核心挑战,亟需推荐去噪技术。然而真实数据缺乏显式噪声标注,现有方法多依赖无监督辅助信息或手工设计启发式规则,导致外部成本高、泛化能力差或依赖不可靠先验,造成噪声误判并破坏真实偏好表示。为此,我们提出推荐去噪的范式革新:不依赖启发式推断,而是通过‘生成-识别’范式主动创建带标签的噪声交互,并训练专用识别器进行监督学习,将去噪从启发式过滤转变为监督学习。基于此,我们提出ANCHOR——一种受大模型作为用户研究启发的代理框架。该框架分两阶段运行:在噪声生成阶段,采用推荐系统内循环的代理架构,合成多样化的非偏好噪声与信息丰富的边界邻近噪声;对于非偏好噪声,设计五种可扩展的模拟机制以逼近主要噪声来源;对于边界邻近噪声,引入对抗性边界精炼机制生成具有挑战性的模糊交互,精准定位决策边界。在噪声识别阶段,利用生成标签训练一个可复用的参数化识别器,融合协同信号与语义表示,在真实交互数据中检测噪声模式。

原文摘要 · Abstract (English)

Distilling accurate user preferences from noisy implicit feedback remains a fundamental bottleneck in recommendation systems, highlighting the need for recommendation denoising. However, real-world data lack explicit noise annotations, forcing existing methods to rely on unsupervised side information or handcrafted heuristics. These approaches often incur high external costs, generalize poorly, or depend on unreliable priors, causing noise misidentification and corrupting true user preference representations. To address these limitations, we propose a paradigm-level reformulation of recommendation denoising. Instead of indirectly inferring noisy interactions through heuristics, our Creation-Recognition paradigm proactively creates labeled noisy interactions and trains a dedicated recognizer to identify them, transforming denoising from heuristic filtering into supervised learning. Based on this paradigm, we present ANCHOR, an agent-based framework inspired by recent LLM-as-User research. ANCHOR simulates user behaviors to generate realistic noise labels and enables supervised denoising through two stages: noise creation and noise recognition. In the noise creation stage, ANCHOR adopts a recommender-in-the-loop agentic architecture to synthesize both diverse out-of-preference noise and informative boundary-adjacent noise. For out-of-preference noise, it implements five extensible simulation mechanisms to approximate major sources of noisy implicit feedback. For boundary-adjacent noise, an adversarial boundary refinement mechanism generates ambiguous interactions that challenge the recognizer and target the decision boundary. In the noise recognition stage, ANCHOR leverages the generated labels to train a reusable parametric recognizer that integrates collaborative signals and semantic representations to detect noise patterns in real interaction data.

推荐系统去噪智能体生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。