arXiv:2504.14147cs.IRcs.AI2025-04被引 4

用AI模拟人类反馈,让推荐解释更懂用户需求。

Explainable Recommendation with Simulated Human Feedback

  • 用大模型模拟人类反馈,动态优化解释生成。
  • 在4个数据集上显著提升解释质量,效果优于传统方法。
  • 适合需要可解释推荐的系统研发与算法设计者。

近年来,可解释推荐通过揭示决策逻辑显著提升了用户体验。然而,现有方法因依赖稀疏交互数据中的传统监督学习范式,难以提供有效反馈信号来改进或淘汰生成的解释。为此,本文提出一种类人反馈驱动的优化框架,通过动态交互机制实现以人为本的解释需求,且无需高人力成本。具体地,利用大语言模型(LLMs)作为人类模拟器,预测类人反馈以指导学习过程。为使LLM深入理解任务本质并满足用户个性化需求,引入人类引导的定制化奖励评分方法,激发其语言理解与逻辑推理能力。针对解释质量多视角间的潜在冲突,提出基于帕累托优化的原则性多目标优化策略,将多维度质量提升转化为多目标优化问题。此外,为实现高效训练,设计了离策略优化管道,通过引入回放缓冲区并缓解数据分布偏差,显著提高数据利用率与模型泛化能力。在四个数据集上的大量实验验证了该方法的优越性。

原文摘要 · Abstract (English)

Recent advancements in explainable recommendation have greatly bolstered user experience by elucidating the decision-making rationale. However, the existing methods actually fail to provide effective feedback signals for potentially better or worse generated explanations due to their reliance on traditional supervised learning paradigms in sparse interaction data. To address these issues, we propose a novel human-like feedback-driven optimization framework. This framework employs a dynamic interactive optimization mechanism for achieving human-centered explainable requirements without incurring high labor costs. Specifically, we propose to utilize large language models (LLMs) as human simulators to predict human-like feedback for guiding the learning process. To enable the LLMs to deeply understand the task essence and meet user's diverse personalized requirements, we introduce a human-induced customized reward scoring method, which helps stimulate the language understanding and logical reasoning capabilities of LLMs. Furthermore, considering the potential conflicts between different perspectives of explanation quality, we introduce a principled Pareto optimization that transforms the multi-perspective quality enhancement task into a multi-objective optimization problem for improving explanation performance. At last, to achieve efficient model training, we design an off-policy optimization pipeline. By incorporating a replay buffer and addressing the data distribution biases, we can effectively improve data utilization and enhance model generality. Extensive experiments on four datasets demonstrate the superiority of our approach.

可解释推荐大模型强化学习反馈优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。