为生成式推荐系统设计全面评估框架,兼顾效果与安全。
Toward Holistic Evaluation of Recommender Systems Powered by Generative Models
- 提出场景化多指标评估体系,覆盖相关性、事实性等维度。
- 识别出生成内容引发的新型风险,如虚构商品和矛盾解释。
- 适合研究生成式推荐系统的安全与可靠性问题者阅读。
由生成模型驱动的推荐系统(Gen-RecSys)突破传统物品排序,生成开放式内容,既提升个性化体验,也带来新风险。一方面,系统可通过动态解释和多轮对话增强用户参与;另一方面,可能产生虚构商品、放大偏见或泄露隐私。传统准确率指标无法衡量事实正确性、内容安全性和意图对齐性。本文首先将评估挑战分为两类:(i) 生成输出加剧的旧问题(如偏见、隐私);(ii) 生成带来的全新风险(如物品幻觉、矛盾解释)。其次,提出包含场景化评估与多指标检测的综合性评估方法,涵盖相关性、事实性、偏见检测和政策合规性。旨在为研究人员与实践者提供全面评估框架,确保推荐系统的有效个性化与负责任部署。
原文摘要 · Abstract (English)
Recommender systems powered by generative models (Gen-RecSys) extend beyond classical item ranking by producing open-ended content, which simultaneously unlocks richer user experiences and introduces new risks. On one hand, these systems can enhance personalization and appeal through dynamic explanations and multi-turn dialogues. On the other hand, they might venture into unknown territory-hallucinating nonexistent items, amplifying bias, or leaking private information. Traditional accuracy metrics cannot fully capture these challenges, as they fail to measure factual correctness, content safety, or alignment with user intent. This paper makes two main contributions. First, we categorize the evaluation challenges of Gen-RecSys into two groups: (i) existing concerns that are exacerbated by generative outputs (e.g., bias, privacy) and (ii) entirely new risks (e.g., item hallucinations, contradictory explanations). Second, we propose a holistic evaluation approach that includes scenario-based assessments and multi-metric checks-incorporating relevance, factual grounding, bias detection, and policy compliance. Our goal is to provide a guiding framework so researchers and practitioners can thoroughly assess Gen-RecSys, ensuring effective personalization and responsible deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。