提出新评估框架,测试推荐系统对自然语言偏好指令的响应能力。
SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation
- 用自然语言指令干预推荐,测试系统可调节性
- 涵盖从类型到内容警告的多样干预,评估更精细控制能力
- 为个性化推荐设计提供实证指导,适合关注可解释推荐的研究者
自然语言用户画像因提升可解释性并增强推荐系统的可调节性而受到关注。通过直接编辑,用户可明确表达难以从历史行为推断的偏好。然而,当前基于自然语言的推荐方法是否能有效响应这些调节指令尚不明确。现有评估主要针对电影类型等明确属性,但无法捕捉推动可调节推荐的复杂用户控制需求。为此,我们提出SteerEval,一个评估框架,通过从类型到内容警告等多样化干预手段,衡量推荐系统在更细微、多样的调节形式下的表现。我们评估了一类预训练自然语言推荐模型的调节能力,分析了在较冷门主题上的调节潜力与局限,并比较不同画像和推荐干预对调节效果的影响。最后,基于发现提出实用设计建议,并讨论未来可调节推荐系统的发展方向。
原文摘要 · Abstract (English)
Natural-language user profiles have recently attracted attention not only for improved interpretability, but also for their potential to make recommender systems more steerable. By enabling direct editing, natural-language profiles allow users to explicitly articulate preferences that may be difficult to infer from past behavior. However, it remains unclear whether current natural-language-based recommendation methods can follow such steering commands. While existing steerability evaluations have shown some success for well-recognized item attributes (e.g., movie genres), we argue that these benchmarks fail to capture the richer forms of user control that motivate steerable recommendations. To address this gap, we introduce SteerEval, an evaluation framework designed to measure more nuanced and diverse forms of steerability by using interventions that range from genres to content-warning for movies. We assess the steerability of a family of pretrained natural-language recommenders, examine the potential and limitations of steering on relatively niche topics, and compare how different profile and recommendation interventions impact steering effectiveness. Finally, we offer practical design suggestions informed by our findings and discuss future steps in steerable recommender design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。