arXiv:2509.07269cs.IRcs.AI2025-09被引 2

构建两个含敏感标签的推荐数据集,助力评估系统对有害内容的推荐风险。

Datasets for Navigating Sensitive Topics in Recommendation Systems

  • 融合影评与内容警告,构建电影推荐敏感标签数据集
  • 整合粉丝创作互动与用户自定义警示,形成文学类敏感数据集
  • 为研究者提供超越点击率的伦理评估工具,适合安全推荐方向

个性化AI系统(如推荐系统、聊天机器人)根据用户偏好分发内容,但存在暴露用户至敏感或有害信息的风险,可能影响整体福祉。为量化评估此类风险,需构建带有敏感性标签的内容数据集,使研究者能超越单纯参与度指标进行评估。本文提出两个新数据集:一是将MovieLens评分数据与Does the Dog Die?社区的内容警告结合;二是整合Archive of Our Own上的粉丝小说互动数据与用户生成的警示信息,均包含敏感性标签分类体系,支持对推荐系统在敏感话题上的表现进行系统性研究。

原文摘要 · Abstract (English)

Personalized AI systems, from recommendation systems to chatbots, are a prevalent method for distributing content to users based on their learned preferences. However, there is growing concern about the adverse effects of these systems, including their potential tendency to expose users to sensitive or harmful material, negatively impacting overall well-being. To address this concern quantitatively, it is necessary to create datasets with relevant sensitivity labels for content, enabling researchers to evaluate personalized systems beyond mere engagement metrics. To this end, we introduce two novel datasets that include a taxonomy of sensitivity labels alongside user-content ratings: one that integrates MovieLens rating data with content warnings from the Does the Dog Die? community ratings website, and another that combines fan-fiction interaction data and user-generated warnings from Archive of Our Own.

推荐系统敏感内容数据集伦理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。