提出新评估框架,让推荐系统不再黑箱,可被用户引导。
Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

- 用多智能体协作评估推荐系统可控性
- 发现长尾内容难引导是核心瓶颈
- 适合研究者、监管者和想掌控推荐的人
推荐系统如黑箱,用户与监管者难以引导其输出或审计行为。本文提出CtrlBench-Rec框架,首次系统评估推荐系统的可控性。通过三个任务——目标内容发现、兴趣画像塑造、流行度偏差缓解,衡量从明确指令到隐式表征控制再到算法偏见克服的可引导能力。在真实数据集和多个推荐模型上实验表明,该框架能有效量化可控性,并暴露关键缺陷,尤其发现系统对长尾内容存在顽固抗拒。本工作提供首个标准化工具包,助力可控推荐研究、算法审计与用户赋权。代码已开源:https://github.com/caskcsg/CtrlBenchRec。
原文摘要 · Abstract (English)
Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。