用户更关注音乐推荐的流行度分布,但未必偏好校准结果。
Robustness and User-Perceived Value of Popularity Calibration in Music Recommendation: A User Study
- 用不同流行度组成的推荐列表做对照实验,测试用户感知差异。
- 用户能分辨流行度差异,但不明显偏爱校准列表,仅在熟悉度高时有轻微偏好。
- 计算流行度与用户判断相关性弱,且受历史数据和熟悉度影响大。
流行度校准在推荐系统中被视为用户中心化个性化或流行度偏差的指标。现有研究多依赖离线指标,假设用户偏好与其历史消费流行度匹配的推荐列表。然而,关于校准的用户研究仍有限,且已有发现表明校准推荐未必显著提升用户体验。此外,尽管先前工作显示校准指标与用户对推荐列表的感知存在关联,但这种关系在不同项目熟悉度和用户历史信息不完整条件下是否稳健尚不明确。本研究通过构建基于用户近期听歌记录的个性化曲目列表,使用一个控制型简单推荐器生成高流行、低流行及校准三类列表,考察用户是否能感知其差异,是否偏好校准列表,JSD-based 流行度校准在不同熟悉度与历史数据可用性下的可靠性,以及计算流行度标签与用户主观判断的一致性。结果表明:用户可感知流行度分布差异,但未表现出对校准列表的明确偏好;JSD 与感知流行度的关系受项目熟悉度、列表构成和可用历史数据的影响;计算流行度与用户判断仅有弱相关性。这些发现有助于更批判性地理解流行度校准作为离线指标与面向用户的构造的有效性。
原文摘要 · Abstract (English)
Popularity calibration in recommender systems has been studied both as a form of user-centered personalization and as an indicator of popularity bias. Most existing work evaluates calibration through offline metrics, often assuming that users prefer recommendation lists whose popularity distribution matches their historical consumption profile. However, user studies on calibration remain limited, and existing findings suggest that calibrated recommendations do not necessarily have a strong effect on user experience. Moreover, although prior work has shown that calibration metrics can correlate with users' perceptions of recommendation lists, the robustness of this relation remains unclear under different levels of item familiarity and incomplete user-history information. In this work, we study the perceived value and measurement reliability of popularity calibration in music recommendation. We construct personalized track lists from users' recent listening histories and use a controlled naive recommender to create lists with different popularity compositions: highpop-heavy, lowpop-heavy, and calibrated. We investigate whether users perceive differences between these lists, whether calibrated lists are preferred, how robust JSD-based popularity calibration is under different familiarity and history-availability conditions, and how computational popularity labels align with users' own popularity judgments. Our results show that users perceive differences in popularity composition, but do not clearly prefer calibrated lists. We further find that the relation between JSD and perceived popularity depends on item familiarity, list composition, and available user history, while computational and user-judged popularity labels only weakly align. These findings contribute to a more critical understanding of popularity calibration as both an offline metric and a user-facing construct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。