研究大模型生成的音乐品味画像是否偏倚,影响推荐可信度。
Biases in LLM-Generated Musical Taste Profiles for Recommendation
- 用用户听歌记录生成自然语言品味画像,评估其准确性
- 发现主流性与口味多样性影响用户对画像的认同度
- 适合关注可解释推荐系统公平性的研究者阅读
大型语言模型(LLM)在推荐系统中的一个有前景应用是基于用户消费数据自动生成自然语言(NL)品味画像。这类画像提供了比黑箱协同过滤更透明、可编辑的替代方案,有助于提升系统可信度与用户控制力。然而,用户是否认为这些画像真实反映自身品味仍不明确,这对信任和可用性至关重要。此外,由于LLM继承社会与数据偏差,画像质量可能在不同用户属性(如主流性、口味多样性)和物品特征(如流派、原产国)下系统性差异。本文聚焦音乐流媒体场景,开展用户研究,让用户评估基于自身听歌历史生成的画像。分析画像认同度是否受用户属性与物品特征影响,并对比其在下游推荐任务中的表现。结果揭示了可解释性LLM画像在个性化系统中的潜力与局限。
原文摘要 · Abstract (English)
One particularly promising use case of Large Language Models (LLMs) for recommendation is the automatic generation of Natural Language (NL) user taste profiles from consumption data. These profiles offer interpretable and editable alternatives to opaque collaborative filtering representations, enabling greater transparency and user control. However, it remains unclear whether users consider these profiles to be an accurate representation of their taste, which is crucial for trust and usability. Moreover, because LLMs inherit societal and data-driven biases, profile quality may systematically vary across user and item characteristics. In this paper, we study this issue in the context of music streaming, where personalization is challenged by a large and culturally diverse catalog. We conduct a user study in which participants rate NL profiles generated from their own listening histories. We analyze whether identification with the profiles is biased by user attributes (e.g., mainstreamness, taste diversity) and item features (e.g., genre, country of origin). We also compare these patterns to those observed when using the profiles in a downstream recommendation task. Our findings highlight both the potential and limitations of scrutable, LLM-based profiling in personalized systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。