解决对话推荐中头部内容过热、尾部内容稀缺的问题。
LumiCRS: Asymmetric Contrastive Prototype Learning for Long-Tail Conversational Recommender Systems
- 用动态加权损失和原型学习缓解头部过拟合与尾部稀疏。
- 在REDIAL和INSPIRED上召回率提升7-15%,尾部推荐显著改善。
- 适合关注推荐公平性与多样性的人士,尤其对冷启动场景有效。
对话推荐系统常面临极端长尾分布问题,导致对高频率热门内容偏倚严重,牺牲多样性并加剧冷启动。对DCRS的分析及REDIAL语料库统计显示,仅10%的热门电影占了近一半提及量,而约70%的尾部电影仅获26%关注度。这种失衡引发三大挑战:头部过拟合、主体表征漂移和尾部稀疏。为此,我们提出LumiCRS,一个端到端框架,通过三层协同机制缓解长尾失衡:(i) 自适应综合焦点损失(ACFL),动态调整类别权重与聚焦因子,抑制头部过拟合;(ii) 长尾推荐原型学习,选取语义、情感与上下文原型指导聚类,稳定主体与尾部表征;(iii) 基于GPT-4o的原型引导对话增强模块,自动生成多样化的尾部对话片段,缓解尾部稀疏与分布偏移。实验表明,该方法在REDIAL与INSPIRED基准上,相较十五个强基线,召回率@10与尾部召回率@10均提升7-15%,人工评估也验证其在流畅性、信息量与尾部相关性上的优势。结果证明多层协作在构建高效公平的长尾对话推荐系统中的有效性。
原文摘要 · Abstract (English)
Conversational recommender systems (CRSs) often suffer from an extreme long-tail distribution of dialogue data, causing a strong bias toward head-frequency blockbusters that sacrifices diversity and exacerbates the cold-start problem. An empirical analysis of DCRS and statistics on the REDIAL corpus show that only 10% of head movies account for nearly half of all mentions, whereas about 70% of tail movies receive merely 26% of the attention. This imbalance gives rise to three critical challenges: head over-fitting, body representation drift, and tail sparsity. To address these issues, we propose LumiCRS, an end-to-end framework that mitigates long-tail imbalance through three mutually reinforcing layers: (i) an Adaptive Comprehensive Focal Loss (ACFL) that dynamically adjusts class weights and focusing factors to curb head over-fitting and reduce popularity bias; (ii) Prototype Learning for Long-Tail Recommendation, which selects semantic, affective, and contextual prototypes to guide clustering and stabilize body and tail representations; and (iii) a GPT-4o-driven prototype-guided dialogue augmentation module that automatically generates diverse long-tail conversational snippets to alleviate tail sparsity and distribution shift. Together, these strategies enable LumiCRS to markedly improve recommendation accuracy, diversity, and fairness: on the REDIAL and INSPIRED benchmarks, LumiCRS boosts Recall@10 and Tail-Recall@10 by 7-15% over fifteen strong baselines, while human evaluations confirm superior fluency, informativeness, and long-tail relevance. These results demonstrate the effectiveness of multi-layer collaboration in building an efficient and fair long-tail conversational recommender.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。