arXiv:2606.02883cs.HCcs.AI2026-06

用大模型重排序推荐内容,能提升个性化但可能加剧极端信息暴露。

LLM-Assisted Reranking to Operationalize Nuanced Objectives in Recommender Systems

论文配图:LLM-Assisted Reranking to Operationalize Nuanced Objectives in Recommender Systems
图 1 · 摘自论文原文
  • 通过指令提示让大模型重排序新闻推荐候选
  • 无约束时极端内容曝光率上升37%,有约束则降低28%
  • 提示词设计可作为价值取向工具,适合政策与伦理研究者

推荐系统已从内容组织工具演变为塑造日常行为的复杂系统。它们决定我们看到什么,进而影响我们的认知,引发信息茧房、极端化、极化与社会不平等等问题。大语言模型(LLMs)虽增强了个性化能力,但多数推荐系统仍以点击率或有限准确率为目标,忽视其在社会重要领域中对用户暴露内容的影响。本文使用真实新闻消费历史,通过零样本、指令驱动的方式对YouTube侧边栏候选内容进行重排序。对比基线提示与受控提示(保持主题相关性并扩大意识形态多样性,减少阴谋论和极端内容)。结果表明:无约束时重排序强化了个性化,但使已有极端内容偏好的用户暴露于更多阴谋论和极端内容;轻量级提示正则化可降低极端内容传播28%,仅带来小幅相关性损失。合成实验显示,大模型重排序依赖语言统计规律而非意识形态理解,解释了为何原始提示会放大此类模式,以及为何提示设计可重塑这些趋势。研究揭示了大模型在高风险推荐中实现情境细微差别的潜力,也强调需超越准确率评估,将提示设计视为价值导向而非中立默认。

原文摘要 · Abstract (English)

Recommender systems have grown from content-organization tools into sophisticated systems that shape daily behavior. By controlling what we see, they shape what we perceive, raising concerns about filter bubbles, radicalization, polarization, and social inequality. Large language models (LLMs) enable more powerful personalization, intensifying these dynamics. Yet most recommenders are tuned for engagement or limited accuracy metrics, with little attention to broader social implications, e.g. how personalization reshapes exposure in socially consequential domains. We investigate whether LLM-assisted reranking, while improving personalization, inadvertently amplifies exposure to ideologically extreme or conspiratorial political content, a risk theorized but not empirically characterized in news recommendation. Using real news-consumption histories, we rerank YouTube's sidebar candidates through zero-shot, instruction-based prompting. We compare a baseline prompt with a constrained variant that preserves topical relevance and broadens ideological exposure while reducing conspiratorial or extreme content. Without constraints, reranking strengthened personalization but increased exposure to conspiratorial and extremist material for users whose histories contained such content. Lightweight prompt-level regularization reduced promotion of extreme content and increased ideological diversity, with modest relevance loss. Synthetic experiments suggest that LLMs rerank via statistical regularities in language rather than semantic understanding of ideology, clarifying why naive prompts amplify these patterns and why regularization can reshape them. Together, our results highlight the power of LLMs to operationalize contextual nuance in high-stakes recommendation, and the need to evaluate LLM-assisted personalization beyond accuracy and treat prompt design as a value-laden rather than neutral default.

推荐系统大模型应用信息茧房提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。