研究发现个性化推荐在选文层面有效,但在句子层面无效。
Selection, Not Salience: The Shape and Limits of Personalization in Social Highlighting
- 用共读控制实验隔离个人偏好,验证选文时历史记录能准确预测标记
- 文档级个人化信号达+0.169,句级则无提升,甚至劣于基础模型
- 适合关注个性化边界与阅读行为建模的研究者
通过社交阅读标注工具与共读身份控制(同一文档被多人标记,固定文档与主题,检验个人历史是否比他人标记更优),我们刻画了个性化在不同阅读层级的形态与极限。在文档层级,获得干净、无泄露的控制测量:个人历史可识别共读邻域中的归属文档,自身与他人对比差距为+0.169,对抗主题匹配硬负样本为+0.119(均显著);内容驱动分析表明信号非仅标题驱动,而是以主题为主。该信号与先前工作句级选择信号(+0.14)量级相当,跨层级稳定,主要源于主题偏好。在句子层级,两阶段个性化自动标注(先生成候选,再个性化重排)未优于无个性基线:两个现成零样本大模型(含前沿模型)预测标注位置反而更差,个性化重排甚至败于显著性排序,即使在最高召回候选池中亦然,说明结果非第一阶段上限所致。可测量的个性化主要存在于选择层:适度(约+0.13)、主题主导,且在显著性层无可靠收益。还发现负样本控制偏差曾将文档差距虚增至+0.227,经审计后修正。超越共享显著性层的路径,或应通过聚合个体而非强化个性化实现。
原文摘要 · Abstract (English)
Does personalizing what a reader sees pay off, and where does it stop? Using a social web highlighter and a co-readership identity control (the same document highlighted by many users, which holds document and topic fixed and asks whether a person's own history predicts their marks better than another reader's does), we map the shape and limits of personalization across reading altitudes. At the document altitude we give the clean, leakage-free, identity-controlled measurement that prior next-document evaluations could only upper-bound: a person's history identifies which documents in a co-reading neighborhood are theirs, with an own-versus-other gap of +0.169 against community negatives and +0.119 against topic-matched hard negatives (both highly significant); a content-based arm suggests the signal is not purely title-driven but is largely thematic. This is comparable to the span-level selection signal (+0.14) from our prior work: the selection signal is of comparable magnitude across altitudes (+0.12 to +0.17), most of it stable topic preference. At the sentence altitude, a two-stage personalized auto-highlight (an impersonal model proposes candidates, a personal model re-ranks them) does not improve on its impersonal baseline: two off-the-shelf zero-shot LLMs, including a frontier model, predict highlight locations worse than a lead baseline, and personal re-ranking is beaten by the salience order even on the highest-recall candidate pool, so the null is not merely a Stage-1 ceiling artifact. Measurable personalization appears primarily at the selection layer: modest (~+0.13), topic-dominated, with no reliable gain at the salience layer. We also surface a control-in-negatives bias that inflated our document gap to a spurious +0.227 until audited. Going beyond the shared salience layer may be better approached by aggregating individuals than by personalizing them harder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。