arXiv:2608.01704cs.IRcs.CL2026-08

机器预测读者选句能力有上限,融合多模型能显著提升效果。

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?

  • 用拼接多个前沿模型+位置先验的方法提升预测性能
  • 融合方法比最强单模型高0.0159 AP,且在多种扰动下仍有效
  • 文档级结构是关键,8B小模型可保留90%融合优势

我们构建了人类读者在120篇网页文档中自发标记句子任务的上下界:下界为简单截断(首段);上界为半群组互预测的虚拟最优模型。两者差距为+0.2028 AP(置信区间[+0.1698, +0.2342],按领域聚类)。语义特征仅解释5%差距。前沿语言模型零样本表现达35-53%,但仍有明显不足;顶级提示压缩器LLMLingua-2反而低于随机选择。五模型无加权融合结合位置先验,达到60%,优于单模型+0.0159 AP(Holm p=0.019),且对最佳成员、分半选择、提示重写、标签/门控/种子扰动均鲁棒。预注册复制实验在217篇独立文档上验证成功(+0.0179, Holm p=0.042)。最终,将融合蒸馏为一个80亿参数的全文档读取学生模型,保留90%优势,与最强单模型统计等价(+0.0070 [-0.0068, +0.0200]),而局部上下文模型仅保留63%——说明群体信号存在于文档级结构中,最廉价提升方式是集成多个模型并平均。

原文摘要 · Abstract (English)

A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for a task with a rare kind of ground truth: predicting which sentences a crowd of readers -- highlighting for their own purposes, unpaid, uninstructed, and blind to each other -- marked in 120 web documents. The floor is naive truncation (lead); the ceiling is a split-half oracle: half the crowd predicting the other half. The gap between them is +0.2028 AP [+0.1698, +0.2342, domain-clustered], and three findings structure it. First, the gap is semantic: position and length features recover 5% of it. Second, frontier language models reach 35-53% of it zero-shot -- far above classical baselines, far below the crowd; a state-of-the-art prompt compressor (LLMLingua-2) lands below the floor, indistinguishable from random selection. Third, an unweighted cross-vendor fusion of five frontier rankings plus a position prior reaches 60%, beating the best single model by +0.0159 [+0.0044, +0.0269; Holm p=0.019] -- a gain that survives ablation of its best member, split-half arm selection, prompt paraphrase, and label, gate, and seed perturbations, and was CONFIRMED by a pre-registered replication on 217 independent documents (+0.0179, Holm p=0.042). Finally, the bracket compresses: distilling the fusion into one open-weight 8B student that reads the whole document retains 90% of the fusion's edge and reaches statistical parity with the strongest single frontier model (+0.0070 [-0.0068, +0.0200]), where a local-context student retains only 63% -- the crowd's signal lives in document-level structure, and the cheapest known improvement is to ask several different models and average.

注意力预测模型融合文本摘要评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。