机器预测读者选句能力有上限,融合多模型能显著提升效果。
Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?
- 用拼接多个前沿模型+位置先验的方法提升预测性能
- 融合方法比最强单模型高0.0159 AP,且在多种扰动下仍有效
- 文档级结构是关键,8B小模型可保留90%融合优势
我们构建了人类读者在120篇网页文档中自发标记句子任务的上下界:下界为简单截断(首段);上界为半群组互预测的虚拟最优模型。两者差距为+0.2028 AP(置信区间[+0.1698, +0.2342],按领域聚类)。语义特征仅解释5%差距。前沿语言模型零样本表现达35-53%,但仍有明显不足;顶级提示压缩器LLMLingua-2反而低于随机选择。五模型无加权融合结合位置先验,达到60%,优于单模型+0.0159 AP(Holm p=0.019),且对最佳成员、分半选择、提示重写、标签/门控/种子扰动均鲁棒。预注册复制实验在217篇独立文档上验证成功(+0.0179, Holm p=0.042)。最终,将融合蒸馏为一个80亿参数的全文档读取学生模型,保留90%优势,与最强单模型统计等价(+0.0070 [-0.0068, +0.0200]),而局部上下文模型仅保留63%——说明群体信号存在于文档级结构中,最廉价提升方式是集成多个模型并平均。
原文摘要 · Abstract (English)
A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for a task with a rare kind of ground truth: predicting which sentences a crowd of readers -- highlighting for their own purposes, unpaid, uninstructed, and blind to each other -- marked in 120 web documents. The floor is naive truncation (lead); the ceiling is a split-half oracle: half the crowd predicting the other half. The gap between them is +0.2028 AP [+0.1698, +0.2342, domain-clustered], and three findings structure it. First, the gap is semantic: position and length features recover 5% of it. Second, frontier language models reach 35-53% of it zero-shot -- far above classical baselines, far below the crowd; a state-of-the-art prompt compressor (LLMLingua-2) lands below the floor, indistinguishable from random selection. Third, an unweighted cross-vendor fusion of five frontier rankings plus a position prior reaches 60%, beating the best single model by +0.0159 [+0.0044, +0.0269; Holm p=0.019] -- a gain that survives ablation of its best member, split-half arm selection, prompt paraphrase, and label, gate, and seed perturbations, and was CONFIRMED by a pre-registered replication on 217 independent documents (+0.0179, Holm p=0.042). Finally, the bracket compresses: distilling the fusion into one open-weight 8B student that reads the whole document retains 90% of the fusion's edge and reaches statistical parity with the strongest single frontier model (+0.0070 [-0.0068, +0.0200]), where a local-context student retains only 63% -- the crowd's signal lives in document-level structure, and the cheapest known improvement is to ask several different models and average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。