arXiv:2609.08698cs.CL2026-09

论文揭示了分组方式如何影响语言模型对证据的权重判断。

Record Grouping Controls Evidence Weight in Language Models

论文配图:Record Grouping Controls Evidence Weight in Language Models
图 1 · 摘自论文原文
  • 通过设计分组策略,实现组内去重并聚合内容
  • 错误分组导致准确率下降10.27至32.66个百分点
  • 适用于需控制证据呈现的AI决策研究

检索记录是呈现单元;给定的分组决定了哪些记录作为单一证据输入语言模型。我们刻画了不变的组内内容状态,该状态可消除组内重复内容,同时保留互补的规范内容,并推导出一个精确的内容感知分组误差上界。在给定分组下,我们的预生成表示在组内去重并聚合内容,同时对每组贡献进行边界约束。在104,402次试验和6个公开检查点中,固定内容的错误分组使准确率下降10.27–32.66个百分点,错误合并则减少9.13–31.79个百分点;匹配的六槽对照组在所有16个单元中均保持正向趋势。在新的48项受控实验面板中,改变分组方式在所有四个模型中均引发可测量且依赖检查点的决策变化,平衡镜像设计揭示了显著的顺序交互效应。理论与实验共同确立了分组为可控的预生成表示变量,并刻画其依赖检查点的行为效应。

原文摘要 · Abstract (English)

Retrieved records are presentation units; a supplied partition determines which records enter a language model as one evidential contribution. We characterize the invariant group-content state that removes within-group copies while retaining complementary canonical content, show that equal group counts can encode different evidence states, and derive a sharp content-aware partition-error bound. Given a supplied partition, our pre-generation representation deduplicates and aggregates content within groups and bounds each group's contribution. Across 104,402 trials and 6 public checkpoints, a central natural-text intervention finds that content-fixed false splits add 10.27-32.66 percentage points and false merges remove 9.13-31.79 points; a matched six-slot control retains the positive direction in all 16 cells. In a new 48-item controlled campaign panel, changing the supplied partition produces measurable, checkpoint-dependent decision shifts across all four models, and the balanced mirror design exposes substantial order interactions. Together, the theory and experiments establish the supplied partition as a controllable pre-generation representation variable and characterize its checkpoint-dependent behavioral effects.

语言模型证据权重分组策略可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。