arXiv:2605.26174cs.SEcs.AI2026-05

大模型协同写作时,跨段落缺陷检测能力普遍骤降,且越对齐越易误报。

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

  • 通过多代理协同框架测试十款大模型,发现跨段落矛盾检测率普遍下降三分之二以上。
  • 仅一家厂商的模型在对齐增强后出现缺陷漏检减少但误报上升的双面效应,且与生成版本强相关。
  • 模型私有记录能准确识别结构错误,但整合报告却盲目认可,暴露协同机制根本缺陷。

生产级大模型系统通过隐藏的代理协作链路处理请求,并整合成一份综合报告。我们研究这种架构对单个代理无法察觉的一类缺陷——文档中远距离段落间的逻辑矛盾——的影响。在保持文档、缺陷、机制、评分标准和种子固定的前提下,我们测试了同一开发者五代共十款模型,以及来自五个不同对齐范式的五家提供商模型。结果发现:第一,存在普遍的检测悬崖现象——所有在单代理下能检测缺陷的模型,在协同架构下其检测能力均显著下降,降幅超过三分之二,且不受规模或长程推理增强影响;第二,坠落后的模型行为差异明显:信号检测分解显示,在六款仍具一定区分能力的模型中,仅一家厂商的模型随对齐强化表现出缺陷漏检减少但干净文档误报增多的特征,该趋势在该厂商内部生成版本间高度显著(p < 0.001),其余厂商基本不存在此现象。在最低水平时,缺陷常未被忽略——模型私有记录可准确重建结构故障,但整合报告却认定其无问题,因关注点转向人工产物与缺席合作者。该现象难以量化:自动评判者精度仅17%-50%,关键词无法将其与普通共识区分。我们将其视为一项关键发现,并公开所有实验运行、探测工具、缺陷密钥、评分提示及脚本。结论表明:整合报告的置信度对跨段落缺陷无效,最对齐的系统并非最安全,且该悬崖是结构性的。

原文摘要 · Abstract (English)

Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated report. We ask what this does to a class of defect no single worker can see: a contradiction in the relation between two distant sections of a document. Holding the documents, defects, mechanism, scoring, and seed fixed, we vary only the model -- ten systems across five generations from one developer and five providers from distinct alignment paradigms. Two layers separate. First, a universal detection cliff: every model that finds these cross-section defects under a single agent loses that ability under orchestration, detection falling two-thirds or more across every paradigm tested. The cliff is mechanism-derived and not closed by scale or extended reasoning. Second, how models behave once fallen. A signal-detection decomposition shows that, among the six models discriminating above chance, only one developer's generations move along the reporting-criterion axis: as alignment is strengthened, the model misses fewer defects yet raises more false alarms on clean documents -- two faces of one criterion shift, scaling with generation within that developer (p < 0.001) and near-absent elsewhere. At the floor the missed defect is often not out of view: the model's private record reconstructs the structural fault accurately, while the integrated report signs off on its soundness, its concern spent on the artifact and an absent collaborator. This resists quantification -- an automated judge is unstable (precision 17-50%) and keywords cannot separate it from ordinary agreement -- a resistance we report as a finding. We release all runs, probes, defect keys, scorer prompts, and scripts. An integrated report's confidence is uninformative about partition-spanning defects, the most aligned systems are not the safest, and the cliff is structural.

大模型安全缺陷检测协同机制对齐风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。