arXiv:2606.07316cs.MAcs.AI2026-06

让大模型团队达成可验证的语义共识,关键在于推理过程的可信性。

Certifiable Semantic Agreement Among LLM Agents: What the Admissibility Instrument Decides

论文配图:Certifiable Semantic Agreement Among LLM Agents: What the Admissibility Instrument Decides
图 1 · 摘自论文原文
  • 通过三类输出协议检测语义一致性,构建可验证的集体判断机制。
  • 词法判断器在80%召回率下达到0.982的AUROC,显著优于微调模型。
  • 发现并修复协议中的平局漏洞,证明语义共识不提升安全边界。

大模型代理团队能否在语义层面而非标签层面达成可验证的一致?我们构建了协议来检验。H-CSC每轮输出三类结果——语义承诺、裁决承诺或类型中止,基于统一的2f+1签名证书。研究发现:当语义核心足够大时,裁决置信度已超f,此时与证书包裹的多数裁定无覆盖差异。实验显示两者承诺任务集完全一致。真正关键的是可接受性工具:面对仅篡改推理但保留裁决的对手,442 MB微调编码器在0-8%真阳性率下获AUROC 0.621-0.744(5%诚实假阳性率);而无需训练的词法谓词在38-80%真阳性率下达0.865-0.982(50任务,200次攻击,400个诚实代理),且完全确定性,推翻了摘要证明依赖的假设。诚实代理间距离大于攻击扰动(90百分位诚实角距离0.992弧度,攻击半径0.65),因此自由文本理由上无法使用嵌入过滤——这是输出结构的固有属性,收紧后分散度降低二十倍。承诺摘要并非静止:在非披露条件下,保序压缩图可区分忠实与替换衍生陈述,AUROC达0.741(n=350),而密码哈希和仅裁决摘要均得0.500。此外报告了自身协议中的平局捕获漏洞及其四行修复方案。我们不主张语义共识带来安全或覆盖优势;该结论由包含引理解释。

原文摘要 · Abstract (English)

Can a committee of LLM agents reach agreement that is certifiable at the level of meaning, not only at the level of a label? We build a protocol to find out. H-CSC emits one of three typed outcomes per round -- semantic commit, verdict commit, or typed abort -- under a common 2f+1 distinct-signer certificate, and we use it to measure what such agreement costs and buys. The answer is conditional, and the condition is not the protocol. We prove a containment lemma: whenever the semantic core is large enough to make the committed verdict deterministically valid, the verdict margin already exceeds f, so at matched deterministic guarantees no coverage separation from certificate-wrapped majority is possible. Measurement agrees: the two rules commit the identical task set. What matters instead is the admissibility instrument. Against adversaries that preserve the verdict and corrupt only the reasoning, a 442 MB fine-tuned encoder reaches AUROC 0.621-0.744 at 0-8% TPR (5% honest FPR), while a training-free lexical predicate reaches 0.865-0.982 at 38-80% (50 tasks, 200 attacks, 400 honest) and is exactly deterministic, discharging an assumption the digest proofs rely on. Honest agents disperse further than attacks displace (90th-percentile honest angular distance 0.992 rad against a 0.65 radius), so on free-text rationales no embedding filter is viable -- a property of the output schema, which when tightened collapses dispersion twenty-fold. The committed digest is not inert: under non-disclosure a similarity-preserving sketch separates faithful from substituted derived statements at AUROC 0.741 (n=350), where a cryptographic hash and a verdict-only digest score exactly 0.500. We also report a tie-break capture vulnerability found in our own protocol, and its four-line fix. We claim no safety or coverage advantage over verdict-only certification; the lemma explains why there is none to claim.

大模型协作语义共识可验证性对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。