拆解大模型对话中信息提取与隐藏的差异,揭示防御强、进攻弱的根源。
AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue
- 将多轮对话建模为博弈,分角色评估信息提取与隐藏能力
- 发现防御表现稳定,攻击表现差异大,约束违规占失败41.3%
- 适合研究模型推理机制与安全对齐的学者参考
多轮大模型评估常以单一胜率指标呈现,掩盖了不同能力。本文提出AIDG(对抗性信息推断游戏),将多轮对抗对话形式化为两人部分可观测随机博弈(POSG),并从探索者(信息提取)和持有者(信息隐藏)角色分解性能。该分解识别出三种失效模式:合作先验泄露、约束推理干扰、假设空间遍历低效。在6个前沿大模型上进行439场游戏测试显示,防御表现高度集中(σ=1.9 ELO),而进攻表现波动显著(σ=53.3 ELO);确认式提问使信息提取概率提升7.75倍(p<0.00001);约束违规导致41.3%的推断失败,且与模型规模无关(rho=0.0)。我们指出,隐藏优于提取的差距并非意外,而是局部可解的防御决策与全局耦合的进攻规划共同作用的结果,并基于分解实现按模型归因。所有设计(如回合衰减权重、Bradley-Terry评分模型)均源于明确假设。
原文摘要 · Abstract (English)
Multi-turn LLM evaluation is typically reported as a single win-rate scalar, conflating distinct capabilities. We introduce AIDG (Adversarial Information Deduction Game), formalizing multi-turn adversarial dialogue as a two-player partially observable stochastic game (POSG) and decomposing performance along Seeker (extraction) and Holder (containment) roles. The decomposition isolates three failure modes: cooperative-prior leakage, constraint-reasoning interference, and inefficient hypothesis-space traversal. Across 439 games over six frontier LLMs, defensive performance is tightly clustered (sigma = 1.9 ELO) while offensive performance varies substantially (sigma = 53.3 ELO); confirmation framing increases extraction odds 7.75x over uninformed deduction (p < 0.00001); and constraint violations account for 41.3% of deductive failures, uncorrelated with scale (rho = 0.0). We position the containment-over-extraction gap not as a surprising finding but as a measurable consequence of locally resolvable defensive decisions versus globally coupled offensive planning, and use the decomposition to attribute the gap per model. All design choices, including turn-decay weighting and the Bradley-Terry rating model, are derived from explicit assumptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。