用信息操控理论细粒度检测大模型欺骗行为
DECOR: Auditing LLM Deception via Information Manipulation Theory

- 将信息拆解为原子单元,从四个维度评分
- 在多轮和单轮任务中均达顶尖效果
- 可解释性强,适合安全审计与模型评估
大语言模型可通过微妙的信息操控——如省略关键事实、转移焦点或模糊意义——实现欺骗,此类行为难以察觉。现有黑箱方法依赖粗粒度判断,缺乏可解释性,且无法定位被扭曲的事实及其方式。我们提出DECOR,一种基于信息操控理论的多智能体框架,用于细粒度审计大模型的策略性欺骗。DECOR将输入上下文分解为原子信息单元,针对每个单元在响应中进行四维操控评分,生成可解释的操控画像,并聚合为全局欺骗指数。我们在涵盖真实场景的单轮与多轮欺骗检测基准上进行全面评估,结果表明DECOR在两项任务中均达到当前最优性能,超越多个竞争基线。该框架在15个前沿模型上具备良好泛化能力,消融实验验证了各设计组件的有效性。研究证明,基于理论的细粒度信息操控审计,是实现高效且可解释的大模型欺骗检测的有效路径。
原文摘要 · Abstract (English)
Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such behavior difficult to detect. Existing black-box methods rely on coarse-grained judgments, offering limited interpretability and failing to pinpoint which facts were distorted and how. We introduce DECOR, a multi-agent framework grounded in Information Manipulation Theory for fine-grained auditing of strategic deception in LLM responses. DECOR decomposes input contexts into atomic informational units and scores each unit against the response across four dimensions of manipulation, producing interpretable manipulation profiles that are aggregated into a global deception index. We comprehensively evaluate DECOR on both single-turn and multi-turn deception detection benchmarks spanning real-world domains, and show that DECOR achieves state-of-the-art performance on both, outperforming competitive baselines. The framework generalizes across 15 frontier models, and ablation studies confirm the contribution of each key design component. Our findings demonstrate that fine-grained, theory-grounded auditing of information manipulation offers an effective and interpretable path for LLM deception detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。