通过信息论分析规则编码如何影响大模型注意力与合规性。
Rule Encoding and Compliance in Large Language Models: An Information-Theoretic Analysis
- 用信息论方法研究规则格式对注意力机制的影响。
- 低熵规则格式可提升指针精确度,但存在冗余与熵的权衡。
- 提出动态验证架构,提升合规输出的长期概率。
基于大语言模型的安全关键型代理设计不能仅依赖简单的提示工程。本文对系统提示中的规则编码如何影响注意力机制和合规行为进行了全面的信息论分析。我们证明,语法熵较低且锚点高度集中的规则格式能降低注意力熵并提升指针精确度,但揭示了先前研究未注意到的锚点冗余与注意力熵之间的根本性权衡。通过对因果、双向、局部稀疏、核化及交叉注意力等多种注意力架构的正式分析,我们建立了指针精确度的边界,并说明锚点布局策略必须兼顾精确度与熵的双重目标。结合动态规则验证架构,我们给出了一个形式化证明:经过验证的规则集热重载可提高合规输出的渐近概率。这些发现强调了原则性锚点设计与双重执行机制的重要性,以保护基于大模型的代理免受提示注入攻击,并在动态环境中维持合规性。
原文摘要 · Abstract (English)
The design of safety-critical agents based on large language models (LLMs) requires more than simple prompt engineering. This paper presents a comprehensive information-theoretic analysis of how rule encodings in system prompts influence attention mechanisms and compliance behaviour. We demonstrate that rule formats with low syntactic entropy and highly concentrated anchors reduce attention entropy and improve pointer fidelity, but reveal a fundamental trade-off between anchor redundancy and attention entropy that previous work failed to recognize. Through formal analysis of multiple attention architectures including causal, bidirectional, local sparse, kernelized, and cross-attention mechanisms, we establish bounds on pointer fidelity and show how anchor placement strategies must account for competing fidelity and entropy objectives. Combining these insights with a dynamic rule verification architecture, we provide a formal proof that hot reloading of verified rule sets increases the asymptotic probability of compliant outputs. These findings underscore the necessity of principled anchor design and dual enforcement mechanisms to protect LLM-based agents against prompt injection attacks while maintaining compliance in evolving domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。