用通信理论统一分析大模型可靠性技术,实现质量与成本的动态平衡。
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability

- 将大模型采样视为信道,用通信理论构建六类可靠性操作框架。
- 在69个难题上验证,动态路由使成本降56%、质量提升7%。
- 无需重训练即可调整质量-成本权衡,适合部署优化场景。
基于大语言模型(LLMs)的智能体依赖多种可靠性技术,如重试、多数投票和自一致,但这些方法缺乏统一分析框架。本文观察到,在温度T下采样的LLM可视为香农编码理论中的离散随机信道p(y|x),以此为切入点构建通信理论基础的分析框架。所提出的框架涵盖六类经典可靠性操作:多样性合并、混合重传、迭代生成-批判解码、无速率采样、结构化冗余验证及难度自适应路由。框架内给出两个闭式结果:噪声方差阈值以上时,均匀平均优于质量加权平均;生成-批判精炼具备收缩性准则,与3B至14B参数模型间观察到的从收缩到发散的转变一致。进一步提出一种成本感知语义最近邻路由器,仅通过一个拉格朗日参数即可遍历质量-成本前沿而无需重新训练。在69个跨本地与云部署的硬任务上,无固定组合占优,支持按任务分配资源。在包含MMLU、GSM8K和HumanEval的300项难题子集上,该路由器达到全经验帕累托前沿:在同等质量下,其归一化成本比最强固定技术低约56%;在相同归一化成本下,质量提升约7%(较单次采样提升26%)。结果表明,应将这些可靠性技术整合为一个由信道编码指导的可调层。
原文摘要 · Abstract (English)
Agents built on large language models (LLMs) rely on a range of reliability techniques, including retry, majority voting, and self-consistency, that have been developed in parallel rather than within a common analytical framework. We observe that an LLM sampled at temperature $T$ is a discrete stochastic channel $p(y \mid x)$ in the sense of Shannon's coding theory, and use this identity as the entry point for such a framework grounded in communication theory. Each of these techniques is a special case of one of six classical reliability operators: diversity combining, hybrid retransmission, iterative generator-critic decoding, rateless sampling, structured redundant verification, and difficulty-adaptive routing. Within the framework we give two closed-form results: a noise-variance threshold above which uniform averaging beats quality-weighted averaging, and a contractivity criterion for generator-critic refinement, consistent with a contractive-to-divergent transition we observe between 3B- and 14B-parameter models. We further introduce a cost-aware semantic-nearest-neighbor router whose single Lagrangian knob traverses the quality-cost frontier without retraining. Across six channel configurations spanning local and cloud models on 69 hard tasks, no fixed model-technique-budget choice dominates, motivating per-task allocation. On a 300-item hard split of MMLU, GSM8K, and HumanEval, our router occupies the full empirical Pareto frontier: at matched quality, its normalized cost is ${\approx}56$\% lower than the strongest fixed technique; at matched normalized cost, it improves quality by ${\approx}7$\% ($26$\% over single-shot decoding). These results argue for consolidating these reliability techniques into a single tunable layer informed by channel coding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。