arXiv:2606.25960cs.AI2026-06

用压缩比特数衡量智能,发现智能系统能更高效编码信息。

Agentic System as Compressor: Quantifying System Intelligence in Bits

论文配图:Agentic System as Compressor: Quantifying System Intelligence in Bits
图 1 · 摘自论文原文
  • 以压缩比特数衡量智能:越聪明的系统,越能用更少比特还原信息。
  • 五类实验中,加入智能组件后编码长度均减少,最高降幅达27%。
  • 适合研究智能系统评估、模型优化与资源分配的科研人员参考。

大语言模型正从孤立预测器演变为智能体系统:调用工具、检索证据、遵守环境约束、使用验证器,并通过搜索和多轮交互完成任务。本文基于「压缩即智能」的视角,提出在固定任务分布、接口和算力预算下,更强的智能体系统能以更少比特重构目标对象。通过算术编码、种子编码与备用方案实现该度量,在五种场景(逆向文本、国际象棋走法、蛋白质序列、检索增强问答、语义故事压缩)中验证,所有智能组件均显著降低编码长度。这些小规模可控实验覆盖了真实智能体系统的典型组件,揭示了组件、观察者和算力预算如何影响残余不确定性,为评估真实智能体系统提供指导。

原文摘要 · Abstract (English)

Large language models are turning from isolated predictors into agentic systems: they call tools, retrieve evidence, obey environment constraints, use verifiers, and complete tasks through search and multi-turn interaction. We adopts an analytical viewpoint based on "compression is intelligence": under a fixed task distribution, interface, and compute budget, a stronger agentic system lets a target object be reconstructed with fewer bits. We operationalize the measure with arithmetic coding, seed coding, and a fallback, and evaluate it in five settings: reversed text, chess moves, protein sequences, retrieval-augmented question answering, and semantic story compression; in all of them agentic components reduce codelength. These small, controlled experiments cover component types typical of real agentic systems, show that codelength can analyze how components, observers, and budgets change residual uncertainty, and offer guidance for evaluating real agent systems.

智能体压缩评估语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。