用大模型压缩上下文,让智能体更高效决策
Learning to Decide with Just Enough: Information-Theoretic Context Summarization for CMDPs
- 用大模型将复杂上下文压缩为低维语义摘要
- 在多个任务中提升奖励与采样效率,降低延迟和内存
- 首次给出CMDP的后悔值上界与延迟-熵权衡分析
上下文马尔可夫决策过程(CMDPs)为外部信号下的序列决策提供了框架,但现有方法在高维或非结构化上下文中泛化能力差,导致计算量过大且性能不稳定。本文提出一种基于信息论的上下文摘要方法,利用大语言模型(LLMs)将上下文输入压缩为低维、语义丰富的摘要,增强状态表示,保留关键决策线索的同时减少冗余。基于近似上下文充分性概念,我们首次为CMDPs提供了后悔值上界及延迟-熵权衡特性分析,阐明了信息量对计算成本的影响。在离散、连续、视觉和推荐等多类基准测试中,该方法显著优于原始上下文和非上下文基线,提升了奖励、成功率和样本效率,同时降低了延迟与内存占用。结果表明,基于LLM的摘要为资源受限的高上下文环境中高效决策提供了可扩展且可解释的解决方案。
原文摘要 · Abstract (English)
Contextual Markov Decision Processes (CMDPs) offer a framework for sequential decision-making under external signals, but existing methods often fail to generalize in high-dimensional or unstructured contexts, resulting in excessive computation and unstable performance. We propose an information-theoretic summarization approach that uses large language models (LLMs) to compress contextual inputs into low-dimensional, semantically rich summaries. These summaries augment states by preserving decision-critical cues while reducing redundancy. Building on the notion of approximate context sufficiency, we provide, to our knowledge, the first regret bounds and a latency-entropy trade-off characterization for CMDPs. Our analysis clarifies how informativeness impacts computational cost. Experiments across discrete, continuous, visual, and recommendation benchmarks show that our method outperforms raw-context and non-context baselines, improving reward, success rate, and sample efficiency, while reducing latency and memory usage. These findings demonstrate that LLM-based summarization offers a scalable and interpretable solution for efficient decision-making in context-rich, resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。