为大模型上下文管理建立结构化框架,解决长文本推理中的信息失真问题。
Context Cartography: Toward Structured Governance of Contextual Space in Large Language Model Systems
- 提出上下文三区模型:黑雾(未见)、灰雾(记忆)、可视域(推理)
- 定义七类操作符,实现信息在不同区域间的精准流转与优化
- 揭示行业系统中隐含的共性机制,可指导模型设计与评测
当前提升大模型推理能力的方法主要依赖扩大上下文窗口,隐含假设是更多令牌带来更好性能。然而实证研究显示,基于Transformer架构的上下文空间存在结构梯度、显著性不对称和熵积累等问题。本文提出上下文制图(Context Cartography),一种系统性治理上下文空间的形式化框架。该框架将信息空间划分为三个区域:黑雾(未观测)、灰雾(存储记忆)、可视域(主动推理表面),并形式化定义七种制图算子——侦察、选择、简化、聚合、投影、位移和分层,用于描述信息在区域间及内部的转换。这些算子源自对所有非平凡区域变换的系统性覆盖分析,按变换类型与作用范围分类组织。框架基于Transformer注意力的显著性几何特性,将算子视为对线性前缀记忆、仅追加状态与上下文扩展引发熵累积的必要补偿。对四个主流系统(Claude Code、Letta、MemOS、OpenViking)的分析表明,这些算子在产业中正独立收敛。框架进一步导出可验证预测,包括算子特异性消融假设,并提出诊断基准以供实证检验。
原文摘要 · Abstract (English)
The prevailing approach to improving large language model (LLM) reasoning has centered on expanding context windows, implicitly assuming that more tokens yield better performance. However, empirical evidence - including the "lost in the middle" effect and long-distance relational degradation - demonstrates that contextual space exhibits structural gradients, salience asymmetries, and entropy accumulation under transformer architectures. We introduce Context Cartography, a formal framework for the deliberate governance of contextual space. We define a tripartite zonal model partitioning the informational universe into black fog (unobserved), gray fog (stored memory), and the visible field (active reasoning surface), and formalize seven cartographic operators - reconnaissance, selection, simplification, aggregation, projection, displacement, and layering - as transformations governing information transitions between and within zones. The operators are derived from a systematic coverage analysis of all non-trivial zone transformations and are organized by transformation type (what the operator does) and zone scope (where it applies). We ground the framework in the salience geometry of transformer attention, characterizing cartographic operators as necessary compensations for linear prefix memory, append-only state, and entropy accumulation under expanding context. An analysis of four contemporary systems (Claude Code, Letta, MemOS, and OpenViking) provides interpretive evidence that these operators are converging independently across the industry. We derive testable predictions from the framework - including operator-specific ablation hypotheses - and propose a diagnostic benchmark for empirical validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。