研究大模型跨会话任务延续的最优信息传递方法
Handover of In-Context Learning State Across Session Boundaries
- 将上下文学习状态作为可传递的中间态,明确传递目标
- 证明了在特定条件下只需固定位数即可保证预测等价性
- 提出三部分记录机制,兼顾精度与记忆成本
本研究探讨大型语言模型应用中会话间状态转移的方法论与理论特性。当上下文超出模型输入限制、应用重启或由其他代理接续任务时,需决定从前一会话传递哪些信息。本文将手交定义为任务相关的上下文学习(ICL)状态转移,区分精确恢复先前内容与保留目标分布。在外生性条件下,预测等价性刻画了最粗粒度的确定性充分手交,并给出固定长度比特需求。分析分离了记忆约束、写入者和延续过程的影响,量化了在知晓下游查询前书写带来的代价。提出三部分记录:精确存储决策与约束,用任务合理的统计量处理重复证据,保留未被统计量覆盖的原始观测。高斯线性回归给出精确有限维手交及有限比特扰动界;非参数回归提供平方预测误差与内存的上下界。结果为确定手交必须保留的内容及其内存依赖提供了理论与方法。
原文摘要 · Abstract (English)
This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. The application must then decide which information from the earlier session to pass on. We formulate handover as the transfer of a task-relative in-context learning (ICL) state and distinguish exact recovery of earlier material from preservation of the target distribution. Under an exogeneity condition, predictive equivalence characterizes the coarsest deterministic sufficient handover and gives a fixed-length bit requirement. The analysis isolates the effects of the memory constraint, the writer, and the continuation procedure, and quantifies the cost of writing before the realized downstream query is known. We propose a three-part record that stores decisions and constraints exactly, uses task-justified statistics for repeated evidence, and retains original observations whose effect is not preserved by those statistics. Gaussian linear regression gives an exact finite-dimensional handover and finite-bit perturbation bounds, while nonparametric regression gives upper and lower bounds that relate memory to squared prediction error. These results provide a theory and method for deciding what a handover must retain and how its memory requirement depends on the continuation task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。