用低秩流形优化压缩推理状态,提升受限记忆下的决策能力
Rate-Distortion Analysis of Compressed Query Delegation with Low-Rank Riemannian Updates
- 将高维推理状态压缩为低秩张量查询,通过流形优化更新
- 在2500项任务中实现与链式思维相当的准确率,且内存占用更低
- 适合研究大模型认知边界与高效推理机制的学者
受限上下文智能体在中间推理超出有效工作记忆预算时会失效。本文研究压缩查询委托(CQD):(i) 将高维潜在推理状态压缩为低秩张量查询,(ii) 将最小查询委托给外部代理,(iii) 通过固定秩流形上的黎曼优化更新潜在状态。提出数学优先的建模:CQD 是一个带查询预算泛函的约束随机规划问题,代理被建模为有噪算子。将 CQD 与经典率失真和信息瓶颈原理关联,证明谱硬阈值对自然的约束二次失真问题是最优的,并在有界代理噪声和光滑性假设下推导出黎曼随机近似的收敛保证。实验上报告:(A) 一个包含2500个项目的受限上下文推理套件(基于BBH任务及精心筛选的悖论实例),在固定计算与上下文条件下对比CQD与链式思维基线;(B) 一项人类“认知镜像”基准测试(N=200),测量现代代理的知性增益与语义漂移。
原文摘要 · Abstract (English)
Bounded-context agents fail when intermediate reasoning exceeds an effective working-memory budget. We study compressed query delegation (CQD): (i) compress a high-dimensional latent reasoning state into a low-rank tensor query, (ii) delegate the minimal query to an external oracle, and (iii) update the latent state via Riemannian optimization on fixed-rank manifolds. We give a math-first formulation: CQD is a constrained stochastic program with a query-budget functional and an oracle modeled as a noisy operator. We connect CQD to classical rate-distortion and information bottleneck principles, showing that spectral hard-thresholding is optimal for a natural constrained quadratic distortion problem, and we derive convergence guarantees for Riemannian stochastic approximation under bounded oracle noise and smoothness assumptions. Empirically, we report (A) a 2,500-item bounded-context reasoning suite (BBH-derived tasks plus curated paradox instances) comparing CQD against chain-of-thought baselines under fixed compute and context; and (B) a human "cognitive mirror" benchmark (N=200) measuring epistemic gain and semantic drift across modern oracles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。