在固定预算下,动态调整访问策略比扩大模型规模更有效提升感知精度。
Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B
- 用四类资源墙定义决策约束,设计无需训练的ASP框架动态分配预算。
- 4096令牌预算下,检索准确率达75%-94%,是固定采样的3-19倍。
- 适合关注高效推理、预算敏感场景的多模态系统研究者。
具身多模态智能体需在固定每步令牌预算下处理持续的观察流。我们通过四类资源墙形式化该约束:感知香农墙(状态受限)、时域墙(查询无关帧选择)、轮次墙(非自适应检索)和条件组合墙(固定深度推理)。提出ASP——一种针对冻结多模态模型的无训练封装,结合截断结构化状态、完整回忆索引与查询条件预算分配,支持迭代访问。遵循预注册协议,在SEW-Bench(一个免授权的合成长时程走读基准)上评估了7个3B至31B参数的开源模型。因自然视频需数据集授权未运行,证据聚焦访问机制而非真实场景感知。在4,096令牌决策预算下,ASP实现75%至94%的回忆准确率,远超等预算下查询无关采样的3%至19%;预算重分配在所有主干模型上优于采样预算翻倍。但完整三组件架构未验证通道对偶性:移除压缩状态使旗舰平均分从35.4升至58.0,ASP在任一主干上均未超越仅使用原始索引的基线,且四个预注册可证伪标准中有两个被触发。结果表明,在固定预算下,查询条件访问机制较参数量或上下文扩展更具决定性,而提示式在线压缩在此设置中未能证明其成本合理性。
原文摘要 · Abstract (English)
Embodied multimodal agents must answer from growing observation streams under a fixed per-decision token budget. We formalize this constraint through four resource walls: a perceptual Shannon wall for bounded state, a horizon wall for query-independent frame selection, a round wall for non-adaptive retrieval, and a conditional composition wall for fixed-depth inference. We introduce ASP, a training-free wrapper for frozen multimodal models that combines a capped structured state, a verbatim episodic index, and query-conditioned budget allocation with iterative access. Following a pre-registered protocol, we evaluate seven open-weight models from 3B to 31B on SEW-Bench, a license-free synthetic long-horizon walkthrough benchmark constructed to instantiate these walls. The registered natural-video benchmarks were not run because their frames require dataset agreements; our evidence therefore concerns access mechanisms, not natural-scene perception. Under a 4,096-token decision budget, ASP reaches 75 to 94% episodic retrieval accuracy, compared with 3 to 19% for equal-budget query-independent sampling, and budget reallocation outperforms quadrupling the sampling budget on every backbone. However, the full three-component architecture does not validate channel duality: removing the compressive state raises the flagship mean from 35.4 to 58.0, ASP does not outperform the verbatim-only baseline on any backbone, and two of four pre-registered falsification criteria fire. These results show that query-conditioned access, rather than parameter count or context growth alone, is decisive under a fixed budget, while prompted online compression does not earn its cost in this setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。