用斐波那契采样提升视觉语言动作模型的时序效率与动作平滑性
FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling

- 采用对数回溯采样减少冗余,高效捕捉长时序依赖
- 斐波那契循环推理使动作规划更连贯,成功率达92.3%
- 无需重训大模型,实时响应优于视频基线
视觉-语言-动作模型(VLAs)通过多模态认知推断物理世界行为,为具身智能提供通用解决方案。传统VLAs多聚焦于当前状态认知,虽有尝试引入时序信息以增强推理能力,但长上下文编码导致效率下降。为此,本文提出FibVLA,一种兼具长时序感知与高效率的框架。具体地,通过对数回溯采样对本体状态与视觉帧进行处理,以最小冗余捕捉长期时序依赖;在动作专家部分,引入流匹配生成动作分布,并采用斐波那契递归推理策略,基于实时闭环反馈生成长程规划步骤。实验表明,FibVLA在不重训练大规模视觉编码器的前提下,显著提升动作平滑性与成功率(达92.3%),效率分析显示其在真实场景中具有更优的实时响应能力,优于视频基线。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes an efficiency-decreasing issue. To reconcile the conflict between capturing temporal information and maintaining inference efficiency in VLAs, this paper introduces FibVLA, an efficient framework featuring temporal perception of long-context history. Specifically, we leverage logarithmic hindsight sampling to both proprioceptive states and visual frames to capture long-term temporal dependencies with minimal redundancy. For the action expert, we introduce the flow matching to produce action distributions, and the Fibonacci recurrent inference strategy to generate long-range planning steps based on real-time closed-loop feedback. Experiments demonstrate that FibVLA significantly improves action smoothness and success rates without retraining large-scale visual encoders. Efficiency analysis demonstrates superior real-time responsiveness compared to video-based baselines in real-world evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。