arXiv:2601.14945cs.ROcs.AI2026-01被引 3

让视觉语言动作模型在边缘设备上实现9赫兹高频控制,突破延迟瓶颈。

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control

  • 分层双频架构:低频推理缓存语义,高频执行实时调整
  • 边缘端达9赫兹控制频率(基线2.4赫兹),反馈频率提升4倍
  • 适合动态环境中的实时控制任务,如移动目标拦截

大规模视觉-语言-动作(VLA)模型具备语义泛化能力,但推理延迟高,仅适用于低频批量执行模式。这种频率不匹配导致执行盲区,在目标运动的动态环境中易失败。我们提出TIDAL(时间交织扩散与动作循环),一种解耦语义推理与高频执行的分层框架。TIDAL作为基于扩散的VLA的通用模块,采用双频架构重新分配计算资源:低频宏观意图循环缓存语义嵌入,高频微观控制循环交错执行单步流整合与动作。该设计使边缘设备实现约9赫兹控制更新(相比基线约2.4赫兹),且边际开销无增加。为应对延迟变化,引入时间错位训练策略,使策略学习利用过时语义意图和实时本体感知进行预测补偿。此外,通过引入微分运动预测器,缓解静态视觉编码器对速度不敏感的问题。实验表明,在动态拦截任务中,性能较开环基线提升2倍。尽管静态成功率略有下降,但反馈频率提升4倍,语义嵌入有效作用时长超过原始动作块大小。在非暂停推理协议下,TIDAL仍保持鲁棒性,而标准基线因延迟失效。

原文摘要 · Abstract (English)

Large-scale Vision-Language-Action (VLA) models offer semantic generalization but suffer from high inference latency, limiting them to low-frequency batch-and-execute paradigm. This frequency mismatch creates an execution blind spot, causing failures in dynamic environments where targets move during the open-loop execution window. We propose TIDAL (Temporally Interleaved Diffusion and Action Loop), a hierarchical framework that decouples semantic reasoning from high-frequency actuation. TIDAL operates as a backbone-agnostic module for diffusion-based VLAs, using a dual-frequency architecture to redistribute the computational budget. Specifically, a low-frequency macro-intent loop caches semantic embeddings, while a high-frequency micro-control loop interleaves single-step flow integration with execution. This design enables approximately 9 Hz control updates on edge hardware (vs. approximately 2.4 Hz baselines) without increasing marginal overhead. To handle the resulting latency shift, we introduce a temporally misaligned training strategy where the policy learns predictive compensation using stale semantic intent alongside real-time proprioception. Additionally, we address the insensitivity of static vision encoders to velocity by incorporating a differential motion predictor. TIDAL is architectural, making it orthogonal to system-level optimizations. Experiments show a 2x performance gain over open-loop baselines in dynamic interception tasks. Despite a marginal regression in static success rates, our approach yields a 4x increase in feedback frequency and extends the effective horizon of semantic embeddings beyond the native action chunk size. Under non-paused inference protocols, TIDAL remains robust where standard baselines fail due to latency.

VLA控制扩散模型高频执行边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。