arXiv:2512.00783cs.LGcs.RO2025-12被引 1

让视觉语言动作模型实现意念式对齐,无需重训练即可精准控制。

Sigma: The Key for Vision-Language-Action Models toward Telepathic Alignment

  • 构建新架构Sigma,融合语义理解与关联推理,打通感知与动作的思维桥梁。
  • 在相同数据下控制误差下降,向量、片段和轨迹级表现均更优。
  • 适合做意图驱动机器人控制的研究者,可直接复现且不改动基础模型。

为解决认知系统中语义与连续控制间缺乏可时序更新的中介思维空间这一根本问题,本文构建并训练了名为Sigma的视觉-语言-动作模型,部署于单张RTX 4090显卡。模型基于开源pi0.5_base骨干网络,将svla_so101_pickplace数据集预处理为结构化训练语料。提出独立设计的VLA架构,融合深度语义理解与关联推理,实现感知与动作间的意念式对齐。训练通过迭代优化数据预处理、基于LoRA的微调及推理阶段适配器设计完成。评估采用离线闭环回放,在相同数据条件下对比未调参的pi0.5_base。实验结果表明,Σigma在向量、片段和轨迹层级均持续降低控制均方误差,同时保持意念范数与语义文本对齐质量的稳定。这些发现证明,仅通过语义与关联架构集成即可定量实现心响应式对齐控制,无需重训练基础模型,为语义对齐与意图驱动行为提供可复现路径。

原文摘要 · Abstract (English)

To address a fundamental limitation in cognitive systems, namely the absence of a time-updatable mediating thought space between semantics and continuous control, this work constructs and trains a vision-language-action model termed Sigma, deployed on a single RTX 4090. The model is built upon the open-source pi0.5_base backbone, with the svla_so101_pickplace dataset preprocessed into a structured training corpus. An independently designed VLA architecture is introduced to integrate deep semantic understanding with associative reasoning, enabling telepathic-style alignment between perception and action. Training proceeds through iterative optimization of data preprocessing, LoRA-based fine-tuning, and inference-stage adapter design. Evaluation is conducted using offline closed-loop replay, comparing Sigma against the untuned pi0.5_base under identical data conditions. Experimental results indicate a consistent reduction in control MSE across vector-, fragment-, and trajectory-level scales, while preserving the stability of the telepathy norm and semantic-text alignment quality. These findings demonstrate that mind-responsive alignment control can be quantitatively achieved through semantic and associative architectural integration without retraining the base model, providing a reproducible pathway for semantic alignment and intention-driven behavior.

视觉语言动作意念对齐控制优化低资源训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。