arXiv:2511.05397cs.ROcs.CV2025-11被引 4

低成本机器人搭配统一视觉语言动作模型,实现在复杂场景下的可靠操作。

EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation

  • 用统一模型同时输出离散与连续动作,实现端到端控制。
  • 真实场景中性能比之前方法提升49%(分布内)和34.9%(分布外)。
  • 整机成本低于300美元,适合家庭与实验室普及使用。

虽然视觉-语言-动作(VLA)模型可将视觉输入与语言指令直接映射为机器人动作,但通常依赖昂贵硬件,在新环境或杂乱场景中表现不佳。我们提出EverydayVLA,一种6自由度机械臂,组装成本低于300美元,具备一定承载能力和工作空间。采用单一统一模型联合输出离散与连续动作,其自适应时域集成机制通过监测运动不确定性,实现运行中动态重规划,确保安全可靠。在LIBERO数据集上,EverydayVLA达到顶尖成功率;真实世界测试中,其在分布内任务上优于先前方法49%,分布外任务上提升34.9%。结合先进VLA模型与低成本硬件,EverydayVLA推动了机器人基础模型的普及,为家庭与科研实验室提供经济可行方案。实验视频与详情见:https://everydayvla.github.io/

原文摘要 · Abstract (English)

While Vision-Language-Action (VLA) models map visual inputs and language instructions directly to robot actions, they often rely on costly hardware and struggle in novel or cluttered scenes. We introduce EverydayVLA, a 6-DOF manipulator that can be assembled for under $300, capable of modest payloads and workspace. A single unified model jointly outputs discrete and continuous actions, and our adaptive-horizon ensemble monitors motion uncertainty to trigger on-the-fly re-planning for safe, reliable operation. On LIBERO, EverydayVLA matches state-of-the-art success rates, and in real-world tests it outperforms prior methods by 49% in-distribution and 34.9% out-of-distribution. By combining a state-of-the-art VLA with cost-effective hardware, EverydayVLA democratizes access to a robotic foundation model and paves the way for economical use in homes and research labs alike. Experiment videos and details: https://everydayvla.github.io/

机器人低成本视觉语言动作端到端控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。