让视觉语言动作模型更懂意图与环境,提升复杂场景决策效率
MAIN-VLA: Modeling Abstraction of Intention and eNvironment for Vision-Language-Action Models
- 分离意图与环境语义,用抽象表征替代原始感知输入
- 在Minecraft和PvP游戏上实现更高决策质量与推理效率
- 无需调参的注意力过滤机制,自动去除冗余信息
尽管视觉-语言-动作(VLA)模型取得显著进展,但在涉及实时不可预测交互的复杂动态环境(如3D开放世界和大规模对战游戏)中,现有方法仍难以从冗余传感器流中有效提取关键行动信号。为此,我们提出MAIN-VLA框架,通过显式建模意图与环境语义抽象,使决策基于深层语义对齐而非表面模式匹配。具体而言,意图抽象(IA)将冗长语言指令及其推理转化为紧凑明确的语义原语;环境语义抽象(ESA)将海量视觉流投影为结构化、拓扑化的可操作性表示。进一步地,两种抽象模态的对齐引发涌现的注意力集中效应,实现无需参数调整的令牌剪枝策略,可在不降低性能的前提下过滤感知冗余。在开放世界Minecraft及大规模对战环境(Game for Peace、Valorant)中的大量实验表明,MAIN-VLA达到新基准,兼具更优决策质量、更强泛化能力与前沿推理效率。
原文摘要 · Abstract (English)
Despite significant progress in Visual-Language-Action (VLA), in highly complex and dynamic environments that involve real-time unpredictable interactions (such as 3D open worlds and large-scale PvP games), existing approaches remain inefficient at extracting action-critical signals from redundant sensor streams. To tackle this, we introduce MAIN-VLA, a framework that explicitly Models the Abstraction of Intention and eNvironment to ground decision-making in deep semantic alignment rather than superficial pattern matching. Specifically, our Intention Abstraction (IA) extracts verbose linguistic instructions and their associated reasoning into compact, explicit semantic primitives, while the Environment Semantics Abstraction (ESA) projects overwhelming visual streams into a structured, topological affordance representation. Furthermore, aligning these two abstract modalities induces an emergent attention-concentration effect, enabling a parameter-free token-pruning strategy that filters out perceptual redundancy without degrading performance. Extensive experiments in open-world Minecraft and large-scale PvP environments (Game for Peace and Valorant) demonstrate that MAIN-VLA sets a new state-of-the-art, which achieves superior decision quality, stronger generalization, and cutting-edge inference efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。