RynnEC让机器人能像人一样看懂视频中物体的局部属性和空间关系。
RynnEC: Bringing MLLMs into Embodied World
- 用区域编码器+掩码解码器实现视频中局部区域的精细理解
- 在物体属性、分割和空间推理上达到当前最优性能
- 适合做具身智能体认知核心,尤其适用于数据稀缺场景
我们提出RynnEC,一种面向具身认知的视频多模态大模型。基于通用视觉-语言基础模型,RynnEC引入区域编码器与掩码解码器,支持灵活的区域级视频交互。尽管架构紧凑,其在物体属性理解、物体分割和空间推理任务上均达到领先水平。该模型提供以区域为中心的视频理解范式,使具身智能体能获得对物理世界的细粒度感知,从而实现更精准的交互。为缓解3D标注数据稀缺问题,我们设计了一种基于第一人称视角视频的数据生成流水线。此外,我们构建了RynnEC-Bench——一个以区域为中心的具身认知评估基准。我们期望RynnEC能推动通用认知核心的发展,并促进在多样化具身任务间的泛化能力。代码、模型权重及基准已开源:https://github.com/alibaba-damo-academy/RynnEC
原文摘要 · Abstract (English)
We introduce RynnEC, a video multimodal large language model designed for embodied cognition. Built upon a general-purpose vision-language foundation model, RynnEC incorporates a region encoder and a mask decoder, enabling flexible region-level video interaction. Despite its compact architecture, RynnEC achieves state-of-the-art performance in object property understanding, object segmentation, and spatial reasoning. Conceptually, it offers a region-centric video paradigm for the brain of embodied agents, providing fine-grained perception of the physical world and enabling more precise interactions. To mitigate the scarcity of annotated 3D datasets, we propose an egocentric video based pipeline for generating embodied cognition data. Furthermore, we introduce RynnEC-Bench, a region-centered benchmark for evaluating embodied cognitive capabilities. We anticipate that RynnEC will advance the development of general-purpose cognitive cores for embodied agents and facilitate generalization across diverse embodied tasks. The code, model checkpoints, and benchmark are available at: https://github.com/alibaba-damo-academy/RynnEC
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。