构建多模态众包平台,让机器人学人类行为更高效
TeachAnything: A Multimodal Crowdsourcing Platform for Training Embodied AI Agents in Symmetrical Reality

- 三阶段演示框架融合视觉、语言等多模态信号
- 支持虚拟与物理环境统一交互,收集多样化任务数据
- 适合研究具身智能与人机协同的开发者使用
对称现实(Symmetrical Reality, SR)正成为人机共存的未来趋势,对智能体提出更高要求,亟需更丰富多样的人类指导。本文提出一种三阶段演示范式,整合多模态演示信号。基于此,我们开发了TeachAnything——一个基于云端、面向众包的演示平台,具备物理仿真能力,可在多种场景、任务和形态下收集多样化的演示数据。通过方法设计与物理仿真的统一,系统实现了虚拟与物理交互的无缝衔接,为符合对称现实理念的具身智能体研发提供了切实可行的基础。
原文摘要 · Abstract (English)
Symmetrical Reality (SR) is emerging as a future trend for human-agent coexistence, placing higher demands on agents to acquire human-like intelligence. It calls for richer and more diverse human guidance. We introduce a three-stage demonstration paradigm integrating multimodal demonstration signals. Building on this paradigm, we developed TeachAnything, a cloud-based, crowdsourcing-oriented demonstration platform with physics simulation capable of collecting diverse demonstration data across varied scenes, tasks, and embodiments. By unifying virtual and physical interactions through both methodological design and physics simulation, the system serves as a practical foundation for developing embodied agents aligned with Symmetrical Reality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。