让机器人跨任务复用视觉理解知识,像带背包一样随时调用技能。
Learning reusable concepts across different egocentric video understanding tasks
- 构建统一框架,生成可迁移的任务视角知识包。
- 在多个自视点视频任务中实现知识共享与性能提升。
- 适合需要多任务协同的机器人视觉系统研究者。
人类理解描绘人类活动的视频流时,能在短时间内同时把握当前行为、识别物体间关联与互动,并预测未来事件。为使自主系统具备这种整体感知能力,学习如何关联概念、在不同任务间抽象通用知识,并利用任务间的协同效应来学习新技能至关重要。本文提出 Hier-EgoPack 框架,能够创建一组可跨下游任务迁移的任务视角,作为额外洞察来源,如同机器人携带的技能背包,按需调用。
原文摘要 · Abstract (English)
Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecast what will happen soon, everything all at once. To endow autonomous systems with such holistic perception, learning how to correlate concepts, abstract knowledge across diverse tasks, and leverage tasks synergies when learning novel skills is essential. In this paper, we introduce Hier-EgoPack, a unified framework able to create a collection of task perspectives that can be carried across downstream tasks and used as a potential source of additional insights, as a backpack of skills that a robot can carry around and use when needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。