首个面向模仿学习中主动视觉与前瞻注视的基准测试平台。
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

- 构建两类任务套件,分别模拟头部与手部摄像头的主动视觉
- 发现主动视觉效果因任务而异,且跨任务泛化能力弱
- 首次量化学习策略的前瞻注视行为,媲美人类操作员表现
主动视觉——即策略自主控制观察视角进行操作——已成为模仿学习的关键能力。然而,缺乏统一的基准来比较不同方法、量化其在各类任务和条件下的贡献。本文提出TAVIS,一个用于主动视觉模仿学习的评估框架,包含两个互补的任务套件:TAVIS-Head(5项任务,通过云台移动实现全局搜索)和TAVIS-Hands(3项任务,通过腕部摄像头处理局部遮挡),基于IsaacLab在两种人形躯干机器人(GR1T2、Reachy2)上实现。TAVIS提供三种评估范式:头戴相机与固定相机的配对实验;GALT(注视-动作领先时间),一种基于认知科学与人机交互的新型指标,量化策略的前瞻注视行为;以及程序性数据集内/外分布划分。基线实验使用Diffusion Policy和$π_0$表明:(i) 主动视觉总体有益,但效果具有任务依赖性而非普适;(ii) 多任务策略在受控分布偏移下性能急剧下降;(iii) 仅靠模仿即可生成前瞻注视,中位领先时间与人类遥控操作员相当。代码、评估脚本、演示数据(LeRobot v3.0;约2200个轨迹)及训练好的基线模型已开源于https://github.com/spiglerg/tavis 和 https://huggingface.co/tavis-benchmark。
原文摘要 · Abstract (English)
Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared benchmark to compare approaches or quantify what active vision contributes, on which task types, and under what conditions. We introduce TAVIS, evaluation infrastructure for active-vision imitation learning, with two complementary task suites -- TAVIS-Head (5 tasks, global search via pan/tilt necks) and TAVIS-Hands (3 tasks, local occlusion via wrist cameras) -- on two humanoid torso embodiments (GR1T2, Reachy2), built on IsaacLab. TAVIS provides three evaluation primitives: a paired headcam-vs-fixedcam protocol on identical demonstrations; GALT (Gaze-Action Lead Time), a novel metric grounded in cognitive science and HRI that quantifies anticipatory gaze in learned policies; and procedural ID/OOD splits. Baseline experiments with Diffusion Policy and $π_0$ reveal that (i) active-vision generally helps, but benefits are task-conditional rather than uniform; (ii) multi-task policies degrade sharply under controlled distribution shifts on both suites; and (iii) imitation alone yields anticipatory gaze, with median lead times comparable to the human teleoperator reference. Code, evaluation scripts, demonstrations (LeRobot v3.0; ~2200 episodes) and trained baselines are released at https://github.com/spiglerg/tavis and https://huggingface.co/tavis-benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。