arXiv:2512.03666cs.CVcs.AI2025-12被引 6

首个面向任务的视角视频定位基准,推动智能体理解动作意图。

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos

  • 基于任务目标而非描述定位物体,模拟真实交互场景。
  • 支持显式与隐式双重定位,100段视频含2704条指令。
  • 评测多目标、隐式推理能力,适配机器人与具身智能研究者。

实现通用具身智能的核心能力在于从第一人称视角定位与任务相关的物体,即时空视频定位(STVG)。尽管近期取得进展,现有研究仍主要局限于以物体为中心的描述性指令,忽视了具身智能体完成目标导向交互所必需的任务推理。为此,我们提出首个面向第一人称视频的任务导向时空定位基准——ToG-Bench。该基准具备三大特性:(1) 任务导向定位,要求根据目标任务而非直接描述识别并定位物体;(2) 显式-隐式双重定位,目标物体可明确提及或通过上下文推理得出;(3) 一对一多定位,一条指令可能对应多个参与任务执行的物体。ToG-Bench基于ScanNet视频构建,包含100个标注片段与2,704条任务导向定位指令,采用半自动流程结合基础模型标注与人工精修生成。同时,我们设计了一套面向多物体与显-隐式定位的任务级评估指标,并系统评测了七种先进多模态大模型。实验揭示任务导向STVG的内在挑战及在显-隐式与多对象定位上的显著性能差距,凸显具身场景中感知与交互融合的难度。数据与代码将公开于:https://github.com/qaxuDev/ToG-Bench。

原文摘要 · Abstract (English)

A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG). Despite recent progress, existing STVG studies remain largely confined to object-centric and descriptive instructions, neglecting the task-oriented reasoning that is crucial for embodied agents to accomplish goal-directed interactions. To bridge this gap, we introduce \textbf{ToG-Bench}, the first task-oriented spatio-temporal video grounding benchmark for egocentric videos. ToG-Bench is characterized by three key features: (1) \textbf{Task-oriented Grounding}, which requires identifying and localizing objects based on intended tasks rather than straightforward descriptions; (2) \textbf{Explicit-Implicit Dual Grounding}, where target objects can be either explicitly mentioned or implicitly inferred by contextual reasoning; (3) \textbf{One-to-Many Grounding}, where a single instruction may correspond to multiple objects involved in task execution. Built upon videos sourced from ScanNet, ToG-Bench comprises 100 annotated clips with 2,704 task-oriented grounding instructions, constructed via a semi-automated pipeline that combines foundation model annotation and human refinement. In addition, we introduce a set of task-level evaluation metrics tailored for multi-object and explicit-implicit object grounding, and systematically benchmark seven state-of-the-art MLLMs. Extensive experiments reveal the intrinsic challenges of task-oriented STVG and substantial performance gaps across explicit-implicit and multi-object grounding, highlighting the difficulty of bridging perception and interaction in embodied scenarios. Data and code will be released at: \href{https://github.com/qaxuDev/ToG-Bench}{https://github.com/qaxuDev/ToG-Bench}..

视频定位具身智能任务导向多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。