构建电视遥控交互基准,提升大模型对电视导航拓扑的感知能力。
TVWorld: Foundations for Remote-Control TV Agents
- 基于真实电视操作构建离线图谱,支持可复现评估
- 提出拓扑感知训练框架,使模型在长序列导航中成功率超68%
- 适合研究智能设备控制、视觉语言模型落地的开发者
近期大型视觉-语言模型(LVLMs)在设备控制方面展现出强大潜力,但现有研究主要聚焦于点选点击(PnC)交互,而日常电视使用中常见的遥控器(RC)交互仍缺乏系统探索。为此,本文提出 extbf{TVWorld},一个基于真实电视导航的离线图结构抽象,实现可复现且无需部署的评估。在此基础上,构建两个互补基准: extbf{TVWorld-N}(拓扑感知导航)与 extbf{TVWorld-G}(焦点感知定位),全面评估电视使用能力。实验揭示现有代理在焦点驱动、长时序电视导航中拓扑感知不足的关键缺陷。针对此问题,提出 extit{Topology-Aware Training} 框架,将拓扑知识注入 LVLM。基于该框架,开发专用于电视导航的 extbf{TVTheseus} 模型,在 TVWorld-N 上达到 68.3% 成功率,超越 Gemini 3 Flash 等闭源强基线,达当前最优(SOTA)水平。额外分析为高效电视使用代理的发展提供了关键洞见。
原文摘要 · Abstract (English)
Recent large vision-language models (LVLMs) have demonstrated strong potential for device control. However, existing research has primarily focused on point-and-click (PnC) interaction, while remote-control (RC) interaction commonly encountered in everyday TV usage remains largely underexplored. To fill this gap, we introduce \textbf{TVWorld}, an offline graph-based abstraction of real-world TV navigation that enables reproducible and deployment-free evaluation. On this basis, we derive two complementary benchmarks that comprehensively assess TV-use capabilities: \textbf{TVWorld-N} for topology-aware navigation and \textbf{TVWorld-G} for focus-aware grounding. These benchmarks expose a key limitation of existing agents: insufficient topology awareness for focus-based, long-horizon TV navigation. Motivated by this finding, we propose a \emph{Topology-Aware Training} framework that injects topology awareness into LVLMs. Using this framework, we develop \textbf{TVTheseus}, a foundation model specialized for TV navigation. TVTheseus achieves a success rate of $68.3\%$ on TVWorld-N, surpassing strong closed-source baselines such as Gemini 3 Flash and establishing state-of-the-art (SOTA) performance. Additional analyses further provide valuable insights into the development of effective TV-use agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。