arXiv:2607.29401cs.CV2026-07

提出一步证据融合方法,提升跨视频场景规划的准确率。

OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

论文配图:OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning
图 1 · 摘自论文原文
  • 构建多源基准,统一评估证据检索与计划生成
  • 在六组可比任务中取得最高性能,胜过现有最强方法2.9-10.7点
  • 适合需要精准视频证据融合的智能助手研发者

视频场景流程规划(VSPP)预先给出起止观测,但未解决证据需主动检索的问题。本文提出跨视频场景流程规划(CVSPP):给定一个答案遮蔽的起止查询和K个候选视频,模型需检索支持视频、定位相关片段并预测动作序列。两大挑战耦合:同类任务演示共享阶段与窗口,早期硬选择会传递错误场景链。我们构建了一个包含11个数据源的基准,包含带类型负样本、防答案泄露门控机制及分离的证据与计划评估指标。在14个源时序单元上,我们对比九种规划器家族与多数序列基线。提出一步证据融合(OSEF),在所有候选视频上对查询条件下的单元-跨度网格进行评分,并通过令牌全局适配器将完整软网格输入规划器,不提前裁剪窗口。OSEF在六个经认证可比较的单元中排名第一。在四组匹配的同任务与跨任务COIN和CrossTask单元中,相比增强版硬选择最先进方法,精确视频与计划成功率提升2.9–10.7个百分点;组件分析表明,令牌全局接口贡献最大单次提升。五组转换源单元接近多数序列基线,体现基准剩余潜力。附录提供模型构造器与评估代码。

原文摘要 · Abstract (English)

Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence must itself be retrieved. We introduce Cross-Video Scene Procedure Planning (CVSPP): given an answer-redacted start-goal query and K candidate videos, a model must retrieve the supporting video, localize the relevant window, and predict the action sequence. Two obstacles couple here. Same-task demonstrations share stages and windows, and an early hard selection passes the wrong scene chain to the planner. We build an eleven-source benchmark with typed negative roles, a fail-closed answer-leakage gate, and separate Evidence- and Plan-axis metrics. On its 14 source-horizon cells we adapt nine planner families against a majority-sequence floor. We then present One-Step Evidence Fusion (OSEF), which scores a query-conditioned cell-and-span lattice over all candidates and feeds the full soft lattice to the planner through a token-global adapter, cropping no window beforehand. OSEF ranks first on all six cells the benchmark certifies as method-rankable. On four matched same-task COIN and CrossTask cells it improves exact-video-and-plan success by 2.9-10.7 points over an enhanced hard-selection SOTA, and a component study assigns the largest single increment to the token-global interface. Five converted-source cells sit at or near the majority-sequence floor, the benchmark's remaining headroom. The supplementary package includes model constructors and evaluation code.

视频理解规划推理证据融合多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。