提出一步证据融合方法,提升跨视频场景规划的准确率。
OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

- 构建多源基准,统一评估证据检索与计划生成
- 在六组可比任务中取得最高性能,胜过现有最强方法2.9-10.7点
- 适合需要精准视频证据融合的智能助手研发者
视频场景流程规划(VSPP)预先给出起止观测,但未解决证据需主动检索的问题。本文提出跨视频场景流程规划(CVSPP):给定一个答案遮蔽的起止查询和K个候选视频,模型需检索支持视频、定位相关片段并预测动作序列。两大挑战耦合:同类任务演示共享阶段与窗口,早期硬选择会传递错误场景链。我们构建了一个包含11个数据源的基准,包含带类型负样本、防答案泄露门控机制及分离的证据与计划评估指标。在14个源时序单元上,我们对比九种规划器家族与多数序列基线。提出一步证据融合(OSEF),在所有候选视频上对查询条件下的单元-跨度网格进行评分,并通过令牌全局适配器将完整软网格输入规划器,不提前裁剪窗口。OSEF在六个经认证可比较的单元中排名第一。在四组匹配的同任务与跨任务COIN和CrossTask单元中,相比增强版硬选择最先进方法,精确视频与计划成功率提升2.9–10.7个百分点;组件分析表明,令牌全局接口贡献最大单次提升。五组转换源单元接近多数序列基线,体现基准剩余潜力。附录提供模型构造器与评估代码。
原文摘要 · Abstract (English)
Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence must itself be retrieved. We introduce Cross-Video Scene Procedure Planning (CVSPP): given an answer-redacted start-goal query and K candidate videos, a model must retrieve the supporting video, localize the relevant window, and predict the action sequence. Two obstacles couple here. Same-task demonstrations share stages and windows, and an early hard selection passes the wrong scene chain to the planner. We build an eleven-source benchmark with typed negative roles, a fail-closed answer-leakage gate, and separate Evidence- and Plan-axis metrics. On its 14 source-horizon cells we adapt nine planner families against a majority-sequence floor. We then present One-Step Evidence Fusion (OSEF), which scores a query-conditioned cell-and-span lattice over all candidates and feeds the full soft lattice to the planner through a token-global adapter, cropping no window beforehand. OSEF ranks first on all six cells the benchmark certifies as method-rankable. On four matched same-task COIN and CrossTask cells it improves exact-video-and-plan success by 2.9-10.7 points over an enhanced hard-selection SOTA, and a component study assigns the largest single increment to the token-global interface. Five converted-source cells sit at or near the majority-sequence floor, the benchmark's remaining headroom. The supplementary package includes model constructors and evaluation code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。