arXiv:2510.04396cs.MMcs.IR2025-10

优化视频检索的关键帧布局,提升用户找目标的效率与准确率。

Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections

  • 对比七种关键帧排布方式,发现视频分组布局最高效。
  • 四列保持排序的网格准确率最高,但可能延迟重要结果出现。
  • 混合设计可兼顾排名位置与浏览便利性,适合多种场景。

多模态深度学习模型通过为文本查询排序关键帧来支持交互式视频检索。尽管如此,用户仍需手动浏览结果以定位目标。关键帧在搜索网格中的布局显著影响浏览效率与用户表现,但该问题尚未得到充分研究。我们对49名参与者进行了实验,评估了七种关键帧布局在视觉已知项搜索任务中的表现。除了效率与准确率外,还分析了遗漏现象与布局特征的关系。结果表明,视频分组布局在效率上最优,而四列保持排序的网格在准确率上最高。排序网格虽便于快速扫描无趣区域,但会将相关项推至不显眼位置,导致首次到达时间延迟并增加遗漏。这些发现推动了混合设计的发展:保留高排名项的位置,同时对余下内容进行排序或分组,为视频检索之外的网格搜索提供指导。

原文摘要 · Abstract (English)

Multimodal deep-learning models power interactive video retrieval by ranking keyframes in response to textual queries. Despite these advances, users must still browse ranked candidates manually to locate a target. Keyframe arrangement within the search grid highly affects browsing effectiveness and user efficiency, yet remains underexplored. We report a study with 49 participants evaluating seven keyframe layouts for the Visual Known-Item Search task. Beyond efficiency and accuracy, we relate browsing phenomena, such as overlooks, to layout characteristics. Our results show that a video-grouped layout is the most efficient, while a four-column, rank-preserving grid achieves the highest accuracy. Sorted grids reveal potentials and trade-offs, enabling rapid scanning of uninteresting regions but down-ranking relevant targets to less prominent positions, delaying first arrival times and increasing overlooks. These findings motivate hybrid designs that preserve positions of top-ranked items while sorting or grouping the remainder, and offer guidance for searching in grids beyond video retrieval.

视频检索界面设计人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。