arXiv:2505.15128cs.IRcs.MM2025-05中稿 · ICMR 2025被引 2

让视频搜索在用户反馈不一致时仍能精准定位目标。

Robust Relevance Feedback for Interactive Known-Item Video Search

  • 用多子感知模型分解用户判断,提升反馈稳定性。
  • 在V3C数据集上将初始排名10-50的目标优化至首位成功率超60%。
  • 适合需要高鲁棒性交互式视频搜索的场景。

已知项搜索(KIS)仅有一个目标,传统相关性反馈因需识别多个正例而难以应用。PicHunter通过让用户从显示集合中选择与唯一目标最相似的前k个实例来解决此问题。理想情况下,当用户感知与机器相似性判断一致时,数次迭代即可将目标推至首位。但现实中,由于嵌入特征缺乏可解释性,要求用户持续一致判断不切实际。为此,我们引入成对相对判断反馈,以减少不一致反馈的影响;并分解用户感知为多个独立嵌入空间的子感知,假设用户更可能与其中一部分而非单一表示对齐。我们构建预测用户模型,基于每次反馈估计子感知组合,并训练其过滤不一致的子感知。在大规模开放域数据集V3C上的实验表明,该模型可将初始排名在10至50之间的超过60%的目标优化至首位;即使初始排名在1,000至5,000之间,也有超40%的成功率提升至首位,验证了在不一致反馈下相关性反馈在KIS中的增强鲁棒性。

原文摘要 · Abstract (English)

Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal conditions, when the user's perception aligns closely with the machine's perception of similarity, consistent and precise judgments can elevate the target to the top position within a few iterations. However, in practical scenarios, expecting users to provide consistent judgments is often unrealistic, especially when the underlying embedding features used for similarity measurements lack interpretability. To enhance robustness, we first introduce a pairwise relative judgment feedback that improves the stability of top-k selections by mitigating the impact of misaligned feedback. Then, we decompose user perception into multiple sub-perceptions, each represented as an independent embedding space. This approach assumes that users may not consistently align with a single representation but are more likely to align with one or several among multiple representations. We develop a predictive user model that estimates the combination of sub-perceptions based on each user feedback instance. The predictive user model is then trained to filter out the misaligned sub-perceptions. Experimental evaluations on the large-scale open-domain dataset V3C indicate that the proposed model can optimize over 60% search targets to the top rank when their initial ranks at the search depth between 10 and 50. Even for targets initially ranked between 1,000 and 5,000, the model achieves a success rate exceeding 40% in optimizing ranks to the top, demonstrating the enhanced robustness of relevance feedback in KIS despite inconsistent feedback.

视频搜索相关性反馈用户建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。