arXiv:2412.15494cs.IR2024-12被引 1

用生成式方法增强视频搜索查询理解,提升检索效果。

PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search

  • 通过文本、图像、图像转文本三步生成扩展查询
  • 融合生成查询后在TV24数据集上表现优于原始查询
  • 适合关注多模态检索与LLM应用的研究者

今年,我们探索了生成式增强检索在TRECVid AVS任务中的应用。具体而言,通过文本到文本、文本到图像、图像到文本三种生成方式增强文本查询的理解,以解决词汇外问题。基于这些生成方式的不同组合以及原始查询的排名列表,我们提交了四组自动运行。对于手动运行,我们使用大语言模型(如GPT-4)根据搜索引擎的概念库重述测试查询,并人工检查确保重述查询中使用的概念均在概念库内。结果表明,原始查询与生成查询的融合在TV24查询集上表现优于原始查询。生成查询检索出的排名列表与原始查询不同。

原文摘要 · Abstract (English)

This year, we explore generation-augmented retrieval for the TRECVid AVS task. Specifically, the understanding of textual query is enhanced by three generations, including Text2Text, Text2Image, and Image2Text, to address the out-of-vocabulary problem. Using different combinations of them and the rank list retrieved by the original query, we submitted four automatic runs. For manual runs, we use a large language model (LLM) (i.e., GPT4) to rephrase test queries based on the concept bank of the search engine, and we manually check again to ensure all the concepts used in the rephrased queries are in the bank. The result shows that the fusion of the original and generated queries outperforms the original query on TV24 query sets. The generated queries retrieve different rank lists from the original query.

视频搜索生成增强多模态检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。