arXiv:2607.09876cs.CVcs.AI2026-07中稿 · ECCV

用结构化文本+专家级查询,精准检索野生动物相机陷阱视频。

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

论文配图:Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
图 1 · 摘自论文原文
  • 将视频动作定位转为结构化文本,再由大模型解析匹配
  • 在135个生态查询上达到34%的集合F1分数
  • 兼顾精度与可解释性,适合生态研究者使用

自动从大规模相机陷阱数据集中检索视频仍具挑战。基于大视频语言模型(VLM)的文本到视频检索(TVR)方法有望通过简单文本查询检索感兴趣事件。然而,现有方法常缺乏时空理解能力,且难以泛化至生态数据。本文提出首个相机陷阱TVR基准——Prompting-MammAlps,并构建细粒度、可解释的TVR方法。具体而言,训练视觉变压器实现时空动作定位,将其输出转化为结构化文本描述视频内容;独立地,采用受动物行为学启发的查询,通过基于大语言模型(LLM)的编码代理解析每段结构化文本并完成检索。利用自定义解析库函数调用,降低大模型幻觉风险并提升可解释性。该方法在包含135个生态相关查询和775个候选视频的测试集上取得34%的集合F1分数,优于最佳零样本VLM的18%表现,且具备可解释性。

原文摘要 · Abstract (English)

Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large video-language models (VLMs) have potential to retrieve events of interest by describing them with simple text queries. However, current methods often lack spatiotemporal understanding and do not generalize well to ecological data. In this work, we introduce Prompting-MammAlps, the first camera-trap TVR benchmark, and propose a fine-grained and interpretable TVR method. Specifically, we trained a vision transformer to perform spatiotemporal action localization, and convert its output to structured text, describing each video. Independently, ethology-inspired queries are processed by a Large-Language Model (LLM) based coding agent to parse the structured text per video and retrieve videos accordingly. We harnessed the LLM to use functions from a custom parsing library to minimize the risk of LLM hallucinations and to improve method interpretability. This retrieval approach applied on the Prompting-MammAlps benchmark achieved a set-based F1-score of 34\% on a test set of 135 ecologically-relevant queries and 775 candidate videos. In comparison the best zero-shot VLM achieved a F1-score of 18\%, while also lacking interpretability. Project page: https://cnai.epfl.ch/prompting-mammalps

视频检索生态研究大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。