LLandMark用多智能体框架实现文化地标感知的跨模态视频检索
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
- 分四阶段协作的多智能体架构,专攻复杂查询理解与推理
- 通过地标重构提示词,使CLIP在越南场景下匹配准确率显著提升
- 支持零样本图像输入,适合需本地化语境理解的文旅应用
随着视频数据日益多样化和规模化,亟需具备多模态理解、自适应推理与领域知识融合能力的检索系统。本文提出LLandMark,一种面向真实复杂查询的地标感知多模态视频检索模块化多智能体框架。该框架包含四个协同阶段:查询解析与规划、地标推理、多模态检索与重排序答案生成。核心组件‘地标知识智能体’可识别文化或空间地标,并将其转化为描述性视觉提示,增强CLIP在越南场景下的语义匹配能力。为扩展能力,引入基于大语言模型(Gemini 2.5 Flash)的图像到图像管道,实现地标自动检测、图像搜索查询生成、代表性图像检索及基于CLIP的视觉相似度匹配,无需人工提供图像输入。此外,结合Gemini与LlamaIndex的OCR精炼模块提升了越南语文本识别效果。实验表明,LLandMark实现了自适应、文化情境化且可解释的检索性能。
原文摘要 · Abstract (English)
The increasing diversity and scale of video data demand retrieval systems capable of multimodal understanding, adaptive reasoning, and domain-specific knowledge integration. This paper presents LLandMark, a modular multi-agent framework for landmark-aware multimodal video retrieval to handle real-world complex queries. The framework features specialized agents that collaborate across four stages: query parsing and planning, landmark reasoning, multimodal retrieval, and reranked answer synthesis. A key component, the Landmark Knowledge Agent, detects cultural or spatial landmarks and reformulates them into descriptive visual prompts, enhancing CLIP-based semantic matching for Vietnamese scenes. To expand capabilities, we introduce an LLM-assisted image-to-image pipeline, where a large language model (Gemini 2.5 Flash) autonomously detects landmarks, generates image search queries, retrieves representative images, and performs CLIP-based visual similarity matching, removing the need for manual image input. In addition, an OCR refinement module leveraging Gemini and LlamaIndex improves Vietnamese text recognition. Experimental results show that LLandMark achieves adaptive, culturally grounded, and explainable retrieval performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。