用大模型动态调整训练策略,提升多模态文档检索准确率
Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document Retrieval
- 通过多视角对齐和双向对比学习增强跨模态匹配
- 在ViDoRe V2和MMEB上nDCG@5达65.2%和77.1%的领先效果
- 适合需要高精度多模态检索的系统开发者
视觉语言模型(VLMs)擅长跨模态数据映射,但现实文档的异构性和非结构化特征会破坏跨模态嵌入的一致性。近期的后期交互方法通过多向量表示提升了图像-文本对齐能力,但传统基于有限样本和静态策略的训练无法适应模型动态演化,导致跨模态检索混淆。为此,我们提出Evo-Retriever,一个基于新型观点-路径协同机制的LLM引导式课程演化检索框架。首先,采用多视角图像对齐,通过多尺度、多方向视角实现细粒度匹配。其次,设计双向对比学习策略生成‘难例查询’,建立互补学习路径以实现视觉与文本的消歧,重新平衡监督信号。最后,将上述协同产生的模型状态摘要输入至LLM元控制器,结合专家知识自适应调整训练课程,推动模型持续进化。在ViDoRe V2和MMEB(VisDoc)数据集上,Evo-Retriever达到当前最优性能,nDCG@5分别为65.2%和77.1%。
原文摘要 · Abstract (English)
Visual-language models (VLMs) excel at data mappings, but real-world document heterogeneity and unstructuredness disrupt the consistency of cross-modal embeddings. Recent late-interaction methods enhance image-text alignment through multi-vector representations, yet traditional training with limited samples and static strategies cannot adapt to the model's dynamic evolution, causing cross-modal retrieval confusion. To overcome this, we introduce Evo-Retriever, a retrieval framework featuring an LLM-guided curriculum evolution built upon a novel Viewpoint-Pathway collaboration. First, we employ multi-view image alignment to enhance fine-grained matching via multi-scale and multi-directional perspectives. Then, a bidirectional contrastive learning strategy generates "hard queries" and establishes complementary learning paths for visual and textual disambiguation to rebalance supervision. Finally, the model-state summary from the above collaboration is fed into an LLM meta-controller, which adaptively adjusts the training curriculum using expert knowledge to promote the model's evolution. On ViDoRe V2 and MMEB (VisDoc), Evo-Retriever achieves state-of-the-art performance, with nDCG@5 scores of 65.2% and 77.1%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。