针对多模态问答挑战,构建了融合检索与微调的高效系统。
DB3 Team's Solution For Meta KDD Cup' 25
- 分任务设计领域专用检索管道,整合图像知识图谱与对话历史
- 采用SFT/DPO/RL联合训练,有效控制大模型幻觉现象
- 在三类任务中分别获第1、2名,擅长第一人称视角理解
本文介绍db3团队在KDD Cup'25 Meta CRAG-MM挑战赛中的获奖方案。针对该挑战独特的多模态、多轮问答评测基准(CRAG-MM),我们构建了集成定制化检索流程与统一LLM微调策略的综合框架。解决方案包含:(1) 针对不同任务设计的领域特定检索管道,涵盖图像索引的知识图谱、网络来源及多轮对话历史;(2) 基于SFT、DPO和强化学习的先进拒答训练机制。系统在任务1中获得第2名,任务2中获第2名,任务3中夺得第1名,凭借对第一人称视角问题的优异处理能力,赢得全赛事优胜大奖。
原文摘要 · Abstract (English)
This paper presents the db3 team's winning solution for the Meta CRAG-MM Challenge 2025 at KDD Cup'25. Addressing the challenge's unique multi-modal, multi-turn question answering benchmark (CRAG-MM), we developed a comprehensive framework that integrates tailored retrieval pipelines for different tasks with a unified LLM-tuning approach for hallucination control. Our solution features (1) domain-specific retrieval pipelines handling image-indexed knowledge graphs, web sources, and multi-turn conversations; and (2) advanced refusal training using SFT, DPO, and RL. The system achieved 2nd place in Task 1, 2nd place in Task 2, and 1st place in Task 3, securing the grand prize for excellence in ego-centric queries through superior handling of first-person perspective challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。