首个面向多轮跨模态交互的偏好数据集,用真人反馈训练模型理解复杂对话。
InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- 构建首个基于人类反馈的多轮跨模态交互偏好数据集
- 包含15.6k提示、52.6k对话实例和32.4k偏好对,覆盖九个维度
- 适合研究多轮对话、人机交互与大模型对齐的学者使用
随着多模态大模型在复杂任务上的持续进步,一个关键问题浮现:哪些核心能力仍缺失?人类学习的关键在于与环境的持续互动——不仅限于语言,还包括多模态理解与生成。为逼近人类智能水平,模型也需支持多轮、多模态交互,能理解交错的多模态上下文并作出连贯回应。本文提出InterMT——首个基于真实人类反馈的多轮多模态交互偏好数据集。通过专家标注引导过程,强调人类监督的重要性,因当前多模态大模型缺乏此类复杂交互能力。InterMT从全局与局部层面捕捉人类偏好,涵盖九个子维度,包含15.6k提示、52.6k多轮对话实例及32.4k人工标注的偏好对。为弥补模型在多模态理解与生成上的不足,引入工具增强的代理工作流,构建多轮问答实例。进一步提出InterMT-Bench,用于评估多模态大模型在辅助判断任务中的表现。通过应用如裁判审核等场景,揭示了裁判模型的多轮扩展规律。我们开源该数据集,以推动多模态大模型迈向下一阶段对齐。项目主页见 https://pku-intermt.github.io。
原文摘要 · Abstract (English)
As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: What essential capabilities are still missing? A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation. To move closer to human-level intelligence, models must similarly support multi-turn, multimodal interaction. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges. In this work, we present an initial exploration through the InterMT -- the first preference dataset for multi-turn multimodal interaction, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. InterMT captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances. To further this goal, we introduce InterMT-Bench to assess the ability of MLLMs in assisting judges with multi-turn, multimodal tasks. We demonstrate the utility of \InterMT through applications such as judge moderation and further reveal the multi-turn scaling law of judge model. We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step. Our project website can be found at https://pku-intermt.github.io .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。