测试大模型与人类互动反馈的能力,发现顶尖模型仍难有效改进。
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
- 设计交互框架InterFeedback,自动评估模型响应人类反馈能力。
- 顶级模型OpenAI-o1在反馈下平均得分不足50%,表现不佳。
- 适合关注人机协作、模型可解释性的研究者与开发者。
现有基准未检验大型多模态模型(LMMs)与人类用户互动智能的能力,而这对通用人工智能助手的发展至关重要。我们设计了InterFeedback交互框架,可应用于任意LMM和数据集,实现该能力的自主评估。在此基础上,提出InterFeedback-Bench,利用两个代表性数据集MMMU-Pro和MathVerse,评估10个开源LMM的交互智能。此外,我们构建了InterFeedback-Human,一个包含120个案例的新标注数据集,用于人工评测如OpenAI-o1和Claude-Sonnet-4等领先模型的交互表现。评估结果显示,即使最先进的模型OpenAI-o1也难以根据人类反馈优化回应,平均得分低于50%。研究揭示了提升LMM理解并利用反馈能力的迫切需求。
原文摘要 · Abstract (English)
Existing benchmarks do not test Large Multimodal Models (LMMs) on their interactive intelligence with human users, which is vital for developing general-purpose AI assistants. We design InterFeedback, an interactive framework, which can be applied to any LMM and dataset to assess this ability autonomously. On top of this, we introduce InterFeedback-Bench which evaluates interactive intelligence using two representative datasets, MMMU-Pro and MathVerse, to test 10 different open-source LMMs. Additionally, we present InterFeedback-Human, a newly collected dataset of 120 cases designed for manually testing interactive performance in leading models such as OpenAI-o1 and Claude-Sonnet-4. Our evaluation results indicate that even the state-of-the-art LMM, OpenAI-o1, struggles to refine its responses based on human feedback, achieving an average score of less than 50%. Our findings point to the need for methods that can enhance LMMs' capabilities to interpret and benefit from feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。