arXiv:2509.01610cs.CV2025-09ICCV被引 1

让大模型互相评阅,不用人标注也能提升性能

Improving Large Vision and Language Models by Learning from a Panel of Peers

  • 用多个模型组成评审团,轮流评估和改进彼此输出
  • 15个基准测试平均得分从48%提升至57%
  • 适合想降低人工标注成本的模型优化场景

传统大视觉语言模型对齐方法依赖人工标注的偏好数据,成本高;机器生成的偏好数据质量有限;自监督数据常引入幻觉。为此,我们提出受人类协作学习启发的同伴评审学习框架。该方法构建一个由多个大视觉语言模型组成的评审团,通过迭代自提升过程,共同生成、评估并优化输出。模拟课堂互评机制,模型在精心设计的提示集上进行协同生成与修正。实验表明,该方法无需大量人工标注数据即可显著提升模型性能。在15个基准测试中,平均得分从48%提升至57%,验证了同伴评估作为可扩展对齐替代方案的潜力。

原文摘要 · Abstract (English)

Traditional alignment methods for Large Vision and Language Models (LVLMs) primarily rely on human-curated preference data. Human-generated preference data is costly; machine-generated preference data is limited in quality; and self-supervised preference data often introduces hallucinations. To overcome these limitations, we propose a novel Panel-of-Peers learning framework inspired by collaborative learning among humans. This approach leverages a panel of LVLMs, each evaluating and learning from their collective outputs through an iterative self-improvement process. By simulating a peer review system, our models generate, assess, and refine outputs in response to a curated set of prompts, mimicking a classroom learning environment. We demonstrate that this methodology enhances model performance without requiring extensive human-labeled datasets. Our experiments show significant improvement across multiple benchmarks, demonstrating the potential of peer evaluations as a scalable alternative to self-supervised alignment. Notably, we show that Panel-of-Peers increases the average score on fifteen benchmarks from 48% to 57%

大模型对齐同伴学习自提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。