arXiv:2412.11453cs.CLcs.AI2024-12被引 1

开发开源医学多模态模型评估工具,自动判断其问答能力。

ACE-$M^3$: Automatic Capability Evaluator for Multimodal Medical Models

  • 采用分支合并架构,结合医学标准生成详细分析与评分
  • 用奖励标记直接偏好优化,训练更快且性能不降
  • 适合需要高效评估医疗大模型的研究者和开发者

随着多模态大语言模型(MLLMs)在医疗领域的应用日益广泛,对其效能进行精准评估变得至关重要。传统评估指标如ROUGE、BLEU仅关注词元重叠,难以反映人类判断。人工评估虽可靠但成本高、难扩展。基于LLM的评估方法虽有前景,但医疗领域仍缺乏开源的多模态评估工具。为此,我们提出ACE-$M^3$,一个专为评估医学MLLM问答能力设计的开源自动评估系统。该系统采用分支-合并架构,依据标准医学评估准则提供详细分析与简洁评分;同时引入基于奖励标记的直接偏好优化(RTDPO)策略,在不牺牲性能的前提下显著减少训练时间。大量实验表明,ACE-$M^3$在评估医学多模态模型能力方面具有有效性。

原文摘要 · Abstract (English)

As multimodal large language models (MLLMs) gain prominence in the medical field, the need for precise evaluation methods to assess their effectiveness has become critical. While benchmarks provide a reliable means to evaluate the capabilities of MLLMs, traditional metrics like ROUGE and BLEU employed for open domain evaluation only focus on token overlap and may not align with human judgment. Although human evaluation is more reliable, it is labor-intensive, costly, and not scalable. LLM-based evaluation methods have proven promising, but to date, there is still an urgent need for open-source multimodal LLM-based evaluators in the medical field. To address this issue, we introduce ACE-$M^3$, an open-sourced \textbf{A}utomatic \textbf{C}apability \textbf{E}valuator for \textbf{M}ultimodal \textbf{M}edical \textbf{M}odels specifically designed to assess the question answering abilities of medical MLLMs. It first utilizes a branch-merge architecture to provide both detailed analysis and a concise final score based on standard medical evaluation criteria. Subsequently, a reward token-based direct preference optimization (RTDPO) strategy is incorporated to save training time without compromising performance of our model. Extensive experiments have demonstrated the effectiveness of our ACE-$M^3$ model\footnote{\url{https://huggingface.co/collections/AIUSRTMP/ace-m3-67593297ff391b93e3e5d068}} in evaluating the capabilities of medical MLLMs.

医学AI模型评估多模态自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。