arXiv:2409.20565cs.CL2024-09被引 4

用少量人工标注构建可靠医学解释生成评估系统

Ranking Over Scoring: Towards Reliable and Robust Automated Evaluation of LLM-Generated Medical Explanatory Arguments

  • 通过代理任务和排序机制实现与人类评价对齐的评估方法
  • 仅需每任务1个标注样例,且对对抗攻击具有鲁棒性
  • 适合医疗领域大模型生成质量评估,降低人工成本

评估大语言模型生成的文本已成为关键挑战,尤其在医学等专业领域。本文提出一种新型评估方法,基于代理任务和排序机制,使评估结果更贴近人类评判标准,克服了以大模型作为裁判时常见的偏差问题。实验表明,该评估器对对抗攻击(包括非论证性文本)具有鲁棒性。此外,用于训练评估器的人工标注论证只需每代理任务一个示例。通过分析多个大模型生成的论证,建立了一套判断代理任务是否适用于医学解释生成评估的方法,仅需五个样本和两名人类专家即可完成验证。

原文摘要 · Abstract (English)

Evaluating LLM-generated text has become a key challenge, especially in domain-specific contexts like the medical field. This work introduces a novel evaluation methodology for LLM-generated medical explanatory arguments, relying on Proxy Tasks and rankings to closely align results with human evaluation criteria, overcoming the biases typically seen in LLMs used as judges. We demonstrate that the proposed evaluators are robust against adversarial attacks, including the assessment of non-argumentative text. Additionally, the human-crafted arguments needed to train the evaluators are minimized to just one example per Proxy Task. By examining multiple LLM-generated arguments, we establish a methodology for determining whether a Proxy Task is suitable for evaluating LLM-generated medical explanatory arguments, requiring only five examples and two human experts.

大模型评估医学AI代理任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。