给大模型评分加权重,让关键信息更受重视。
Prompting a Weighting Mechanism into LLM-as-a-Judge in Two-Step: A Case Study
- 用提示词设计显式加入重要性权重机制。
- 评分与人工判断一致率平均提升6%。
- 适合需要精准评估生成文本的场景。
尽管大型语言模型(LLMs)已成为自然语言生成(NLG)任务评估的有力工具,但其有效性受限于无法恰当地权衡不同主题的重要性,常过度关注次要细节而忽视关键信息,导致评估结果误导。本文提出一种高效的提示词设计机制,以解决这一特定问题,并开展案例研究。通过包含显式重要性权重机制的战略性提示工程,提升了使用LLM作为评判者时优先处理相关信息的能力,表现为人类对齐率(HAR)指标平均提升6%。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have emerged as promising tools for evaluating Natural Language Generation (NLG) tasks, their effectiveness is limited by their inability to appropriately weigh the importance of different topics, often overemphasizing minor details while undervaluing critical information, leading to misleading assessments. Our work proposes an efficient prompt design mechanism to address this specific limitation and provide a case study. Through strategic prompt engineering that incorporates explicit importance weighting mechanisms, we enhance using LLM-as-a-Judge ability to prioritize relevant information effectively, as demonstrated by an average improvement of 6% in the Human Alignment Rate (HAR) metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。