自动挖掘网络内容,发现语言模型输出的隐含评估标准。
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
- 从专家在线指南中提取证据,生成任务相关的评估标准。
- 发现的标准虽未明说但精准具体,能指导模型改进。
- 适合想提升模型输出质量的研究者和开发者使用。
在结构化写作任务中,语言模型输出的评估通常依赖于预设的显性标准。然而,高质量响应还需满足隐含的高级要求,如学术演讲应有引人入胜的开场、明确的研究问题和结论。为识别这些隐含标准,我们提出EvalAgent框架,通过挖掘专家撰写的在线指导材料,自动生成基于可靠外部来源的多样化、长尾评估标准。实验表明,这些由EvalAgent生成的标准往往是用户提示中未直接提及的(隐含),但具有高度词汇精确性(具体)。初始模型响应常不满足这些标准,但可通过优化实现达标。此外,结合大模型生成与EvalAgent的标准,比单独使用大模型能发现更多人类认可的评估维度。
原文摘要 · Abstract (English)
Evaluation of language model outputs on structured writing tasks is typically conducted with a number of desirable criteria presented to human evaluators or large language models (LLMs). For instance, on a prompt like "Help me draft an academic talk on coffee intake vs research productivity", a model response may be evaluated for criteria like accuracy and coherence. However, high-quality responses should do more than just satisfy basic task requirements. An effective response to this query should include quintessential features of an academic talk, such as a compelling opening, clear research questions, and a takeaway. To help identify these implicit criteria, we introduce EvalAgent, a novel framework designed to automatically uncover nuanced and task-specific criteria. EvalAgent first mines expert-authored online guidance. It then uses this evidence to propose diverse, long-tail evaluation criteria that are grounded in reliable external sources. Our experiments demonstrate that the grounded criteria produced by EvalAgent are often implicit (not directly stated in the user's prompt), yet specific (high degree of lexical precision). Further, EvalAgent criteria are often not satisfied by initial responses but they are actionable, such that responses can be refined to satisfy them. Finally, we show that combining LLM-generated and EvalAgent criteria uncovers more human-valued criteria than using LLMs alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。