用逆向学习自动生成适合大模型的评测提示,提升评估效率与稳定性。
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
- 通过逆向映射从生成结果反推输入指令,自动构造评测提示。
- 仅需一个样本即可生成有效提示,无需人工调参。
- 适用于追求高效、稳定评测的大模型研究者。
自然语言生成系统的评估因输出多样性而困难。人工评估虽为金标准,却存在不一致、缺乏标准化及群体偏差问题,影响可复现性。基于大模型的评估方法虽具可扩展性,但对提示设计极为敏感,微小改动可能导致显著差异。本文提出一种逆向学习方法,通过学习从模型输出到输入指令的反向映射,实现针对特定模型的高效评测提示自动生成。该方法仅需单个评估样本,无需耗时的人工提示工程,显著提升评估效率与鲁棒性。本工作推动了更稳健、高效的基于大模型评估的新方向。
原文摘要 · Abstract (English)
Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardisation, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。