arXiv:2505.13972cs.CL2025-05中稿 · INLG 2025, camera-…被引 5

选对评估模型才能可靠判断大模型反事实生成效果

Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals

  • 用不同关系的模型做评估,发现独立非微调的判别器最可靠
  • 独立判别器下标签翻转率最高,与人工评估结果最接近
  • 自动化评估仍有差距,反事实数据增强需人工干预

反事实样本广泛用于通过反事实数据增强(CDA)提升大语言模型性能与鲁棒性。然而,用于评估标签翻转(衡量反事实有效性主要指标)的判别模型选择会导致结果不一致。本文定义了生成器与判别器之间的四种关系:同一模型、同家族模型、独立模型、蒸馏关系。通过两个前沿方法、三个数据集、四类生成器与十五个判别器的实验,结合用户研究(n=90),发现与生成器无关联且未微调的判别器能提供最可靠的标签翻转评估。当生成器与判别器关系与用户研究中的认知一致时,模型性能与鲁棒性更优。但最优判别器结果与人工评估仍存在显著差距,表明全自动的CDA流程可能不足,需引入人工干预。

原文摘要 · Abstract (English)

Counterfactual examples are widely employed to enhance the performance and robustness of large language models (LLMs) through counterfactual data augmentation (CDA). However, the selection of the judge model used to evaluate label flipping, the primary metric for assessing the validity of generated counterfactuals for CDA, yields inconsistent results. To decipher this, we define four types of relationships between the counterfactual generator and judge models: being the same model, belonging to the same model family, being independent models, and having an distillation relationship. Through extensive experiments involving two state-of-the-art LLM-based methods, three datasets, four generator models, and 15 judge models, complemented by a user study (n = 90), we demonstrate that judge models with an independent, non-fine-tuned relationship to the generator model provide the most reliable label flipping evaluations. Relationships between the generator and judge models, which are closely aligned with the user study for CDA, result in better model performance and robustness. Nevertheless, we find that the gap between the most effective judge models and the results obtained from the user study remains considerably large. This suggests that a fully automated pipeline for CDA may be inadequate and requires human intervention.

大模型反事实生成评估方法人类评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。