arXiv:2506.04516cs.CL2025-06被引 2

用小模型引导大模型,提升对话评估的准确性

DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation

  • 小模型指导大模型生成初始评分,再由小模型修正分数
  • 在多个基准上与人工判断更一致,优于现有方法
  • 适合需要可靠对话评估的开放域任务

大型语言模型(LLMs)在多数任务中表现优异,但在存在多种合理回应的模糊场景中表现不稳定,结果不可靠;小型语言模型(SLMs)在这些场景中更具鲁棒性,但易受误导或对抗性输入影响。我们发现LLM擅长处理负例,而SLM在正例上表现更好。为此提出SLIDE(Small and Large Integrated for Dialogue Evaluation),通过自适应加权融合大小模型。在此基础上进一步提出双精炼评估(DRE):(1) SLM生成的见解指导LLM生成初始评估;(2) SLM推导的调整量用于优化LLM评分,提升准确性。实验表明,DRE在多个基准上均优于现有方法,与人类判断更一致。本工作展示了融合大小模型可构建更可靠的对话评估工具,尤其适用于开放域任务。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel at many tasks but struggle with ambiguous scenarios where multiple valid responses exist, often yielding unreliable results. Conversely, Small Language Models (SLMs) demonstrate robustness in such scenarios but are susceptible to misleading or adversarial inputs. We observed that LLMs handle negative examples effectively, while SLMs excel with positive examples. To leverage their complementary strengths, we introduce SLIDE (Small and Large Integrated for Dialogue Evaluation), a method integrating SLMs and LLMs via adaptive weighting. Building on SLIDE, we further propose a Dual-Refinement Evaluation (DRE) method to enhance SLM-LLM integration: (1) SLM-generated insights guide the LLM to produce initial evaluations; (2) SLM-derived adjustments refine the LLM's scores for improved accuracy. Experiments demonstrate that DRE outperforms existing methods, showing stronger alignment with human judgment across diverse benchmarks. This work illustrates how combining small and large models can yield more reliable evaluation tools, particularly for open-ended tasks such as dialogue evaluation.

对话评估模型融合LLMSLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。