arXiv:2501.12975cs.CL2025-01被引 6

针对小模型幻觉问题,提出多层评估框架与新指标

OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models

  • 设计分层评估框架,引入上下文影响得分
  • 发现小模型事实分析强但上下文推理弱
  • 简单思维链策略可显著提升实用性能

大语言模型虽能力强,但训练和推理需大量计算资源。小型语言模型(参数少于100亿)在多个任务上表现良好,但同样存在幻觉问题。现有评估基准大多未专门针对小模型,且其在不同基准上表现差异大。本文提出OnionEval,一种多层结构评估框架,包含专门的上下文影响得分(CI)指标,用于评估小模型在不同上下文层级上的事实冲突型幻觉倾向。实验显示,小模型在事实分析方面表现优异,但在上下文推理方面存在不足。进一步研究发现,采用简单的思维链(Chain-of-Thought)策略能显著缓解此问题,提升小模型在真实场景中的实用性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are highly capable but require significant computational resources for both training and inference. Within the LLM family, smaller models (those with fewer than 10 billion parameters) also perform well across various tasks. However, these smaller models share similar limitations to their larger counterparts, including the tendency to hallucinate. Despite the existence of many benchmarks to evaluate hallucination in LLMs, few have specifically focused on small LLMs (SLLMs). Additionally, SLLMs show widely varying performance across different benchmarks. In this paper, we introduce OnionEval, a multi-layer structured framework with a specific metric called the context-influence score (CI), designed to effectively assess the fact-conflicting hallucination tendencies of small LLMs across different contextual levels. Our experimental results reveal a key feature of SLLMs: they excel in factual analysis but face challenges with context reasoning. Further investigation shows that a simple Chain-of-Thought strategy can significantly reduce these limitations, improving the practical usefulness of SLLMs in real-world applications.

小模型幻觉评估思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。