arXiv:2505.23843cs.CLcs.LG2025-05被引 2

揭示大模型多轮推理评估中的幻觉问题,提出更可靠的评测标准

Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks

  • 通过分析推理路径发现模型存在捷径行为和过早终止问题
  • 现有评测方法常误导结果,无法真实反映模型的横向思维能力
  • 引入人类对比与多样化指标,提升评估可信度,适合模型评测研究者

多轮不完整信息任务对评估大语言模型(LLMs)的横向思维能力至关重要。当前研究主要依赖多个基准测试和自动化评估指标。然而,我们的研究表明,现有方法存在显著局限,常产生误导性结果,难以揭示关键问题,如捷径取巧、僵化模式和过早任务终止。这些问题掩盖了模型真实的推理能力,削弱了评估的可靠性。为此,我们提出一套优化的评估标准,包括推理路径审查、多样化评估指标以及与人类表现的对比分析。

原文摘要 · Abstract (English)

Multi-round incomplete information tasks are crucial for evaluating the lateral thinking capabilities of large language models (LLMs). Currently, research primarily relies on multiple benchmarks and automated evaluation metrics to assess these abilities. However, our study reveals novel insights into the limitations of existing methods, as they often yield misleading results that fail to uncover key issues, such as shortcut-taking behaviors, rigid patterns, and premature task termination. These issues obscure the true reasoning capabilities of LLMs and undermine the reliability of evaluations. To address these limitations, we propose a refined set of evaluation standards, including inspection of reasoning paths, diversified assessment metrics, and comparative analyses with human performance.

大模型评测横向思维推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。