arXiv:2510.17921cs.CLcs.AI2025-10NeurIPS

无需人工评估,用注意力窗口识别大模型数学解法的创意水平。

CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections

  • 通过分析提示与输出中各部分的注意力权重,划分解法类型。
  • 在4545道竞赛题上,对5个7-8B模型的创意检测准确率超越现有方法。
  • 适合关注大模型推理创意性评估的研究者与开发者使用。

大语言模型(LLM)在强化学习训练下,于数学与编程等挑战性任务中表现出色,即使模型规模较小。然而,尽管任务准确率提升显著,推理任务中的创意评估仍被忽视,远不如写作任务受关注。这主要源于两大挑战:一是创意范围难以界定,二是评估过程常需人工介入。为此,我们提出CLAWS,一种无需人工评价的方法,通过分析提示与输出中各部分的注意力权重,将数学解法分为典型、创意和幻觉三类。CLAWS在5个7-8B规模的数学强化学习模型(DeepSeek、Qwen、Mathstral、OpenMath2、Oreal)上,优于五种现有白盒检测方法(Perplexity、Logit Entropy、Window Entropy、Hidden Score、Attention Score)。我们在181场数学竞赛(如AJHSME、AMC、AIME)中收集的4545道题目上验证了该方法的有效性。

原文摘要 · Abstract (English)

Recent advances in enhancing the reasoning ability of large language models (LLMs) have been remarkably successful. LLMs trained with reinforcement learning (RL) for reasoning demonstrate strong performance in challenging tasks such as mathematics and coding, even with relatively small model sizes. However, despite these improvements in task accuracy, the assessment of creativity in LLM generations has been largely overlooked in reasoning tasks, in contrast to writing tasks. The lack of research on creativity assessment in reasoning primarily stems from two challenges: (1) the difficulty of defining the range of creativity, and (2) the necessity of human evaluation in the assessment process. To address these challenges, we propose CLAWS, a method that defines and classifies mathematical solutions into typical, creative, and hallucinated categories without human evaluation, by leveraging attention weights across prompt sections and output. CLAWS outperforms five existing white-box detection methods (Perplexity, Logit Entropy, Window Entropy, Hidden Score, and Attention Score) on five 7-8B math RL models (DeepSeek, Qwen, Mathstral, OpenMath2, and Oreal). We validate CLAWS on 4545 math problems collected from 181 math contests (AJHSME, AMC, AIME).

大模型评估创意检测注意力机制数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。