arXiv:2507.00092cs.AIcs.CL2025-07被引 1

让大模型反向解释自己怎么想的,提升推理透明度。

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models

  • 用逆向注意力机制让模型回溯并解释自身推理路径。
  • 在AQUA-RAT上准确率达74.6%,解释质量获92.1%人类偏好分。
  • 适合关注AI可解释性、安全与教育的应用场景。

大型语言模型(LLMs)在链式思维(CoT)提示下展现出解决复杂推理任务的能力,但其决策过程仍属黑箱。本文提出文本逆向推理新范式,使LLMs能事后分解并解释自身的推理链条。所提出的SAGE-nano模型(40亿参数)采用元认知结构,通过注意力机制回溯识别关键决策点,并生成推理选择说明。不同于传统前向推理,逆向推理揭示为何选择特定推理路径。在AQUA-RAT、CommonsenseQA及定制化基准测试中,SAGE-nano在逻辑谜题、数学问题和伦理困境上表现优异,推理准确率达74.6%(AQUA-RAT),解释质量人类偏好得分高达92.1%,性能接近Claude-3.5 Sonnet与GPT-4o。贡献包括:(i) 首个基于逆向推理的自省框架;(ii) 反向注意力流的新元学习机制;(iii) 推理透明性综合评估体系;(iv) 证明逆向推理可同步提升可解释性与推理能力。本工作为透明人工智能开辟新路径,填补了人工智能安全、教育与科学发现中的重要空白。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities at solving complex reasoning tasks with Chain-of-Thought (CoT) prompting, but their decision-making processes remain somewhat blackbox. We introduce textbfinverse reasoning, a novel paradigm enabling LLMs to decompose and explain their own reasoning chains post-hoc. Our approach, used in SAGE-nano, a 4-billion-parameter reasoning model, employs a metacognitive structure that reflects back via attention processes to identify major decision points and generate explanations of reasoning choices. While typical CoT approaches are directed towards forward reasoning generation, inverse reasoning provides insight into why specific reasoning chains were selected over others. Through thorough testing of logical reasoning puzzles, math problems and ethical dilemmas from AQUA-RAT, CommonsenseQA, and customized benchmarks, we demonstrate that SAGE-nano is at the cutting edge both on reasoning accuracy (74.6% on AQUA-RAT) and explanation quality (92.1% human preference score) for its task, and offers performance almost on par with models like Claude-3.5 Sonnet or GPT-4o. Our contributions are: (i) the first rigorous framework for LLM self-reflection via inverse reasoning, (ii) a novel metalearning framework to reverse the attention flow, (iii) comprehensive evaluation frameworks for reasoning transparency, and (iv) evidence that increasing reasoning using inverse reasoning improves interpretability along with reasoning performance. Our work creates new avenues for transparent AI systems and closes significant gaps in AI safety, education, and scientific discovery.

可解释AI自我反思逆向推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。