构建图形化常识知识库,量化大模型的常识推理能力。
Towards Quantifying Commonsense Reasoning with Mechanistic Insights

- 用图形结构标注37种日常活动的隐含知识。
- 可生成约10^17个常识问题,支持严格评估。
- 发现大模型推理集中在特定模块,适合研究模型内部机制。
常识推理涉及人类通过与世界互动获得的隐含知识。近期,大语言模型(LLMs)的常识理解多通过文本任务评估。本文提出,可通过图形结构表征这种理解,并创建针对37种日常活动的隐含知识标注方案。该资源可生成约10^17个常识问题,支持对LLMs常识推理能力的严格评估。同时,近期大模型的卓越表现引发对其是否真具备野外推理能力、以及推理如何在模型内部发生等问题的质疑。本资源论文通过提出设计机制,弥合这一空白。研究发现,当面对常识查询时,大模型中的推理组件具有显著局部化特征,在决策中起关键作用。
原文摘要 · Abstract (English)
Commonsense reasoning deals with the implicit knowledge that is well understood by humans and typically acquired via interactions with the world. In recent times, commonsense reasoning and understanding of various LLMs have been evaluated using text-based tasks. In this work, we argue that a proxy of this understanding can be maintained as a graphical structure that can further help to perform a rigorous evaluation of commonsense reasoning abilities about various real-world activities. We create an annotation scheme for capturing this implicit knowledge in the form of a graphical structure for 37 daily human activities. We find that the created resource can be used to frame an enormous number of commonsense queries (~ 10^{17}), facilitating rigorous evaluation of commonsense reasoning in LLMs. Moreover, recently, the remarkable performance of LLMs has raised questions about whether these models are truly capable of reasoning in the wild and, in general, how reasoning occurs inside these models. In this resource paper, we bridge this gap by proposing design mechanisms that facilitate research in a similar direction. Our findings suggest that the reasoning components are localized in LLMs that play a prominent role in decision-making when prompted with a commonsense query.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。