arXiv:2508.19988cs.CL2025-08ACL被引 2

测试大模型在真实场景中混合常识与数学推理的能力,发现性能大幅下降。

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

  • 设计新基准AgentCoMa,要求同时完成常识和数学推理
  • 大模型组合任务平均准确率比单任务低30%以上
  • 人类非专家能轻松应对,揭示模型在跨类型推理上的脆弱性

大型语言模型(LLMs)在需要多步推理的复杂常识和数学问题上已取得高准确率。然而,现有组合性基准大多只聚焦于常识或数学推理,而解决真实世界任务的LLM代理需要两者结合。本文提出一个面向代理的常识与数学推理基准(AgentCoMa),每个任务均需完成一次常识推理和一次数学推理。我们在61个不同规模、模型族和训练策略的LLM上进行了测试。结果显示,尽管模型在单独处理两类任务时表现良好,但当两者结合时,平均准确率下降近30%。这一性能差距显著大于以往仅混合同类型推理步骤的基准。相比之下,非专家人类标注者在组合问题及单个步骤上均能以高准确率完成。我们还通过可解释性研究分析了神经元模式、注意力图和成员推断,揭示了模型在异质推理组合中的严重脆弱性,并为未来改进提供了测试平台。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple reasoning steps. However, current compositional benchmarks testing these skills tend to focus on either commonsense or math reasoning, whereas LLM agents solving real-world tasks would require a combination of both. In this work, we introduce an Agentic Commonsense and Math benchmark (AgentCoMa), where each compositional task requires a commonsense reasoning step and a math reasoning step. We test it on 61 LLMs of different sizes, model families, and training strategies. We find that LLMs can usually solve both steps in isolation, yet their accuracy drops by nearly 30% on average when the two are combined. This is a substantially greater performance gap than the one we observe in prior compositional benchmarks that combine multiple steps of the same reasoning type. In contrast, non-expert human annotators can solve the compositional questions and the individual steps in AgentCoMa with similarly high accuracy. Furthermore, we conduct a series of interpretability studies to better understand the performance gap, examining neuron patterns, attention maps and membership inference. Our work underscores a substantial degree of model brittleness in the context of mixed-type compositional reasoning and offers a test bed for future improvement.

常识推理数学推理大模型评估组合推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。