arXiv:2601.09274cs.AI2026-01被引 1

评测科学推理中的记忆激活机制,揭示知识调用对推理稳定性的影响。

$A^3$-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation

  • 基于锚点与吸引子的双尺度记忆激活框架
  • 构建2198个跨领域科学问题数据集,标注知识激活路径
  • 提出AAUI指标量化记忆使用率,适合研究认知推理模型者

科学推理不仅依赖逻辑推导,更需激活已有知识与经验结构。记忆能高效复用知识,提升推理的一致性与稳定性。然而现有基准多聚焦最终答案或步骤连贯性,忽视人类推理中关键的「记忆驱动」机制——即通过锚点与吸引子激活,并融入多步推理。为此,我们提出A³-Bench(https://a3-bench.github.io),一个基于锚点与吸引子激活(Anchor and Attractor Activation)的科学推理评测基准。首先,采用SAPM流程(主题、锚点与吸引子、问题、记忆构建)标注2,198个跨领域的科学推理问题;其次,设计双尺度记忆评估框架,引入AAUI(锚点-吸引子利用率指数)衡量记忆激活率;最后,通过多种基础模型与范式实验,验证A³-Bench的有效性,并分析记忆激活对推理性能的影响,为理解记忆驱动的科学推理提供洞见。

原文摘要 · Abstract (English)

Scientific reasoning relies not only on logical inference but also on activating prior knowledge and experiential structures. Memory can efficiently reuse knowledge and enhance reasoning consistency and stability. However, existing benchmarks mainly evaluate final answers or step-by-step coherence, overlooking the \textit{memory-driven} mechanisms that underlie human reasoning, which involves activating anchors and attractors, then integrating them into multi-step inference. To address this gap, we propose $A^3$-Bench~ https://a3-bench.github.io, a benchmark designed to evaluate scientific reasoning through dual-scale memory-driven activation, grounded in Anchor and Attractor Activation. First, we annotate 2,198 science reasoning problems across domains using the SAPM process(subject, anchor & attractor, problem, and memory developing). Second, we introduce a dual-scale memory evaluation framework utilizing anchors and attractors, along with the AAUI(Anchor--Attractor Utilization Index) metric to measure memory activation rates. Finally, through experiments with various base models and paradigms, we validate $A^3$-Bench and analyze how memory activation impacts reasoning performance, providing insights into memory-driven scientific reasoning.

科学推理记忆机制评测基准认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。