研究人机协作如何提升物理理论推理,发现反馈策略效果取决于角色搭配。
When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

- 构建评审-执行循环框架,让AI分步迭代改进物理问题解答。
- 轻量级执行者配合强模型评审时,建设性反馈能显著提升解题得分。
- 模型规模增大对难题帮助有限,反馈策略在异构搭配中更关键。
随着大语言模型在科研级物理推理任务中表现日益突出,代理型AI应用愈发普遍,一个核心问题是:研究人员与代理之间的交互如何影响结果?本文通过SCALAR(用于智能体推理的结构化评审-执行循环)框架,研究量子场论与弦论问题中的互动机制。该框架包含执行者提出方案、评审者提供迭代反馈、独立裁判依据参考答案评估对话内容。我们调整执行者身份、评审反馈策略以及模型家族与规模。多轮对话始终优于单次尝试,但改进机制与提示策略的价值高度依赖于执行者-评审者配对。在同一模型家族内增加规模(如从80亿参数的DeepSeek-R1升级至700亿参数的DeepSeek-R1)可改善部分简单问题表现,但无法突破最困难瓶颈。在异构配对中(如轻量级Haiku执行者由更强的Sonnet评审者指导),建设性反馈明显提升平均得分;同家族情况下,策略影响较弱,宽松反馈有时更优,而严格或对抗式反馈无益。总体而言,SCALAR为评估科学问题中人机协作结构的有效性提供了可控实验平台。
原文摘要 · Abstract (English)
As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results? We study this using SCALAR (Structured Critic--Actor Loop for Agentic Reasoning), an Actor--Critic--Judge pipeline applied to quantum field theory and string theory problems. The Actor proposes solutions, the Critic provides iterative feedback, and an independent Judge evaluates the transcript against reference solutions. We vary the Actor persona, the Critic feedback strategy, and the Actor model family and scale. Multi-turn dialogue improves over single-shot attempts throughout, but both the mechanism of improvement and the value of different prompting choices depend strongly on the Actor--Critic pairing. Increasing the scale within one model family (e.g. from the 8B-parameter DeepSeek-R1 variant to DeepSeek-R1 70B) improves some easier-problem behavior, but does not remove the hardest bottleneck we observe. Critic feedback strategy matters most clearly in the asymmetric Actor--Critic setting (e.g, a lightweight Haiku Actor guided by a stronger Sonnet Critic), where constructive feedback improves mean-score outcomes. In same-family Actor--Critic settings, strategy effects are weaker: lenient feedback is sometimes favored, while strict and adversarial feedback are not beneficial. Taken together, SCALAR provides a controlled testbed for evaluating which interaction structures help or hinder AI-assisted reasoning on known-answer scientific problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。