让专家、开发者和AI助手协同设计可验证的科学领域AI代理。
Collaborative Agent Reasoning Engineering (CARE): A Three-Party Design Methodology for Systematically Engineering AI Agents with Subject Matter Experts, Developers, and Helper Agents

- 三角色协作:专家提需求,开发者写代码,AI助手转译成可审查规范。
- 分阶段生成交互要求、推理策略等具体成果,提升开发效率与查询准确率。
- 适合需要严谨验证的科研场景,尤其适合非专家主导的复杂任务设计。
我们提出协同代理推理工程(CARE),一种在科学领域系统化构建大语言模型(LLM)代理的规范化方法。不同于随意试错的方式,CARE通过可复用的产物和分阶段、有门禁的流程,明确指定行为、知识约束、工具调度与验证方式。该方法采用三方协作流程:领域专家(SMEs)、开发者与基于LLM的辅助代理。这些辅助代理作为协作基础设施,将非正式的领域意图转化为结构化、可审查的规范,供人类在关键节点审核。CARE应对“技术前沿不均衡”问题,弥合新手与专家在领域约束和验证实践上的差距。通过生成具体产物,如交互需求、推理策略与评估标准,确保代理行为具备可说明性、可测试性和可维护性。在一项科学应用场景中的评估显示,这种分阶段、产物驱动的方法显著提升了开发效率与复杂查询性能。
原文摘要 · Abstract (English)
We present Collaborative Agent Reasoning Engineering (CARE), a disciplined methodology for engineering Large Language Model (LLM) agents in scientific domains. Unlike ad-hoc trial-and-error approaches, CARE specifies behavior, grounding, tool orchestration, and verification through reusable artifacts and systematic, stage-gated phases. The methodology employs a three-party workflow involving Subject-Matter Experts (SMEs), developers, and LLM-based helper agents. These helper agents function as facilitation infrastructure, transforming informal domain intent into structured, reviewable specifications for human approval at defined gates. CARE addresses the "jagged technological frontier", characterized by uneven LLM performance, by bridging the gap between novice and expert analysts regarding domain constraints and verification practices. By generating concrete artifacts, including interaction requirements, reasoning policies, and evaluation criteria, CARE ensures agent behavior is specifiable, testable, and maintainable. Evaluation results from a scientific use case demonstrate that this stage-gated, artifact-driven methodology yields measurable improvements in development efficiency and complex-query performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。