arXiv:2603.06394cs.AIcs.LG2026-03被引 1

用规则框架让AI对话自由执行确定,解决科研流程的灵活与可复现矛盾。

Talk Freely, Execute Strictly: Schema-Gated Agentic AI for Flexible and Reproducible Scientific Workflows

  • 通过结构化规则限制执行范围,确保每步操作可验证
  • 20个系统评估显示灵活与确定不可兼得,存在最优平衡区
  • 适合需要稳定科研流程的工业研发团队使用

大型语言模型(LLM)能将研究人员的自然语言目标转化为可执行计算,但科学工作流需要确定性、可追溯性和治理能力,而当LLM自主决定执行内容时难以保障。对18位来自10家工业研发机构的专家进行半结构化访谈,揭示出两个核心矛盾需求:确定性约束执行与对话灵活性不僵化,并提出人类介入控制和透明性为必须满足的边界条件。本文提出‘模式约束编排’作为解决原则:以可机器校验的模式作为组合工作流级别的强制执行边界,只有完整动作(含跨步骤依赖)通过验证才允许运行。将两种需求量化为执行确定性(ED)与对话灵活性(CF),并基于多模型协议(15次独立会话,3种LLM家族)对20个系统进行评分,结果显示ED一致性达Krippendorff α=0.80,CF达α=0.98,证明多模型评分可替代人工专家小组用于架构评估。结果揭示一个经验帕累托前沿——无系统同时具备高灵活性与高确定性,但在生成式与工作流中心范式间出现收敛区域。文章主张采用模式约束架构,分离对话权与执行权,可打破该权衡,并提炼出三项操作原则:执行前澄清、受限计划-执行编排、工具到工作流层级的准入控制。

原文摘要 · Abstract (English)

Large language models (LLMs) can now translate a researcher's plain-language goal into executable computation, yet scientific workflows demand determinism, provenance, and governance that are difficult to guarantee when an LLM decides what runs. Semi-structured interviews with 18 experts across 10 industrial R&D stakeholders surface 2 competing requirements--deterministic, constrained execution and conversational flexibility without workflow rigidity--together with boundary properties (human-in-the-loop control and transparency) that any resolution must satisfy. We propose schema-gated orchestration as the resolving principle: the schema becomes a mandatory execution boundary at the composed-workflow level, so that nothing runs unless the complete action--including cross-step dependencies--validates against a machine-checkable specification. We operationalize the 2 requirements as execution determinism (ED) and conversational flexibility (CF), and use these axes to review 20 systems spanning 5 architectural groups along a validation-scope spectrum. Scores are assigned via a multi-model protocol--15 independent sessions across 3 LLM families--yielding substantial-to-near-perfect inter-model agreement (Krippendorff a=0.80 for ED and a=0.98 for CF), demonstrating that multi-model LLM scoring can serve as a reusable alternative to human expert panels for architectural assessment. The resulting landscape reveals an empirical Pareto front--no reviewed system achieves both high flexibility and high determinism--but a convergence zone emerges between the generative and workflow-centric extremes. We argue that a schema-gated architecture, separating conversational from execution authority, is positioned to decouple this trade-off, and distill 3 operational principles--clarification-before-execution, constrained plan-act orchestration, and tool-to-workflow-level gating--to guide adoption.

AI工作流LLM科研自动化确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。