arXiv:2605.17746cs.AIcs.HC2026-05

用结构化图谱设计人机协作实验,让AI实验更可比可审计。

Agents for Experiments, Experiments for Agents: A Design Grammar for AI-Enabled Experimental Science

  • 将实验条件建模为带类型的交互图谱,支持流程可视化
  • 在医疗分诊任务中,SEED生成的设计更清晰体现角色分工与治理规则
  • 适合研究人机协作、AI工作流设计的学者与工程师

AI系统正深度参与组织与知识工作,与人类协作、协调流程,并以多智能体形式运行。理解其影响需超越输出准确率,还需考察机制、委托、反馈与控制。实验仍是核心手段,但面临循环挑战:我们需要实验来研究智能体协作,又需智能体帮助探索海量设计空间。当前的人机与智能体工作流仍以自然语言描述,难以比较、复用或审计。本文提出SEED(结构化实验发现编码),将实验条件表示为类型化的参与者-流程图谱,支持三类功能:描述交互结构、评估新设计相对于已有设计的创新性、在可行性与治理约束下生成候选方案。通过轻量级实证测试,对比了无图谱与基于SEED的生成在医疗分诊任务中的表现。结果表明,SEED引导的设计在角色流程变化、假设显化和治理检查方面更清晰,验证了该框架作为设计辅助工具的可行性。文章最后指出新颖性、可复现性、有效性、探究多样性与问责之间的治理张力。

原文摘要 · Abstract (English)

AI systems are becoming active participants in organizational and knowledge work. They increasingly interact with humans, coordinate workflows, and operate in multi-agent arrangements. Understanding their effects therefore requires more than measuring output accuracy; it requires evidence about mechanisms, delegation, feedback, and control. Experiments remain central to this task, but they also face a recursive challenge: we need experiments for agents to study these arrangements, and we may need agents for experiments to help search the expanding space of possible designs. Yet experimental conditions for human-AI and agentic workflows are still largely specified in prose, making them difficult to compare, reuse, or audit. We frame this as a problem of workflow representation, traceability, and governance in AI-enabled knowledge production. We introduce SEED (Structural Encoding for Experimental Discovery), a framework that represents experimental conditions as typed actor-flow graphs. SEED supports three design functions: describing conditions as interaction structures, evaluating structural novelty relative to encoded prior designs, and generating candidate designs under feasibility and governance constraints. We report a lightweight empirical feasibility test that compares graph-blind and SEEDguided generation in a medical-triage design task. In this diagnostic contrast, SEED-guided candidate designs show clearer actor-flow changes, assumptions, and governance checks, supporting the feasibility of the grammar as a design aid. The commentary closes by identifying governance tensions around novelty, replication, validity, diversity of inquiry, and accountability.

人机协作实验设计智能体结构建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。