arXiv:2608.05235cs.IRcs.AI2026-08

让实验轨迹变成可审计的证据,提升工业研究可靠性

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

  • 构建生成-验证-修复闭环,确保成果可追溯、无漏洞
  • 实测显示最终结果常不如早期最优,揭示轨迹非单调演化
  • 适合需合规审计的工业级机器学习研究团队使用

研究代理在工业推荐场景中开展多轮机器学习实验,保留轨迹以指导后续决策。然而,完成的轨迹不等于有效证据:生成结果可能缺乏支撑或不完整,执行回合可能无效或受干扰,后期修改可能掩盖前期发现。本文研究「轨迹转证据」问题,提出基于证据的框架,结合对关键成果的边界验证与事后主张定性。通过上下文隔离的生成-验证-修复流程,在发布前检查成果是否违反证据规则或缺少下游需求。执行后,通过有效性与归属性检查,整合多轮证据,将干预主张分类为可操作修复、诊断保护或暂存发现,并以明确来源和适用范围保存可接受主张作为可审计记录。混合式大模型辅助控制器根据目标证据选择采纳、延迟或拒绝记录。记录审计揭示哪些主张经受住检验,下游诊断则指出正面适用性判断是控制器的瓶颈。跨论文到目标的适配中,后期回合常优于首轮,但最终回合常弱于早期最佳,暴露轨迹演化的非单调性。经完整工作流产出的候选方案相较部署基线带来正向线上提升。

原文摘要 · Abstract (English)

Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions. Yet a completed trajectory is not automatically evidence: generated artifacts may be unsupported or incomplete, executed rounds may be invalid or confounded, and later modifications may obscure earlier findings. We study \textbf{trajectory-to-evidence conversion}, asking what a completed research process has actually established. We introduce an evidence-grounded framework that couples bounded verification of consequential artifacts with post-execution claim qualification. A context-isolated generate--verify--repair process checks artifacts for evidence violations and missing downstream requirements before release. After execution, validity and attribution checks consolidate evidence across rounds, qualify intervention-level claims as actionable repairs, diagnostic guards, or withheld findings, and preserve admitted claims as auditable records with explicit provenance and applicability boundaries. A hybrid LLM-assisted controller subsequently applies, defers, or rejects records based on available target evidence. Record audits characterize which claims survive qualification, while downstream diagnostics identify affirmative applicability judgment as a bottleneck for the tested controller. Across paper-to-target adaptations, later rounds often improve on the first, while final rounds frequently underperform an earlier best, exposing non-monotonic trajectory evolution. Candidates produced through the complete workflow also yielded positive online lifts relative to deployed baselines.

实验审计研究代理证据可信工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。