优化数据代理在分支湖仓中的技能,提升代码生成效果。
"Skill Issues'': Data-Centric Optimization of Lakehouse Agents

- 以湖仓状态变化为验证依据,构建数据驱动的技能优化流程。
- 在数百个任务上,优化后技能使奖励提升最高达28.6%。
- 适合关注数据代理训练与评估的开发者与研究者。
编码代理正成为数据基础设施的使用者,其成功不仅依赖模型质量,还取决于教会代理使用系统的技能和环境文件。本文研究如何在分支湖仓系统Bauplan中优化这些产物。在该系统中,无头API和类Git的数据原语通过代码、分支、提交和合并暴露数据工作流。核心观察是:分支湖仓将代理评估从输出匹配问题转变为状态验证问题——代理生成的流水线代码会引发可检查的湖仓变更。我们提出一种数据中心化的优化流程,生成任务-验证器对,在隔离沙箱中执行候选技能,并基于追踪信号与湖仓状态的程序化检查评分轨迹。在数百个任务的初步评估中,优化后的技能使保留奖励提升最高达28.6%。结果表明,写路径数据工作流为超越只读任务的代理技能优化提供了有效基础。
原文摘要 · Abstract (English)
Coding agents are becoming users of data infrastructure, but their success depends not only on model quality: it also depends on the skills and environment files that teach agents how to use a system. We study how to optimize these artifacts for agents operating on a branching lakehouse, Bauplan. In our setting, headless APIs and Git-like data primitives expose data workflows through code, branches, commits, and merges. Our central observation is that a branching lakehouse turns data-agent evaluation from an output-matching problem into a state-verification problem: agent-generated pipeline code induces concrete, inspectable lakehouse changes. We present a data-centric optimization pipeline that generates task-verifier pairs, executes candidate skills in isolated sandboxes, and scores trajectories using both trace-level signals and programmatic checks over lakehouse state. In a preliminary evaluation on hundreds of tasks, optimized skills improve held-out reward by up to 28.6%. These results suggest that write-path data workflows provide a useful substrate for optimizing agent skills beyond read-only tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。