无需标注数据,自动发现可复用的数据分析技能。
Unsupervised Skill Discovery for Agentic Data Analysis

- 通过探索轨迹自动生成验证信号,迭代优化技能。
- 报告式任务提升9.71%,推理式任务提升32.30%。
- 适合构建无需人工标注的智能数据分析代理。
推理时技能增强为数据分析师代理提供了轻量级改进方式,无需更新模型参数即可注入可复用的过程知识。然而,从无标签探索中发现有效的数据分析技能仍具挑战性,因可靠监督成本高且成功标准随分析形式变化。本文提出DataCOPE,一种无监督验证器引导的技能发现框架。该框架从探索轨迹中提取验证信号,用于刻画轨迹间的相对质量或一致性。它迭代协调数据分析师代理生成轨迹、无监督验证器提取信号、技能管理器进行对比式技能提炼。针对报告式分析,验证器实例化为自适应清单验证器,基于可验证覆盖度评分并迭代优化清单;针对推理式分析,实例化为答案一致性验证器,按答案一致性分组轨迹,并以自一致作为辅助信号。在Deep Data Research(报告式)和DABStep(推理式)上评估,跨两种场景均优于基线。四种模型设置下,报告式任务平均提升9.71%,推理式任务平均提升32.30%。
原文摘要 · Abstract (English)
Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updating model parameters. However, discovering effective skills for data analysis remains challenging, as reliable supervision is expensive and success criteria vary across analytical formats. This raises the key question of how to discover reusable data-analysis skills from unlabeled exploration alone. We propose DataCOPE, an unsupervised verifier-guided skill discovery framework for data-analytic agents. DataCOPE derives verifier signals from the exploration trajectories and uses them to characterize relative quality or aggreement among trajectories. It iteratively coordinates a Data-Analytic Agent for trajectory generation, an Unsupervised Verifier for signal extraction, and a Skill Manager for contrastive skill distillation. For report-style analysis, we instantiate the verifier as an Adaptive Checklist Verifier that derives task-specific criteria, scores reports by verifiable coverage, and iteratively refines the checklist. For reasoning-style analysis, we instantiate it as an Answer Agreement Verifier that groups trajectories by answer agreement and uses self-consistency as an auxiliary signal. We evaluate DataCOPE on report-style analysis from Deep Data Research and reasoning-style analysis from DABStep. Across both settings, DataCOPE consistently improves held-out performance over baselines. Averaged across four model settings, DataCOPE improves the mean score by 9.71% and 32.30% on report-style and reasoning-style tasks respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。