为百亿参数机器人多任务模型设计高效影响函数,加速数据清洗。
ATHENA: Accelerated Multi-Task Heterogeneous Influence Functions for Robot Data Curation

- 利用梯度的克罗内克结构与低秩近似,降低计算开销。
- 仅用一半模拟数据或三分之二真实数据,效果媲美全量训练。
- 适合大规模多任务机器人学习的数据筛选与优化场景。
在机器人模仿学习中,影响函数可量化每条示范对任务结果的影响,但扩展到百亿参数的视觉-语言-动作(VLA)模型时面临计算和多任务瓶颈。为此,我们提出针对百亿参数多任务VLA数据清洗的ATHENA框架。具体而言,它利用线性层梯度的克罗内克结构降低投影成本,并通过秩-r随机截断近似实现密集海森矩阵求逆,使影响计算速度提升约313.4倍。此外,ATHENA提出全局与局部交互影响机制,平衡50个联合训练任务间的数据清洗。在RoboTwin 2.0和真实机器人部署上的大量评估显示,模拟场景仅需50%的示范数据(9.34小时),真实任务仅需66.7%的数据(6.90小时),效果即匹配或超过全量数据联合微调。整体表明,ATHENA在百亿参数多任务VLA微调中的数据清洗中具有显著有效性。
原文摘要 · Abstract (English)
In robot imitation learning, influence functions provide a principled approach to quantify each demonstration's effect on robot task outcomes, yet scaling them to billion-parameter Vision-Language-Action (VLA) models is limited by computational and multitask bottlenecks. To this end, we propose ATHENA, an influence function framework tailored for multitask VLA data curation at a billion-parameter scale. Concretely, it leverages the Kronecker structure of linear-layer gradients to reduce projection cost, and approximates dense Hessian inversion with a rank-r Random Truncated Approximation, achieving about a 313.4x speedup in influence computation. Furthermore, ATHENA formulates global and local interactive influence to balance data curation across 50 jointly trained tasks. Extensive evaluations on RoboTwin 2.0 and real-robot deployment, covering 9.34 and 6.90 hours of demonstrations, respectively, show that ATHENA matches or exceeds full-data joint fine-tuning using only 50% of demonstrations in simulation and 66.7% of data across six real-robot tasks. Overall, ATHENA demonstrates its effectiveness for data curation in billion-parameter multitask VLA fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。