用行为克隆模拟科研标注中的专家操作,提升自动化效率。
A Systematic Study of Behavioral Cloning for Scientific Data Annotation

- 构建9个仿真任务,还原人类标注时的探索与纠错策略。
- 大模型在多任务训练下更省数据,且能自动修正错误。
- 发现跨任务通用的错误表征,适合迁移到新标注任务。
科学数据标注(如视频追踪动物或神经重构校对)仍受制于‘最后一公里’难题:尽管自动化水平高,验证与修正仍需大量人力。现有方法直接预测标注结果,忽略了专家操作过程中的丰富监督信息。本文提出一个行为克隆研究框架,包含9个合成任务及对应合成标注,模拟真实人类的探索、纠错与策略决策。实验发现:1)技能呈层次化涌现,模型先掌握界面操作,再掌握关键决策,出错率低于训练数据,同时保留纠错能力;2)在多任务行为克隆中,更大模型更具数据效率;3)多任务预训练可实现高效微调,而从头训练完全失败;4)线性探针显示模型内隐表示了任务阶段与数据位置等潜在变量,且存在跨任务通用的错误表征。该框架建立系统基准,揭示关键瓶颈,为真实科研标注中的行为克隆扩展提供基础。
原文摘要 · Abstract (English)
Scientific data annotation, such as tracking animals in video or proofreading neural reconstructions, remains bottlenecked by the "last mile" problem: even with strong automation, verification and correction consume substantial human effort. Standard approaches train models to directly predict annotations, discarding the rich supervision in how experts navigate, click, verify, and correct. We introduce a framework for studying behavioral cloning on scientific annotation: 9 synthetic tasks paired with synthetic annotations that simulate realistic human strategies including exploration, mistake correction, and strategic decision-making. Our experiments reveal several findings. First, skills emerge hierarchically: models learn GUI mechanics before task-critical decisions, and commit fewer mistakes than the training data while retaining the ability to correct errors when they occur. Second, scaling models on multi-task behavioral cloning shows that larger models are more data efficient within our scale range. Third, multi-task pretraining enables efficient fine-tuning to new tasks, while training from scratch fails entirely. Fourth, linear probes reveal that models internally represent latent variables of the annotation process such as task phase and data position; interestingly, we find a shared mistake representation that generalizes across different annotation tasks. Overall, our framework establishes systematic benchmarks and identifies key bottlenecks, providing a foundation for scaling behavioral cloning to real-world scientific data annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。