测试分阶段筛选无效,全局统一指标反而更优。
Phase-Localized Curation Does Not Help: A Negative Result on Per-Phase Metric Selection for Demonstration Filtering
- 将演示按阶段分段评分再聚合,试图提升筛选效果
- 三任务中分阶段策略均非最优,甚至两任务表现最差
- 缺陷信号集中时,跨阶段平均会稀释有效信息
操纵演示具有时间阶段结构,一种自然假设是筛选指标应分阶段应用而非全局使用。我们将在三个接触密集的LIBERO抓取放置任务上测试该假设,引入可控的提前释放结构缺陷,比较分阶段筛选与统一指标及强全局指标的表现。在所有三任务和每条件五组随机种子下,分阶段筛选从未成为最佳策略,且在两个任务中表现最差(任务1:86.0 vs. 92.0 全局;任务3:22.7 vs. 48.0 统一)。失败原因在于:当缺陷信号集中在单一阶段时,跨阶段排序聚合会稀释其信号,导致选择劣质演示子集。此外,分阶段指标选择不具跨任务泛化性,任意两任务间无共享最优指标,因此必须针对每任务进行噪声较大的重新搜索。这些结果限制了一种此前未验证的合理方法,建议实践者优先寻找单一缺陷敏感指标,而非按阶段分解筛选。我们已发布完整流程、所有指标实现及每种子结果。
原文摘要 · Abstract (English)
Manipulation demonstrations have temporal phase structure, and a natural hypothesis is that demonstration-curation metrics should be applied within phases rather than globally. The idea is to segment each trajectory into phases, score each phase with the metric that is locally most informative, and then aggregate. This follows directly from prior work showing that a single global metric can be the best detector of a defect and yet the worst curator of the resulting policy. We test the per-phase hypothesis on three contact-rich LIBERO pick-and-place tasks with a controlled early-release structural defect, comparing phase-gated curation against the same metrics applied uniformly and against a strong single global metric. Across all three tasks and five random seeds per condition, phase-gated curation is never the best curation strategy, and it is the worst of the three on two of the three tasks (Task 1: 86.0 vs. 92.0 for global; Task 3: 22.7 vs. 48.0 for uniform). We trace the failure to a concrete mechanism. When the defect signal is concentrated in a single phase, rank-aggregating across phases dilutes that signal with uninformative scores from defect-free phases, selecting a worse demonstration subset than simply applying the defect-informative metric everywhere. We further show that the per-phase metric selection does not transfer across tasks, since no phase shares a winning metric between any two tasks, so the selection cannot be reused and must be re-derived per task from a noisy sweep. These results bound a plausible and previously untested method, and they argue that practitioners should prefer identifying a single defect-informative metric over decomposing curation by phase. We release the full pipeline, all metric implementations, and per-seed results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。