通过原子级探针实现模块更新后的高效回归测试,降低验证成本。
Regression Test Selection for Updated Capability Modules in Compositional ML Systems via Atomic-Quality Probes
- 用跨版本替换法分离单个模块更新的独立影响。
- 发现一个主导模块可使组合成功率提升52个百分点,相关性达0.94。
- 新方法在半成本下匹配全量验证效果,适合快速迭代的系统集成。
组合式机器学习系统由独立重训练的能力模块动态组装而成。更换任一模块引发回归测试难题:静态依赖分析无法判断哪些现有组合仍有效,且测试成本高昂。本文将模块更新建模为回归测试选择(RTS),提出四项成果:第一,设计配对跨版本替换协议,精准隔离单模块更新的影响;第二,在两个高交互操作任务中发现主导技能效应——单一模块原子级成功率达88.0%,其他模块不超过32.0%,其引入使组合成功率最高提升52个百分点;通过权重空间插值验证,组合成功率与原子质量呈高度线性关系(合并皮尔逊相关系数r=0.94),该现象在第二个任务中再次复现,且主导模块位于相位序列关键路径上;第三,基于策略外行为距离的度量方法无法识别主导模块;第四,提出的边缘门控混合选择器在零决策测试成本下达到与全量重验证相同的75.0%黄金标签一致率(无显著差异),在半成本下达成81.25%匹配率,优于同成本随机预算(蒙特卡洛p=0.039)。分辨率分析显示粗粒度评估高估了全量验证的优势。原子质量探针为组合式ML系统的模块更新回归测试提供了原则性筛选准则。
原文摘要 · Abstract (English)
Compositional machine-learning (ML) systems assemble runtime behavior from libraries of independently re-trained capability modules. Replacing one module raises a regression-testing question that static dependence analysis cannot answer: which existing compositions stay valid, and at what test cost? We frame capability updates as regression test selection (RTS) and contribute four results. First, a paired cross-version swap protocol isolates the marginal effect of a single module update. Second, on two contact-rich manipulation tasks we characterize a dominant-skill effect: one capability module reaches 88.0% atomic success while siblings stay at or below 32.0%, and its inclusion shifts composition success by up to 52 percentage points; a controlled weight-space interpolation tracks composition success against atomic quality point-by-point (pooled Pearson r=0.94), and the effect replicates on a second task, where the governing module must lie on the critical path of the phase sequence. Third, off-policy behavioral-distance metrics fail to identify the dominant module. Fourth, a margin-gated Hybrid Selector matches full revalidation at zero per-decision test cost (75.0% gold-label agreement, with no detectable difference) and reaches 81.25% match at half of full-revalidation cost, beating a cost-matched random budget (Monte-Carlo p=0.039). A resolution analysis shows that coarse evaluation overstates the apparent advantage of full revalidation. The atomic-quality probe gives a principled test-selection criterion for capability-update regression testing in compositional ML systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。