用实验设计方法诊断视觉模型失败原因,精准生成补救数据。
Synthetic Designed Experiments for Diagnosing Vision Model Failure

- 基于实验设计理论,系统分析模型对场景因素的敏感性。
- 发现两类失败:数据覆盖不足和依赖虚假关联,针对性补数据提升准确率。
- 适合研究模型鲁棒性与合成数据优化的研究者使用。
当前计算机视觉的合成数据生成流程缺乏对下游模型真实需求的诊断。本文提出合成实验设计(SDRS),将模型视为黑箱,合成生成器作为实验装置,利用部分因子设计通过方差分析(ANOVA)高效审计模型对各场景因素的敏感性。该方法将失败分为两类:类型I(低频因子水平覆盖不足)和类型II(依赖虚假干扰因素)。审计结果可指导生成针对性合成数据以修复缺陷。在三个实验中验证:(1)在带人为偏差的dSprites上,准确率从49.9%提升至79.0%;(2)在程序化场景分割任务中,mIoU从0.948升至0.998;(3)检测到生成器中跨因子混淆问题。此外,发现因子不变性惩罚可转移敏感性,揭示表示层面修正的开放问题。
原文摘要 · Abstract (English)
Current synthetic data pipelines for computer vision generate images without diagnosing what the downstream model actually needs. This open-loop paradigm treats synthetic data as cheap real data, randomly sampling the generator's output space and hoping to cover the model's failure modes. We argue this fundamentally misuses synthetic data's unique property: the controllable, independent variation of scene factors.Drawing on the statistical theory of Design of Experiments (DoE), we propose Synthetic Designed Experiments for Representational Sufficiency (SDRS). SDRS treats the downstream model as a black-box system and the synthetic generator as an experimental apparatus. Using fractional factorial designs, SDRS efficiently audits a model's factor-sensitivity profile via ANOVA decomposition. It classifies failures into two actionable types: Type I gaps (coverage failures on underrepresented factor levels) and Type II gaps (reliance on spurious nuisance dependencies). The audit then prescribes targeted synthetic data to address each gap type. We validate SDRS on three experiments: (1) a controlled diagnostic on dSprites with planted biases, where the audit correctly identifies both gap types and targeted data improves accuracy from 49.9% to 79.0%; (2) a dense segmentation task on procedural scenes, where detecting background-complexity shortcuts and applying targeted data improves mIoU from 0.948 to 0.998; and (3) an entanglement detection experiment showing that the ANOVA audit identifies cross-factor contamination in imperfect generators. Finally, we show that per-factor invariance penalties can transfer sensitivity between factors, identifying an open problem for representation-level correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。