检测加速版文生图模型的语义默认偏移,确保生成结果风格一致。
DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models

- 通过成对样本标注与概率质量迁移分析,量化语义默认变化。
- 14组对比中颜色偏差达0.054至0.303,人类评估验证排序一致性。
- 提出离线校准方法,降低偏移35.1%~10.3%,提升模型公平性与准确性。
少步文生图模型正逐步替代传统慢速生成器,但加速过程可能悄然改变未指定属性的分布,即使单个输出仍合理且与文本对齐。我们将这些分布称为语义默认,其在替换时的变化称为语义默认偏移。现有质量、偏好与多样性评估无法检验替换模型是否保持参考模型的语义默认。我们提出DefaultShift,一种成对审计方法:对重复样本进行封闭语义词汇标注,测量概率质量移动,并将可解释排序与确认性交叉拟合推断分离。在14组参考与替换模型对中,调整后颜色差异范围为0.054至0.303,方向依赖具体配方。1,000张图像的人类审计复现了该排序。我们进一步引入DefaultShift-Select,一种离线校准方法,在Turbo、DMD2和FLUX上使人类测量的偏移减少10.3%至35.1%,无明显质量损失。在均衡评估下,校准数据恢复4.3准确率点和7.5最差群体点。DefaultShift使加速下的语义保真度可测量、可行动。
原文摘要 · Abstract (English)
Few-step text-to-image models increasingly replace slower generators, yet acceleration can silently change distributions over unspecified attributes even when individual outputs remain plausible and aligned. We call these distributions semantic defaults and their change under replacement semantic default shift. Existing quality, preference, and diversity evaluations do not test whether a replacement preserves its reference model's semantic defaults. We introduce DefaultShift, a paired audit that labels repeated samples with closed semantic vocabularies, measures probability-mass movement, and separates interpretable ranking from confirmatory cross-fit inference. Across 14 reference and replacement pairs, adjusted color discrepancies range from 0.054 to 0.303 with recipe-specific directions. A 1,000-image human audit reproduces the ordering. We further introduce DefaultShift-Select, an offline calibration method that reduces human-measured shift by 10.3 percent to 35.1 percent across Turbo, DMD2, and FLUX without material quality loss. Under balanced evaluation, selected data recover 4.3 accuracy points and 7.5 worst-group points over uncalibrated replacement data. DefaultShift makes semantic preservation under acceleration measurable and actionable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。