为文本生成系统设计新校准方法,确保输出风险可控。
Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation

- 通过重加权提示和候选选择,修正搜索过程中的风险偏差
- 在GenEval2上将释放图像风险从0.310降至0.162
- 适合关注生成安全性的部署型模型开发者
验证器引导的文生图系统在推理时采用测试时搜索来选择、优化或终止多个候选图像,但现有释放阈值通常基于单张图像校准。这导致候选级与策略级之间的校准错位:搜索改变了哪些提示获得输出以及最终释放哪个候选,因此候选级风险控制未必能保证释放输出的风险可控。本文通过提示重加权与提示内选择的形式化分析,提出SHIP(Selection-aware Held-out calibration of Inference Policies)方法。SHIP在保留提示上运行或回放完整部署策略,用独立目标判别器评估实际释放的图像,并选取最宽松的阈值,使其风险上界满足预设预算。对于可回放策略及预设阈值网格,同时置信控制提供有限样本有效性。在固定、序列与自适应三种文生图推理流程中实验表明,策略级校准可恢复更低风险的运行点,并揭示风险、覆盖率与计算成本间的策略依赖权衡。在GenEval2上使用FLUX模型且N=16时,基于候选池的阈值导致释放风险为0.310,而SHIP将其降低至0.162;在200个缓存流分割中,固定网格证书未出现目标越界。可靠推理缩放要求对部署策略产生的输出分布进行校准。
原文摘要 · Abstract (English)
Verifier-guided text-to-image systems increasingly use test-time search to select, refine, or stop among multiple candidates, yet release thresholds are often calibrated on individual images. This creates a candidate-to-policy calibration mismatch: search changes both which prompts receive an output and which candidate is released, so candidate-level risk control need not imply control of released-output risk. We formalize this estimand shift through prompt reweighting and within-prompt selection, and introduce SHIP, Selection-aware Held-out calibration of Inference Policies. SHIP runs or replays the complete deployed policy on held-out prompts, evaluates the image it actually releases using an independent target judge, and selects the most permissive threshold whose risk upper bound satisfies a prescribed budget. For replayable policies with a prespecified threshold grid, simultaneous confidence control provides finite-sample validity. Experiments across fixed, sequential, and adaptive T2I inference procedures show that policy-level calibration recovers lower-risk operating points while exposing policy-dependent tradeoffs among risk, coverage, and compute. On GenEval2 with FLUX at N=16, a pooled-candidate threshold yields released risk 0.310, whereas SHIP reduces it to 0.162. Across 200 cached-stream splits, the fixed-grid certificate has no target crossing. Reliable inference-time scaling therefore requires calibrating the output distribution induced by the complete deployed policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。