提出可评估图像生成空间指令遵循的不确定性感知基准
SpatialBench-UC: Uncertainty-Aware Evaluation of Spatial Prompt Following in Text-to-Image Generation
- 设计包含200个提示对的可复现基准,支持空间关系判断
- 引入置信度与放弃决策机制,提升评估结果可解释性
- 适合研究视觉语言模型空间对齐的开发者与评测者
自动评估文本到图像模型是否遵循明确的空间指令仍具挑战。目标检测器可能遗漏目标或产生多个合理检测结果,而简单的几何测试在边界情况下易产生歧义。空间评估本质上是选择性预测问题:当证据不足时,评估器可选择不作答,并报告置信度,使结果体现风险与覆盖的权衡,而非单一分数。我们提出SpatialBench-UC,一个小型、可复现的空间关系评估基准,包含200个提示(50个物体对 × 4种关系),分为100组通过交换物体角色构建的反事实对。我们发布了基准包、版本化提示、固定配置、每样本检查器输出及报告表格,支持跨模型的可复现与可审计比较。还包含轻量级人工审核,用于校准检查器的放弃阈值与置信度阈值。我们评估了三个基线模型:Stable Diffusion 1.5、SD 1.5 BoxDiff 和 SD 1.4 GLIGEN。检查器报告通过率、覆盖率以及在已决定样本上的条件通过率。结果表明,接地方法显著提升通过率与覆盖率,但因漏检问题,放弃决策仍是主要影响因素。
原文摘要 · Abstract (English)
Evaluating whether text-to-image models follow explicit spatial instructions is difficult to automate. Object detectors may miss targets or return multiple plausible detections, and simple geometric tests can become ambiguous in borderline cases. Spatial evaluation is naturally a selective prediction problem, the checker may abstain when evidence is weak and report confidence so that results can be interpreted as a risk coverage tradeoff rather than a single score. We introduce SpatialBench-UC, a small, reproducible benchmark for pairwise spatial relations. The benchmark contains 200 prompts (50 object pairs times 4 relations) grouped into 100 counterfactual pairs obtained by swapping object roles. We release a benchmark package, versioned prompts, pinned configs, per-sample checker outputs, and report tables, enabling reproducible and auditable comparisons across models. We also include a lightweight human audit used to calibrate the checker's abstention margin and confidence threshold. We evaluate three baselines, Stable Diffusion 1.5, SD 1.5 BoxDiff, and SD 1.4 GLIGEN. The checker reports pass rate and coverage as well as conditional pass rates on decided samples. The results show that grounding methods substantially improve both pass rate and coverage, while abstention remains a dominant factor due mainly to missing detections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。