测试图像编辑器能否从单张照片生成物理地图,提出新评估标准。
PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors

- 用固定交互协议评估编辑器在五类物理图上的表现
- 专业模型仍优于编辑器,但部分指标上编辑器可匹敌
- 针对光照、结构等设计了压力测试数据集
通用图像编辑器能否从单张RGB图像预测物理地图?与专用密集预测模型不同,编辑器需通过提示、示例或图文线索引导。为此,我们提出PhysEditBench,一个协议约束的基准,用于评估图像编辑器在深度、法线、反照率、粗糙度和金属度五类物理图上的表现。数据来自OpenRooms-FF、InteriorVerse及新生成的程序化金属度数据集,经质量检查、有效区域掩码、场景级采样和基于光照的压力子集处理,确保评估可靠多样。每个目标定义固定输入、输出格式与评分流程,分数反映特定协议下的性能,而非最佳可能表现。实验显示,专用模型在深度、法线、反照率上仍占优;优秀编辑器可生成更合理的图谱输出。在粗糙度和金属度上,编辑器部分指标可达或超越基线,但仍存在结构错误、稀疏性问题和对光照敏感等缺陷。
原文摘要 · Abstract (English)
Can general-purpose image editors predict physical maps from a single RGB image? General-purpose image editors differ from standard task-specific dense-prediction models: they do not directly take an image and output a physical map. Instead, they must be guided by prompts, examples, or image-based textual cues. To this end, we introduce PhysEditBench, a novel protocol-conditioned benchmark to evaluate and standardize image editors in dense physical-map prediction that covers five targets: depth, normal, albedo, roughness, and metallic maps. For evaluation data, we build a target-dependent benchmark substrate. We use OpenRooms-FF for depth, surface normal, albedo, and roughness, InteriorVerse as an additional source for depth, normal, albedo, and a new procedurally generated source for metallic maps. We curate the data with quality checks, valid-region masks, scene-level sampling, and lighting-based stress subsets to ensure reliable and diverse evaluation. For each target, PhysEditBench defines a fixed protocol that specifies the allowed input, expected output format, and scoring procedure. Each score, therefore, reflects the performance of a model under a specified protocol, rather than its best possible performance under all prompts or interaction modes. Experimental results show that specialized models remain much stronger on depth, normal, and albedo, and stronger image editors can produce more reasonable map-like outputs. For roughness and metallic, image editors can match or outperform specialized baselines on some scalar metrics, but they still suffer from structural errors, sparsity effects, and sensitivity to lighting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。