首个无参考的材质替换评测基准,解决室内设计中材质更换的公平评估难题。
MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes

- 构建无参考评测框架,按四个维度评估材质替换效果。
- 指令驱动下顶尖闭源模型表现超越参照物,但图像参考反而降低性能。
- 专家评分验证结果可信,适合研究视觉材质生成与编辑的学者使用。
材质替换是常见的室内设计操作:在保持几何、环境和光照不变的前提下,更换指定表面的材质。尽管具有商业价值,目前尚无公开基准能独立评估此任务。基于参考的评价指标在此类一对多任务中存在缺陷,会惩罚有效输出,偏袒参考生成器风格,无法公平比较接收不同引导形式的编辑器。本文提出 MatReplace,一个无参考的评测基准,从局部材质正确性、全局光照协调性、外部场景保留和内部结构一致性四个可验证维度评估编辑效果。定义了三个测试路径:(A) 仅指令,(B) 指令加区域掩码,(C) 用材质参考图代替指令。结果表明,命名材质渲染已由最强闭源模型基本解决,但基于像素的材质定位仍是开放挑战。在路径A中,领先闭源模型在主聚合指标上超越参照锚点;路径B中,掩码仅帮助掩码兼容模型,单种子实验效果波动在+0.137至-0.090之间;路径C中,参考图条件使所有模型性能下降,降幅达-0.031至-0.508,最差情况甚至将参考图重绘,表现劣于原图。专家评分验证排名可靠性(Kendall's tau = 0.68),优于基于真实标注或CLIP的基线。
原文摘要 · Abstract (English)
Material replacement is a common interior-design operation: changing the material of a selected surface while preserving its geometry, surroundings, and illumination. Despite its commercial relevance, no public benchmark isolates this task, and evaluating it is challenging. Reference-based metrics penalize valid outputs in this inherently one-to-many setting, favor the style of the reference generator, and cannot fairly compare editors that receive different forms of guidance. We introduce MatReplace, a reference-free benchmark that evaluates edits along four verifiable dimensions: local material correctness, global lighting harmony, outside preservation, and inside structure. It defines three tracks that vary one conditioning signal at a time: (A) instruction only, (B) instruction plus region mask, and (C) material reference image instead of instruction. Our results reveal a clear divide between naming and visually grounding materials. In Track A, leading closed-source editors achieve exemplar-level material rendering and surpass the exemplar anchor under our primary aggregate. In Track B, masks help only mask-compatible models with weak scene preservation, with task-paired, single-seed effects ranging from +0.137 to -0.090 across aligned model families. In Track C, reference-image conditioning degrades every family under both aggregates, by -0.031 to -0.508; in the worst cases, models repaint the reference image itself and perform worse than returning the input unchanged. Thus, named-material rendering is largely solved by the strongest closed editors on this distribution, but grounding materials from pixels remains an open challenge. Expert ratings validate our ranking (Kendall's tau = 0.68) and align with our aggregates more closely than GT-referenced or CLIP-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。