针对隆鼻生成,提出新评估方法与管线,解决局部编辑评价失准问题。
Envisage: Diffusion-Based Rhinoplasty Goal Visualization with Mask-Decomposed Evaluation

- 用掩码分解策略评估局部编辑效果,避免全脸身份指标误导。
- 在211个样本上,新方法在关键指标上领先,但全脸相似度仍为负增长。
- 适用于医美生成、算法评估研究者,推动局部编辑评价标准化。
局部生成编辑需要局部化评估:全图身份度量在硬掩码合成下结构上存在偏差。本文提出Envisage,一个基于FLUX.1-Fill的隆鼻目标可视化参考流程,仅需一张正面照片即可生成。该流程结合8种临床隆鼻预设(另含8种眼整形和8种除皱预设)、MediaPipe掩码及硬掩码合成。由于合成保留掩码外像素,全脸身份分数主要受复制像素主导,而非扩散模型本身。因此,全脸身份指标无法有效评估局部编辑。为此,我们提出SurgicalScore,一种0-1评分协议,分解评估编辑方向、幅度、掩码内LPIPS、真实感与掩码外保持性;其原始得分在理想控制组中达0.919 [0.918, 0.920],作为上限基准。在N=211样本中,所有方法的配对ArcFace增益(输出到真实值减输入到真实值)均为负(Envisage -0.048最小,对比ICEdit -0.139,Kontext -0.242,InstructPix2Pix -0.294;p < 1e-4),外部验证在457对ASPS/PCA语料中显示更大负差。使用SurgicalScore,Envisage得分为0.599 [0.579, 0.619],在两项指标上均最优,但全脸身份持续为负,表明全脸相似度与局部手术精度不匹配。5种子真实值代理(上界,不可部署)使残余弧面差距缩小73%(-0.054至-0.015),且33.9%案例出现正增益,表明学习排序器仍有提升空间。对于局部编辑,应以编辑区域保真度衡量进展。我们发布Envisage、SurgicalScore、预设定义及匹配划分清单。
原文摘要 · Abstract (English)
Localized generative editing needs localized evaluation: full-image identity metrics are structurally confounded under hard-composited edits. We present Envisage, a FLUX.1-Fill inpainting reference pipeline for rhinoplasty goal visualization from a single frontal photograph. The pipeline combines 8 rhinoplasty clinical presets (the released framework also includes 8 blepharoplasty and 8 rhytidectomy presets), MediaPipe masks, and hard-mask compositing. The composite preserves outside-mask pixels by construction, so full-face identity scores are dominated by copied pixels rather than by the diffusion backbone. Because full-face identity metrics cannot grade localized edits, we introduce SurgicalScore, a mask-decomposed 0-1 protocol scoring edit direction, edit magnitude, masked LPIPS, realism, and outside-mask preservation; SS_raw assigns 0.919 [0.918, 0.920] to a perfect-predictor control , anchoring the ceiling. On N=211, the paired ArcFace gain (output-to-GT minus input-to-GT) is negative for all methods (Envisage -0.048 smallest, vs. ICEdit -0.139, Kontext -0.242, InstructPix2Pix -0.294; p < 1e-4), with external validation on a 457-pair ASPS/PCA corpus showing a larger negative gap. With SurgicalScore, Envisage achieves the highest score (0.599 [0.579, 0.619]) and leads on both metrics, but the all-negative ArcFace gap shows that full-face identity is poorly aligned with localized surgical accuracy under hard compositing. A 5-seed GT-oracle (an upper bound, not a deployable result) reduces the residual ArcFace gap by 73% (-0.054 to -0.015), with positive output-to-GT gain on 33.9% of cases, indicating candidate-space headroom for a learned ranker. For localized edits, progress should be measured with edit-region fidelity rather than full-face identity metrics. We release Envisage, SurgicalScore, preset definitions, and matched split manifests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。