arXiv:2604.03061cs.CV2026-04中稿 · CVPR被引 4

纳米香蕉2在图像修复任务中表现优异,但细节过增强问题仍需改进。

Can Nano Banana 2 Replace Traditional Image Restoration Models? An Evaluation of Its Performance on Image Restoration Tasks

  • 通过简洁提示与显式保真约束提升重建与感知质量平衡
  • 全参考指标表现媲美传统模型,用户偏好度高
  • 适合追求视觉效果的修复场景,但需注意保真度控制

生成式AI的发展引发了一个问题:通用图像编辑模型能否成为统一的图像修复解决方案?我们对Nano Banana 2在多种场景和退化类型下的表现进行了系统评估。结果表明,提示设计至关重要,简洁提示配合明确保真度约束可实现重建质量与感知质量间的更好平衡。纳米香蕉2在全参考指标上表现具有竞争力,用户研究中也持续更受欢迎,且在复杂场景中表现出强泛化能力。然而,模型倾向于生成视觉丰富但细节过度增强、存在不一致的结果,这一问题未被现有IQA指标或用户研究充分捕捉。总体而言,通用模型从感知角度展现出作为统一修复方案的潜力,但需提升可控性与保真度感知评价能力。更多对比与详细分析见项目仓库:https://github.com/yxyuanxiao/NanoBanana2TestOnIR。

原文摘要 · Abstract (English)

Recent advances in generative AI raise the question of whether general-purpose image editing models can serve as unified solutions for image restoration. We conduct a systematic evaluation of Nano Banana 2 across diverse scenes and degradations. Our results show that prompt design is critical, with concise prompts and explicit fidelity constraints achieving a better balance between reconstruction and perceptual quality. Nano Banana 2 achieves competitive full-reference performance and is consistently preferred in user studies, while showing strong generalization in challenging scenarios. However, we observe a gap between perceptual quality and restoration fidelity, as the model tends to produce visually rich results with over-enhanced details and inconsistencies. This issue is not well captured by existing IQA metrics or user studies. Overall, general-purpose models show promise as unified IR solvers from a perceptual perspective, but require improved controllability and fidelity-aware evaluation. Further comparisons and detailed analyses are available in our project repository: https://github.com/yxyuanxiao/NanoBanana2TestOnIR.

图像修复生成模型感知质量保真度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。