arXiv:2512.15110cs.CV2025-12被引 22

纳米香蕉Pro在14项低层视觉任务中表现惊艳,但定量指标仍不及专业模型。

Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets

  • 仅用文本提示零样本测试,无需微调
  • 主观画质超越专家模型,但客观指标落后
  • 适合追求创意效果的用户,不推荐高精度场景

文本到图像生成模型的快速发展彻底改变了视觉内容创作。尽管商业产品如纳米香蕉Pro广受关注,其作为传统低层视觉任务通用求解器的潜力却尚未被充分探索。本文深入探讨核心问题:纳米香蕉Pro是否是低层视觉通才?我们在40个不同数据集上的14项独立低层视觉任务中进行了全面的零样本评估。通过简单文本提示,不进行微调,将纳米香蕉Pro与最先进的专用模型进行对比。大量分析显示显著性能分化:虽然纳米香蕉Pro在主观视觉质量上表现优异,常生成比专用模型更逼真的高频细节,但在传统的参考基准量化指标上仍显不足。我们归因于生成模型固有的随机性,难以满足传统指标所需的严格像素级一致性。本报告确认纳米香蕉Pro是低层视觉任务的有力零样本候选者,同时指出实现领域专家级保真度仍是重大挑战。

原文摘要 · Abstract (English)

The rapid evolution of text-to-image generation models has revolutionized visual content creation. While commercial products like Nano Banana Pro have garnered significant attention, their potential as generalist solvers for traditional low-level vision challenges remains largely underexplored. In this study, we investigate the critical question: Is Nano Banana Pro a Low-Level Vision All-Rounder? We conducted a comprehensive zero-shot evaluation across 14 distinct low-level tasks spanning 40 diverse datasets. By utilizing simple textual prompts without fine-tuning, we benchmarked Nano Banana Pro against state-of-the-art specialist models. Our extensive analysis reveals a distinct performance dichotomy: while \textbf{Nano Banana Pro demonstrates superior subjective visual quality}, often hallucinating plausible high-frequency details that surpass specialist models, it lags behind in traditional reference-based quantitative metrics. We attribute this discrepancy to the inherent stochasticity of generative models, which struggle to maintain the strict pixel-level consistency required by conventional metrics. This report identifies Nano Banana Pro as a capable zero-shot contender for low-level vision tasks, while highlighting that achieving the high fidelity of domain specialists remains a significant hurdle.

图像生成零样本视觉质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。