arXiv:2602.07014cs.CVcs.AI2026-02

首个面向电商图文翻译的视觉质量评估框架,解决无参考、细粒度评价难题。

Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation

  • 基于多模态大模型构建无参考评分体系,分解视觉质量为14个可解释维度。
  • 在2000张真实商品图上实现与人工评分高度相关,优于GPT-5和Gemini-3。
  • 适合电商AI系统优化者、图像质量评估研究者使用。

图文机器翻译(IIMT)支撑跨境电商业务;现有研究聚焦于机器翻译评估,而视觉呈现质量对用户参与至关重要。面对密集上下文的商品图像和多模态缺陷,现有基于参考的方法(如SSIM、FID)缺乏可解释性,模型自评方法则缺少领域相关的细粒度奖励信号。为此,我们提出Vectra,据我们所知首个面向电商IIMT的无参考、多模态大模型驱动的视觉质量评估框架。Vectra包含三部分:(1) Vectra Score,一个多维度质量度量系统,将视觉质量分解为14个可解释维度,并引入空间感知的缺陷区域占比(DAR)量化以减少标注歧义;(2) Vectra Dataset,通过多样性采样从110万真实商品图像构建,包含2000张基准测试集、3万条推理标注用于指令微调,以及3.5万条专家标注偏好用于对齐与评估;(3) Vectra Model,一个40亿参数的MLLM,可生成定量分数与诊断性推理。实验表明,Vectra在与人工评分的相关性上达到当前最优,其模型在评分性能上超越主流MLLM,包括GPT-5和Gemini-3。数据集与模型将在论文接受后公开。

原文摘要 · Abstract (English)

In-Image Machine Translation (IIMT) powers cross-border e-commerce product listings; existing research focuses on machine translation evaluation, while visual rendering quality is critical for user engagement. When facing context-dense product imagery and multimodal defects, current reference-based methods (e.g., SSIM, FID) lack explainability, while model-as-judge approaches lack domain-grounded, fine-grained reward signals. To bridge this gap, we introduce Vectra, to the best of our knowledge, the first reference-free, MLLM-driven visual quality assessment framework for e-commerce IIMT. Vectra comprises three components: (1) Vectra Score, a multidimensional quality metric system that decomposes visual quality into 14 interpretable dimensions, with spatially-aware Defect Area Ratio (DAR) quantification to reduce annotation ambiguity; (2) Vectra Dataset, constructed from 1.1M real-world product images via diversity-aware sampling, comprising a 2K benchmark for system evaluation, 30K reasoning-based annotations for instruction tuning, and 3.5K expert-labeled preferences for alignment and evaluation; and (3) Vectra Model, a 4B-parameter MLLM that generates both quantitative scores and diagnostic reasoning. Experiments demonstrate that Vectra achieves state-of-the-art correlation with human rankings, and our model outperforms leading MLLMs, including GPT-5 and Gemini-3, in scoring performance. The dataset and model will be released upon acceptance.

视觉评估电商AI多模态无参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。