构建首个中文电商海报细粒度评估框架,支持自动化质量检测
E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thought
- 基于专家标注的思维链,建立多维度评分数据集
- 模型与人工评估一致率达0.87,显著优于现有方法
- 适合电商AI生成内容质检、设计工具优化研究者
生成式AI广泛用于商业海报创作,但生成技术进步快于自动质量评估。现有模型侧重通用美学或低级失真,缺乏电商设计所需的实用性标准。尤其对中文内容,复杂字符常引发细微但关键的文本瑕疵,现有方法难以识别。为此,我们提出E-comIQ-ZH框架,构建首个包含多维评分与专家校准思维链(CoT)解释的18,000样本数据集E-comIQ-18k。基于此,训练出与人类专家判断高度一致的评估模型E-comIQ-M。该框架实现首个可扩展的中文电商海报自动生成评测基准E-comIQ-Bench。大量实验表明,E-comIQ-M在专家标准下相关性达0.87,显著优于基线模型,支持高效自动化评估。所有数据集、模型与评估工具将公开,代码已发布于https://github.com/4mm7/E-comIQ-ZH。
原文摘要 · Abstract (English)
Generative AI is widely used to create commercial posters. However, rapid advances in generation have outpaced automated quality assessment. Existing models emphasize generic esthetics or low level distortions and lack the functional criteria required for e-commerce design. It is especially challenging for Chinese content, where complex characters often produce subtle but critical textual artifacts that are overlooked by existing methods. To address this, we introduce E-comIQ-ZH, a framework for evaluating Chinese e-commerce posters. We build the first dataset E-comIQ-18k to feature multi dimensional scores and expert calibrated Chain of Thought (CoT) rationales. Using this dataset, we train E-comIQ-M, a specialized evaluation model that aligns with human expert judgment. Our framework enables E-comIQ-Bench, the first automated and scalable benchmark for the generation of Chinese e-commerce posters. Extensive experiments show our E-comIQ-M aligns more closely with expert standards and enables scalable automated assessment of e-commerce posters. All datasets, models, and evaluation tools will be released to support future research in this area.Code will be available at https://github.com/4mm7/E-comIQ-ZH.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。