arXiv:2508.10116cs.IR2025-08

用视觉语言对齐技术,从商品图自动生成高质量描述。

Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment

  • 用多模态大模型+指令微调生成图文一致的电商描述
  • 在真实数据集上提升描述质量与结构化字段完成率
  • 适合需要批量生成商品详情的电商平台或卖家

商品信息(如标题和属性)对电商用户参与至关重要。然而,人工或半自动输入结构化商品信息常导致质量不一、错误频发且响应慢,尤其对个人卖家而言。直接从商品图像生成准确描述是一种有前景的替代方案。现有基于检索的方法虽缓解部分问题,但仍难以捕捉细粒度视觉细节,且在小众或专业品类中表现不佳。我们提出优化偏好式商品描述生成框架OPAL,利用微调的多模态大语言模型(MLLM),从图像生成符合结构化模板、高质量的商品描述。OPAL通过两种数据优化方法解决关键挑战:MLLM辅助一致性增强,确保与结构化字段要求对齐;LLM辅助上下文理解,提升对视觉输入中细微信息的捕捉能力。采用视觉指令微调结合直接偏好优化(DPO)对MLLM进行微调,有效减少幻觉并增强不同骨干网络下的鲁棒性。我们在真实电商数据集上评估,结果表明OPAL在描述质量与结构化字段完成率上均持续优于基线方法。实验验证了其在弥合视觉与文本模态差距方面的有效性,能够生成更丰富、准确、一致的商品描述。本工作推动了自动化商品列表优化,支持电商平台实现可扩展的高质量内容生成。

原文摘要 · Abstract (English)

Item information, such as titles and attributes, is essential for effective user engagement in e-commerce. However, manual or semi-manual entry of structured item specifics often produces inconsistent quality, errors, and slow turnaround, especially for Customer-to-Customer sellers. Generating accurate descriptions directly from item images offers a promising alternative. Existing retrieval-based solutions address some of these issues but often miss fine-grained visual details and struggle with niche or specialized categories. We propose Optimized Preference-Based AI for Listings (OPAL), a framework for generating schema-compliant, high-quality item descriptions from images using a fine-tuned multimodal large language model (MLLM). OPAL addresses key challenges in multimodal e-commerce applications, including bridging modality gaps and capturing detailed contextual information. It introduces two data refinement methods: MLLM-Assisted Conformity Enhancement, which ensures alignment with structured schema requirements, and LLM-Assisted Contextual Understanding, which improves the capture of nuanced and fine-grained information from visual inputs. OPAL uses visual instruction tuning combined with direct preference optimization to fine-tune the MLLM, reducing hallucinations and improving robustness across different backbone architectures. We evaluate OPAL on real-world e-commerce datasets, showing that it consistently outperforms baseline methods in both description quality and schema completion rates. These results demonstrate that OPAL effectively bridges the gap between visual and textual modalities, delivering richer, more accurate, and more consistent item descriptions. This work advances automated listing optimization and supports scalable, high-quality content generation in e-commerce platforms.

多模态电商生成大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。