专为电商设计的图像编辑模型,提升商品一致性与细节保真度。
TBStar-Edit: From Image Editing Pattern Shifting to Consistency Enhancement
- 分阶段训练:先学编辑模式切换,再强化外观一致性。
- 在自建电商数据集上VIE得分超越通用模型。
- 适合需要高保真商品图编辑的电商平台使用。
近期图像生成与编辑技术虽在通用领域表现优异,但在电商场景中仍面临一致性不足的问题。为此,我们提出针对电商领域的图像编辑模型TBStar-Edit。通过严谨的数据工程、架构设计与训练策略,实现精确且高保真的编辑,同时保持商品外观与布局完整性。数据工程方面,构建涵盖采集、构建、过滤与增强的全流程数据管道,获取高质量、指令遵循且高度一致的编辑数据。模型架构采用分层框架,包含基础模型、模式切换模块与一致性增强模块。训练采用两阶段策略:第一阶段专注于编辑模式切换,第二阶段聚焦一致性增强,各阶段使用独立数据集分别训练。在自建的电商基准上进行广泛评估,结果表明,TBStar-Edit在客观指标(VIE Score)和主观用户偏好上均优于现有通用编辑模型。
原文摘要 · Abstract (English)
Recent advances in image generation and editing technologies have enabled state-of-the-art models to achieve impressive results in general domains. However, when applied to e-commerce scenarios, these general models often encounter consistency limitations. To address this challenge, we introduce TBStar-Edit, an new image editing model tailored for the e-commerce domain. Through rigorous data engineering, model architecture design and training strategy, TBStar-Edit achieves precise and high-fidelity image editing while maintaining the integrity of product appearance and layout. Specifically, for data engineering, we establish a comprehensive data construction pipeline, encompassing data collection, construction, filtering, and augmentation, to acquire high-quality, instruction-following, and strongly consistent editing data to support model training. For model architecture design, we design a hierarchical model framework consisting of a base model, pattern shifting modules, and consistency enhancement modules. For model training, we adopt a two-stage training strategy to enhance the consistency preservation: first stage for editing pattern shifting, and second stage for consistency enhancement. Each stage involves training different modules with separate datasets. Finally, we conduct extensive evaluations of TBStar-Edit on a self-proposed e-commerce benchmark, and the results demonstrate that TBStar-Edit outperforms existing general-domain editing models in both objective metrics (VIE Score) and subjective user preference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。