用分层语义优化提升专业领域图像生成的准确与多样
AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework
- 分层框架融合提示优化与多视角理解,捕捉全局与局部语义关系
- 仅需每类16张图训练,在40个领域实现高质量、高一致性的生成
- 跨模态适配减少幻觉,适合医疗、工业等专业场景图像生成
专用领域图像生成旨在为特定专业领域生成高质量视觉内容,同时保证语义准确性和细节保真度。然而,现有方法存在两大局限:其一,提示工程与模型适配被分开处理,忽视了专业领域中语义理解与视觉表征的内在关联;其二,生成过程中缺乏对领域语义约束的有效整合,导致出现幻觉和语义偏差。为此,我们提出 AdaptaGen,一种分层语义优化框架,结合基于矩阵的提示优化与多视角理解,从全局与局部双重角度捕捉完整的语义关系。为缓解专业领域的幻觉问题,设计跨模态适配机制,与智能内容合成结合,可在保留核心主题元素的同时引入多样化细节。此外,在生成阶段引入两阶段字幕语义转换,维持语义连贯性的同时提升视觉多样性,确保生成图像符合领域特定约束。实验结果表明,本方法在来自多个数据集的40个类别上仅使用每类16张图像即取得显著性能提升,大幅改善图像质量、多样性与语义一致性。
原文摘要 · Abstract (English)
Domain-specific image generation aims to produce high-quality visual content for specialized fields while ensuring semantic accuracy and detail fidelity. However, existing methods exhibit two critical limitations: First, current approaches address prompt engineering and model adaptation separately, overlooking the inherent dependence between semantic understanding and visual representation in specialized domains. Second, these techniques inadequately incorporate domain-specific semantic constraints during content synthesis, resulting in generation outcomes that exhibit hallucinations and semantic deviations. To tackle these issues, we propose AdaptaGen, a hierarchical semantic optimization framework that integrates matrix-based prompt optimization with multi-perspective understanding, capturing comprehensive semantic relationships from both global and local perspectives. To mitigate hallucinations in specialized domains, we design a cross-modal adaptation mechanism, which, when combined with intelligent content synthesis, enables preserving core thematic elements while incorporating diverse details across images. Additionally, we introduce a two-phase caption semantic transformation during the generation phase. This approach maintains semantic coherence while enhancing visual diversity, ensuring the generated images adhere to domain-specific constraints. Experimental results confirm our approach's effectiveness, with our framework achieving superior performance across 40 categories from diverse datasets using only 16 images per category, demonstrating significant improvements in image quality, diversity, and semantic consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。