arXiv:2503.22019cs.CV2025-03CVPR被引 4

用注意力引导扩散模型,实现跨域植物性状识别的精准图像与标签转换。

AGILE: A Diffusion-Based Attention-Guided Image and Label Translation for Efficient Cross-Domain Plant Trait Identification

  • 基于扩散模型,通过优化文本嵌入和注意力图控制对象位置。
  • 在跨域植物数据集上提升目标检测性能,保持图像真实感与语义一致性。
  • 适合需要跨域数据增强的农业植物性状识别研究者使用。

语义一致的跨域图像翻译可通过跨域转移标签生成训练数据,对农业植物性状识别尤为有用。然而,现有生成模型在域间差异显著时难以保持物体级精度。本文提出AGILE(Attention-Guided Image and Label Translation for Efficient Cross-Domain Plant Trait Identification),一种基于扩散模型的框架,利用优化的文本嵌入和注意力引导来语义约束图像翻译。AGILE采用预训练扩散模型和公开农业数据集,在保持图像真实性的同时提升翻译图像的语义保真度。方法通过优化文本嵌入强化源域与目标域图像间的对应关系,并在去噪过程中引导注意力图以控制对象位置。在多个跨域植物数据集上的实验表明,AGILE能生成语义准确的翻译图像,显著提升目标域的目标检测性能,优于先前图像翻译方法,尤其在对象差异大或域差距显著的情况下表现更优。

原文摘要 · Abstract (English)

Semantically consistent cross-domain image translation facilitates the generation of training data by transferring labels across different domains, making it particularly useful for plant trait identification in agriculture. However, existing generative models struggle to maintain object-level accuracy when translating images between domains, especially when domain gaps are significant. In this work, we introduce AGILE (Attention-Guided Image and Label Translation for Efficient Cross-Domain Plant Trait Identification), a diffusion-based framework that leverages optimized text embeddings and attention guidance to semantically constrain image translation. AGILE utilizes pretrained diffusion models and publicly available agricultural datasets to improve the fidelity of translated images while preserving critical object semantics. Our approach optimizes text embeddings to strengthen the correspondence between source and target images and guides attention maps during the denoising process to control object placement. We evaluate AGILE on cross-domain plant datasets and demonstrate its effectiveness in generating semantically accurate translated images. Quantitative experiments show that AGILE enhances object detection performance in the target domain while maintaining realism and consistency. Compared to prior image translation methods, AGILE achieves superior semantic alignment, particularly in challenging cases where objects vary significantly or domain gaps are substantial.

图像翻译扩散模型植物识别注意力引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。