用图文生成+图像编辑结合,提升罕见物体分割效果
TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation

- 图文生成提供多样场景,图像编辑精准插入高置信度目标
- 在LVIS上整体准确率提升4.0点,罕见类提升9.5点
- 适合长尾分布下实例分割的训练数据增强
大规模词汇实例分割受限于类别分布长尾和细粒度类间模糊性。数据合成虽有潜力,但现有方法各有局限:图文生成(T2I)继承噪声伪标签,对稀有类别表现差;拷贝粘贴方法牺牲上下文真实性。为此,我们提出一种混合流程,融合T2I生成与上下文感知的图像到图像(I2I)编辑。T2I分支提供广泛的类别与场景多样性,教师-学生机制通过仅保留提示指定类别确保标签可靠性。为强化稀有类别监督,引入VRAIN(经指令编辑验证的稀有类增强),该I2I编辑器将高置信度实例精准插入真实场景语义位置,生成语义一致且视觉自然的修改,减少域差距并实现针对性增强。在LVIS基准上,本方法超越现有基线,整体平均精度(AP)最高提升4.0点,稀有类AP最高提升9.5点,且随主干网络容量有效扩展。
原文摘要 · Abstract (English)
Large-vocabulary instance segmentation is constrained by long-tailed category distributions and fine-grained inter-class ambiguity. While data synthesis offers a promising alternative, current paradigms have complementary limitations: text-to-image (T2I) methods inherit noisy pseudo-labels and struggle on rare classes, whereas copy-paste methods compromise contextual realism. To address these issues, we propose a hybrid pipeline coupling T2I generation with context-aware image-to-image (I2I) editing. The T2I branch provides broad category and scene diversity, while a teacher-student scheme ensures label reliability by selectively retaining only prompt-specified categories. To strengthen supervision for rare classes, we introduce VRAIN (Verified Rare-class Augmentation via INstructed editing), a novel I2I editor. VRAIN inserts high-confidence instances at semantically appropriate locations within in-the-wild scenes, yielding semantically coherent and visually natural edits that reduce domain gaps and enable targeted augmentation. On the LVIS benchmark, our method surpasses existing baselines, improving overall AP by up to +4.0 points and rare-class AP by up to +9.5 points, while scaling effectively with backbone capacity. Our project page is available at https://seokhunchoi.github.io/TMI
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。