用多智能体生成细粒度图文负样本,提升视觉理解模型性能
FineGen: A VLM-based Multi-Agent Framework for Fine-Grained Image-Text Dataset Construction

- 基于VLM的多智能体协作生成-验证-修正闭环流程
- 构建14.7万条细粒度负样本,属性有效率达96.7%
- 在硬样本上提升14.4%准确率,适合细粒度识别研究者
当前视觉语言数据集中硬负样本稀缺,严重制约细粒度感知能力。为此,我们提出FineGen,一种基于视觉语言模型的多智能体框架,实现自动化数据集构建。通过生成-验证-修正的协同闭环机制,确保合成的硬负样本语义合理但与图像内容严格矛盾。将其应用于ImageNet,构建了FineGen-100K,一个包含超过14.7万条特定属性硬负样本的分层数据集,正负样本比为1:10。大量评估显示属性有效率达96.7%。关键的是,在FG-OVD基准上的下游验证表明,使用FineGen-100K微调可使硬样本准确率提升14.4%,显著优于现有方法。
原文摘要 · Abstract (English)
The scarcity of hard negative samples in current vision-language datasets significantly hinders fine-grained perception. To address this, we propose FineGen, a VLM-based Multi-Agent framework for automated dataset construction. By employing a collaborative Generation-Verification-Correction pipeline with a closed-loop feedback mechanism, FineGen ensures synthesized hard negatives are semantically valid yet strictly contradictory to visual content. Applying this to ImageNet, we construct FineGen-100K, a hierarchical dataset containing over 147,000 attribute-specific hard negatives with a rigorous 1:10 positive-to-negative ratio. Extensive evaluations confirm a 96.7% attribute validity rate. Crucially, downstream validation on the FG-OVD benchmark shows that fine-tuning on FineGen-100K yields a substantial +14.4% accuracy improvement on hard samples, significantly outperforming state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。