arXiv:2503.17224cs.CVcs.AI2025-03ICML

用符号化场景图提升合成图像数据质量,让模型更懂物体间关系。

Neuro-Symbolic Scene Graph Conditioning for Synthetic Image Dataset Generation

  • 用场景图结构显式编码物体关系,指导合成图像生成。
  • 合成数据加真实数据后,召回率最高提升2.59%。
  • 适合需要复杂视觉推理的合成数据生成任务。

随着机器学习模型规模和复杂度增加,获取足够训练数据成为瓶颈,受采集成本、隐私限制及专业领域数据稀缺影响。合成数据生成虽是可行替代方案,但与真实数据训练模型相比仍存在显著性能差距,尤其在任务复杂度上升时。神经符号方法结合神经网络的学习能力与符号推理的结构化表示,在多种认知任务中展现出潜力。本文探索神经符号条件对合成图像数据集生成的作用,聚焦于提升场景图生成模型性能。研究验证了以场景图为形式的结构化符号表示能否通过显式编码关系约束来改善合成数据质量。结果表明,采用神经符号条件进行数据增强,标准召回率提升最高达+2.59%,无图约束召回率提升+2.83%。这些发现证明,融合神经符号与生成方法可生成包含互补结构信息的合成数据,与真实数据结合后能有效提升模型性能,为克服复杂视觉推理任务中的数据稀缺问题提供新路径。

原文摘要 · Abstract (English)

As machine learning models increase in scale and complexity, obtaining sufficient training data has become a critical bottleneck due to acquisition costs, privacy constraints, and data scarcity in specialised domains. While synthetic data generation has emerged as a promising alternative, a notable performance gap remains compared to models trained on real data, particularly as task complexity grows. Concurrently, Neuro-Symbolic methods, which combine neural networks' learning strengths with symbolic reasoning's structured representations, have demonstrated significant potential across various cognitive tasks. This paper explores the utility of Neuro-Symbolic conditioning for synthetic image dataset generation, focusing specifically on improving the performance of Scene Graph Generation models. The research investigates whether structured symbolic representations in the form of scene graphs can enhance synthetic data quality through explicit encoding of relational constraints. The results demonstrate that Neuro-Symbolic conditioning yields significant improvements of up to +2.59% in standard Recall metrics and +2.83% in No Graph Constraint Recall metrics when used for dataset augmentation. These findings establish that merging Neuro-Symbolic and generative approaches produces synthetic data with complementary structural information that enhances model performance when combined with real data, providing a novel approach to overcome data scarcity limitations even for complex visual reasoning tasks.

合成数据场景图神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。