arXiv:2603.18719cs.CVcs.AI2026-03被引 1

用知识图谱指导扩散模型,让仿真图像自动变真实。

Ontology-Guided Diffusion for Zero-Shot Visual Sim2Real Transfer

  • 将真实感拆解为光照、材质等可解释特征,构建知识图谱。
  • 图嵌入+符号规划,使生成图像在多个基准上更接近真实。
  • 无需真实数据,适合零样本仿真转现实任务。

弥合仿真到现实(sim2real)的差距仍具挑战,因真实世界标注数据稀缺。现有基于扩散的方法依赖无结构提示或统计对齐,未能捕捉使图像显得真实的结构性因素。本文提出本体引导扩散(OGD),一种神经符号式的零样本sim2real图像迁移框架,将真实感表示为结构化知识。OGD将真实感分解为一组可解释特征(如光照、材质属性),并在知识图谱中编码其关系。从合成图像出发,OGD推断特征激活状态,并通过图神经网络生成全局嵌入;同时,符号规划器利用本体特征计算出缩小真实感差距的一致性视觉编辑序列。图嵌入通过交叉注意力条件化预训练指令引导扩散模型,计划编辑转化为结构化指令提示。在多个基准测试中,基于图的嵌入比基线更好地区分真实与合成图像,且OGD在sim2real图像转换上超越当前最优扩散方法。整体表明,显式编码真实感结构能实现可解释、数据高效且泛化性强的零样本sim2real迁移。

原文摘要 · Abstract (English)

Bridging the simulation-to-reality (sim2real) gap remains challenging as labelled real-world data is scarce. Existing diffusion-based approaches rely on unstructured prompts or statistical alignment, which do not capture the structured factors that make images look real. We introduce Ontology- Guided Diffusion (OGD), a neuro-symbolic zero-shot sim2real image translation framework that represents realism as structured knowledge. OGD decomposes realism into an ontology of interpretable traits -- such as lighting and material properties -- and encodes their relationships in a knowledge graph. From a synthetic image, OGD infers trait activations and uses a graph neural network to produce a global embedding. In parallel, a symbolic planner uses the ontology traits to compute a consistent sequence of visual edits needed to narrow the realism gap. The graph embedding conditions a pretrained instruction-guided diffusion model via cross-attention, while the planned edits are converted into a structured instruction prompt. Across benchmarks, our graph-based embeddings better distinguish real from synthetic imagery than baselines, and OGD outperforms state-of-the-art diffusion methods in sim2real image translations. Overall, OGD shows that explicitly encoding realism structure enables interpretable, data-efficient, and generalisable zero-shot sim2real transfer.

图像生成扩散模型零样本知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。