用视觉常识优化场景图生成,提升稀疏标注下的准确性。
Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation

- 从数据中自动挖掘空间、功能等关系规律,构建常识约束
- 推理时无需重训练,即刻提升预测准确率,跨模型通用
- 适合需要可靠常识推理的场景理解任务
学习驱动的场景图生成(SGG)模型在常见关系类型上表现优异,但在标注稀疏情况下性能急剧下降,难以捕捉可靠的视觉常识知识。我们提出一种模型无关、语义引导的知识精炼框架,通过系统性地从训练数据中挖掘基于常识的关系约束——包括空间、功能和定性关系规律——并利用通用的声明式常识推理,在推理阶段对排序后的SGG预测结果进行修正与优化。该框架无需人工规则编写,无需模型重训练,且可跨数据集与架构迁移。在三个标准基准上,均持续优于强基线,表明对深度场景语义进行结构化视觉常识推理,是纯学习方法的有效补充。
原文摘要 · Abstract (English)
Learning-driven Scene Graph Generation (SGG) models excel on frequent relation types but degrade sharply under annotation sparsity, failing to capture reliable visual commonsense knowledge. We propose a model-agnostic, semantically-guided knowledge refinement framework that systematically mines commonsense-grounded constraints from training data - capturing spatial, functional, and qualitative relational regularities - and uses general declarative commonsense reasoning to correct and refine ranked SGG predictions at inference time. The framework requires no manual rule authoring, no model retraining, and transfers across datasets and architectures. On three standard benchmarks, we obtain consistent improvements over strong baselines, demonstrating that structured visual commonsense reasoning over deep scene semantics is a practical and effective complement to purely learning-based scene graph generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。