arXiv:2605.18132cs.CVcs.AI2026-05

识别3D生成资产来源,能精准定位是哪个模型生成的。

Who Generated This 3D Asset? Learning Source Attribution for Generative 3D Models

论文配图:Who Generated This 3D Asset? Learning Source Attribution for Generative 3D Models
图 1 · 摘自论文原文
  • 用多视角多模态Transformer融合几何与频域特征,捕捉分散指纹。
  • 全监督下准确率97.22%,仅1%数据时仍达77.17%。
  • 适用于游戏、机器人等领域,保障3D内容可信溯源。

生成式3D模型广泛应用于游戏、机器人和沉浸式创作,源属性识别至关重要:给定一个3D资产,能否判断其由哪个生成模型创建?该问题面临两大挑战:指纹信号分散,分布在多视角、几何和频域特征中;以及真实部署限制,包括标签稀缺、提示退化及真实与合成资产混合导致的可靠性下降。为此,我们构建了首个面向现代生成资产的被动源属性基准,涵盖22个代表性3D生成器,在标准、少样本和真实部署协议下进行评估。基于该基准,我们发现生成3D模型留下两类稳定指纹:跨视角不一致性,以及反映在几何统计与频域线索中的结构伪影。为捕捉这些分散信号,我们提出一种分层多视图多模态Transformer,融合每个视图内的外观、几何与频域特征,并建模视图间的全局关系。大量实验表明,该方法表现优异,在全监督下准确率达97.22%,仅使用1%训练数据(每类少于5个样本)时仍达到77.17%。结果证明现代3D生成器留下可追踪的稳定指纹,建立了新的基准与方法论基础,推动可信3D内容溯源发展。

原文摘要 · Abstract (English)

Generative 3D models are deployed in gaming, robotics, and immersive creation, making source attribution critical: given a 3D asset, can we identify whether and which generative model created it? This problem faces two core challenges: dispersed attribution signals, where 3D fingerprints are distributed across multi-view, geometric, and frequency-domain cues; and realistic deployment constraints, where scarce labels, degraded prompts, and mixed real/synthetic assets undermine attribution reliability. To systematically study this problem, we construct, to the best of our knowledge, the first passive source attribution benchmark for modern generated assets, covering 22 representative 3D generators under standard, few-shot, and realistic deployment protocols. Based on this benchmark, we find that generative 3D models leave two types of stable fingerprints: cross-view inconsistency and structural artifacts reflected in geometric statistics and frequency-domain cues. To capture these dispersed signals, we propose a hierarchical multi-view multi-modal Transformer that fuses appearance, geometric, and frequency-domain features within each view and models global relationships across views. Extensive experiments demonstrate strong performance, achieving 97.22% accuracy under full supervision and 77.17% accuracy with only 1% training data, corresponding to fewer than five samples per generator. These results show that modern 3D generators leave stable and attributable fingerprints, establishing a new benchmark and methodological foundation for trustworthy 3D content provenance.

3D生成内容溯源指纹识别多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。