现有图像转3D模型易生成危险几何结构,且多数无法被现有防护机制识别。
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models

- 构建三类危险几何类别,通过多场景测试评估模型生成能力
- 超99.7%的危险结构未触发商业平台内容审核,风险极低可见性
- 提出分层防御策略,可将有害内容留存降至1%以下
图像到3D模型的最新进展显著提升了3D内容创作的保真度与可及性。然而,这种强大的重建能力也可能被恶意利用,生成可通过3D打印制造的危险几何结构,带来现实世界风险。目前此类风险尚未充分研究:当前图像到3D模型对危险几何的生成能力如何,现有防护机制是否可靠仍不明确。为此,我们开展系统性测量研究,定义三类危险类别:直接物理危害、高风险模板/组件、欺骗性复制品,并以代表性物体实例化。评估开源与商用模型在原始、退化、视角偏移及语义伪装输入下的表现,采用几何有效性、多视图VLM语义评分、人工定向验证与可控物理打印等多重指标。结果表明,当前模型能有效重建危险几何,而少于0.3%的生成物触发商业平台内容审核。作为初步缓解措施,我们评估了三类典型防护方案:输入端审核、模型级良性对齐、输出端过滤,发现各具局限。进一步提出分层防御机制,可将有害内容保留率降至<1%,但整体误报率仍达11%。研究揭示当前系统存在显著风险,呼吁发展更精细的几何感知内容安全机制。
原文摘要 · Abstract (English)
Recent advances in image-to-3D models have significantly improved the fidelity and accessibility of 3D content creation. Such a powerful reconstruction capability that enables creative design can also be misused by the adversary to generate harmful geometries, which can be further fabricated via 3D printers and pose real-world risks. However, such risks are largely underexplored: it remains unclear how well current image-to-3D models can produce these harmful geometries, and whether existing safeguards can reliably prevent such generation. To fill this gap, we conduct a systematic measurement study of harmful geometry generation and mitigation. We first describe this risk through three kinds of unsafe categories: direct-use physical hazards, risky templates or components, and deceptive replicas. Each category is instantiated with representative objects. We evaluate both open-source and commercial image-to-3D models under original, degraded, viewpoint-shifted, and semantically camouflaged inputs. We consider different evaluation metrics, including geometric validity, multi-view VLM-based semantic scoring, targeted human validation, and controlled physical fabrication. The results reveal a concerning reality that current image-to-3D models can effectively reconstruct the harmful geometries, while fewer than 0.3% of such geometries trigger commercial moderation flags. As a first step toward mitigation, we evaluate three representative safeguard families, including input moderation, model-level benign alignment, and output-level filtering. We find that existing safeguards have distinct weaknesses. We further develop a stacked defense that can reduce harmful retention to <1%, but still at 11% overall false-positive cost. Taken together, our findings demonstrate that the risk in current system and encourage better geometry-aware safeguards for moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。