利用人物眼神一致性检测AI生成图像,突破传统低级痕迹检测瓶颈。
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

- 通过眼神方向、头眼对齐等语义一致性设计新检测机制。
- 在多模型数据集上提升准确率,最高达71.5%平衡精度。
- 适用于多种视觉模型,适合图像真实性验证研究者。
近期生成模型已显著减少低层伪影(如像素指纹、频率异常、插值痕迹),尤其在人物中心或局部编辑场景中,篡改区域小且周围内容逼真。本文提出社会凝视一致性这一高层语义线索,指交互个体间眼神方向、头眼对齐与瞳孔位置的相互协调性,并证明其构成与现有低层检测范式正交的新型检测轴。通过三项耦合机制实现:(i) 设计可控诊断数据集,对眼神一致图像进行区域特异性扰动,严格成对分组以避免生成器指纹记忆作为优化捷径;(ii) 块组合描述监督,保持1,250个宏观组合描述中单一5块推理骨架不变,解耦推理一致性与表面多样性;(iii) 跨架构验证显示,相同监督使视觉语言骨干(FakeVLM)在COCOAI交互子集上准确率提升+3.7个百分点(67.8→71.5),在人物子集上提升+1.3个百分点(83.0→84.3),视觉单干骨干(Effort)亦有持续增益,证明该线索具备骨干无关性。真实与虚假类别召回率同步上升,排除‘全判为假’的伪影。四步机制解释:成对编辑捷径阻断、难到易难度迁移、CLIP先验保留、扩散族共有的眼周结构谱弱点,说明仅用一个修补器(FLUX.1-Fill)训练即可泛化至多生成器套件。代码将于录用后公开,以促进可复现性。
原文摘要 · Abstract (English)
Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularly in person-centric and partial-edit settings where the manipulated region is small and surrounded by photometrically authentic content. We introduce Social Gaze Consistency, a high-level semantic cue defined as the mutual coherence of gaze direction, head-eye alignment, and pupil placement between interacting individuals, and show that it constitutes a previously underutilized detection axis orthogonal to existing low-level paradigms. We instantiate this insight through three coupled mechanisms: (i) a controlled diagnostic dataset with region-specific perturbations of gaze-consistent imagery, where strict pair-level grouping forecloses generator-fingerprint memorization as an optimization-time shortcut rather than relying on augmentation; (ii) Block-Compositional Caption Supervision, which holds a single 5-block reasoning skeleton invariant across 1,250 macro-combined captions, decoupling reasoning consistency from surface diversity; (iii) Cross-architecture validation showing the same supervision improves a vision-language backbone (FakeVLM) by +3.7 pp on the COCOAI Interaction subset (balanced accuracy 67.8 -> 71.5) and +1.3 pp on the COCOAI Person subset (83.0 -> 84.3), with consistent gains on a vision-only backbone (Effort), evidencing a backbone-agnostic cue. Real- and fake-class recalls rise simultaneously, ruling out a "predict-all-fake" artifact. A four-step mechanistic account - paired-edit shortcut blocking, hard-to-easy difficulty transfer, CLIP prior preservation, and diffusion-family shared spectral weakness in periocular structure - explains why training on a single inpainter (FLUX.1-Fill) transfers to multi-generator suites. We will release the code upon acceptance to facilitate reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。