人类能否发现AI图像编辑,分注意捕捉和判断两阶段,影响因素不同。
Attention Capture Is Not Detection: A Two-Stage Account of How Humans Miss Localized AI Image Edits

- 区分注意捕捉与判断准确,前者由编辑区域决定,后者由语义合理性决定。
- 实验显示注意捕捉率与语义合理性均显著影响假图识别准确率(p<0.001)。
- 用生成式眼动模型预测漏看现象,效果优于传统方法,但仍有局限。
随着AI生成图像编辑泛滥,平台通常将可检测性视为单一属性:要么被警告,要么不被警告。我们证明这是错误的模型。在一项受控的眼动追踪研究中(N=59,拉丁方设计,四个条件交叉编辑区域与语义合理性),混合效应分析显示,是否注意到编辑与是否正确判断为伪造是可分离的两个阶段,受不同因素影响:编辑区域决定注意捕获(p<0.001),语义合理性决定判断准确率和‘看了却没察觉’(LBFS)错误率(p<0.001)。这一分离在多重比较校正后仍成立;两因素间无显著交互作用。该双阶段模型拓展了视觉注意研究中‘前注意捕获’与‘努力识别’的长期区分,应用于AI编辑可检测性领域。我们进一步测试生成式眼动模型能否计算实现注意捕获阶段:基于Transformer的扫描路径生成模型在预测每张图像注意力分布时表现优异(皮尔逊相关r=0.77–0.82,跨保留刺激),在预测LBFS发生率上,即使未使用语义合理性标签,也略优于双参数线性基线(r=0.52 vs. r=0.48)。我们如实报告该方法的比较结果、消融实验及局限性(仅单次固定训练/验证划分,非留一被试),以负责任方式传达机器学习系统在遏制AI误导信息中的能力边界。
原文摘要 · Abstract (English)
As AI-generated image edits proliferate, the platforms meant to curb the resulting disinformation treat detectability as a single, undifferentiated property: an edit either gets a warning or it does not. We show this is the wrong model. Across a controlled eye-tracking study ($N=59$, Latin-square design, four conditions crossing edit area and semantic plausibility), a mixed-effects analysis reveals that whether an edit is noticed and whether it is correctly judged as fake are dissociable stages, governed by different factors: edit area drives attention capture ($p<0.001$) while semantic plausibility drives judgment accuracy and look-but-fail-to-see (LBFS) error rates ($p<0.001$). This dissociation survives correction for multiple comparisons; a secondary interaction between the two factors does not. This two-stage account extends a long-standing distinction in visual attention research (between pre-attentive capture and effortful recognition) into the new domain of AI-edit detectability. We then test whether a generative eye-movement model can computationally operationalize the attention-capture stage: a Transformer trained to generate scanpaths tracks per-image attention with strong discriminative power (Pearson $r=0.77$--$0.82$ across held-out stimuli) and, on the harder task of predicting LBFS incidence, modestly outperforms a two-parameter linear baseline even without access to the plausibility label ($r=0.52$ vs. $r=0.48$). We report this comparison, our ablations, and our method's limitations (a single fixed train/validation split, not leave-one-subject-out) without inflation, consistent with responsibly communicating what a machine learning system can and cannot do to help curb AI-driven disinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。