arXiv:2509.04889cs.CVcs.AI2025-09被引 1

用视觉模型从恐蛛图像中预测人类恐惧,支持个性化心理治疗。

SpiderNets: Vision Models Predict Human Fear From Aversive Images

  • 用预训练视觉模型迁移学习,直接从图像预测恐惧值。
  • 跨人跨图预测误差低于10分(0-100分制),性能稳定。
  • 模型关注蜘蛛部位,适合开发自适应数字心理疗法。

恐惧行为普遍且具有破坏性,暴露疗法通过呈现引发恐惧的视觉刺激来治疗,是目前最有效的手段。实现可扩展的计算机化暴露疗法需能自动根据图像内容预测恐惧程度,以动态调整刺激与治疗强度。然而,这种预测是否可靠并具备跨个体、跨图像的泛化能力尚不明确。本研究发现,经迁移学习微调的卷积与变压器视觉模型,能准确预测蜘蛛相关图像在群体层面的感知恐惧,即使在新人群和新图像上评估,均实现0-100恐惧量表下的平均绝对误差(MAE)低于10。视觉解释分析表明,预测主要依赖图像中的蜘蛛特定区域。学习曲线分析显示,变压器模型数据效率高,在约300张可用图像下即接近性能饱和。预测误差在极低和极高恐惧水平及特定图像类别中升高。这些结果建立了透明、数据驱动的图像恐惧估计方法,为自适应数字心理健康工具奠定基础。

原文摘要 · Abstract (English)

Phobias are common and impairing, and exposure therapy, which involves confronting patients with fear-provoking visual stimuli, is the most effective treatment. Scalable computerized exposure therapy requires automated prediction of fear directly from image content to adapt stimulus selection and treatment intensity. Whether such predictions can be made reliably and generalize across individuals and stimuli, however, remains unknown. Here we show that pretrained convolutional and transformer vision models, adapted via transfer learning, accurately predict group-level perceived fear for spider-related images, even when evaluated on new people and new images, achieving a mean absolute error (MAE) below 10 units on the 0-100 fear scale. Visual explanation analyses indicate that predictions are driven by spider-specific regions in the images. Learning-curve analyses show that transformer models are data efficient and approach performance saturation with the available data (~300 images). Prediction errors increase for very low and very high fear levels and within specific categories of images. These results establish transparent, data-driven fear estimation from images, laying the groundwork for adaptive digital mental health tools.

视觉模型恐惧预测数字疗法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。