arXiv:2605.03863cs.AIcs.CV2026-05

用视觉大模型量化日常所见对心理的影响,发现绿色越多心情越好。

Quantifying the human visual exposome with vision language models

  • 结合实时拍照与视觉语言模型,分析人眼真实看到的环境
  • 33%的视觉特征与情绪、压力显著相关,绿色最有益
  • 自动挖掘上百万文献,找到近1000个影响心理的环境因素

视觉环境是影响心理健康的关键但尚未量化的因素。现有方法依赖粗略地理数据或主观报告,无法捕捉日常生活中的第一视角视觉体验。我们通过生态瞬时评估结合视觉语言模型(VLMs),量化人类视觉体验的语义丰富度。在2674张参与者拍摄的照片中,VLM估算的绿色程度能稳健预测即时情绪和慢性压力,符合已有研究基准。随后,我们开发了一种半自动的大语言模型(LLM)管道,从七百多万篇科学论文中提取出近1000个与心理健康相关的环境特征。当应用于真实世界图像时,最多有33%的VLM提取的上下文评分与情绪和压力显著相关。这些结果建立了一个可扩展的客观视觉暴露组学范式,实现高通量解码可见世界与心理健康的关系。

原文摘要 · Abstract (English)

The visual environment is a fundamental yet unquantified determinant of mental health. While the concept of the environmental exposome is well established, current methods rely on coarse geospatial proxies or biased self reports, failing to capture the first person visual context of daily life. We addressed this gap by coupling ecological momentary assessment with vision language models (VLMs) to quantify the semantic richness of human visual experience. Across 2674 participant generated photographs, VLM derived estimates of greenness robustly predicted momentary affect and chronic stress, consistent with established benchmarks. We then developed a semi autonomous large language model (LLM) based pipeline that mined over seven million scientific publications to extract nearly 1000 environmental features empirically linked to mental health. When applied to real world imagery, up to 33 percent of VLM extracted context ratings significantly correlated with affect and stress. These findings establish a scalable objective paradigm for visual exposomics, enabling high throughput decoding of how the visible world is associated with mental health.

视觉暴露组心理健康大模型语义分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。