AI视频模型描绘抑郁时,平台不同导致叙事差异显著。
Depictions of Depression in Generative AI Video Models: A Preliminary Study of OpenAI's Sora 2
- 用单关键词'抑郁'生成100段视频,对比消费者端与开发者API输出
- 消费者端78%视频展现好转趋势,开发者端仅14%,且亮度和运动量显著更高
- 画面重复使用帽子、窗户、雨等符号,人物多为20-30岁独居青年
生成式视频模型正日益能呈现复杂的心理健康体验,但对抑郁等状况的呈现方式知之甚少。本研究分析OpenAI Sora 2如何描绘抑郁,并比较消费者App与开发者API两种访问方式的差异。使用单字提示词“抑郁”生成100个视频,每种方式各50个。两名训练编码员独立评估叙事结构、视觉环境、物体、人物特征与状态。提取并对比了视觉美学、音频、语义内容及时间动态等计算特征。结果显示,消费者端视频存在显著康复倾向:78%(39/50)呈现从抑郁到缓解的叙事弧线,而开发者端仅14%(7/50)。消费者端视频随时间变亮(斜率=2.90亮度单位/秒,对比API的-0.18;d=1.59,q<.001),且运动量高出三倍(d=2.07,q<.001)。两类输出均局限于狭窄的视觉词汇库,反复出现的物体包括连帽衫(n=194)、窗户(n=148)和雨(n=83)。人物多为20-30岁年轻人(88%),几乎总是独自一人(98%)。性别分布因访问方式而异:消费者端偏男性(68%),开发者端偏女性(59%)。Sora 2并未创造新的抑郁视觉语法,而是压缩重组文化符号,平台级约束显著影响最终叙事呈现。临床人员应意识到,AI生成的心理健康内容反映训练数据与平台设计,而非临床知识,患者在脆弱期可能接触此类内容。
原文摘要 · Abstract (English)
Generative video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions like depression. This study characterizes how OpenAI's Sora 2 generative video model depicts depression and examines whether depictions differ between the consumer App and developer API access points. We generated 100 videos using the single-word prompt "Depression" across two access points: the consumer App (n=50) and developer API (n=50). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Computational features across visual aesthetics, audio, semantic content, and temporal dynamics were extracted and compared between modalities. App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution, compared with 14% (7/50) of API outputs. App videos brightened over time (slope = 2.90 brightness units/second vs. -0.18 for API; d = 1.59, q < .001) and contained three times more motion (d = 2.07, q < .001). Across both modalities, videos converged on a narrow visual vocabulary and featured recurring objects including hoodies (n=194), windows (n=148), and rain (n=83). Figures were predominantly young adults (88% aged 20-30) and nearly always alone (98%). Gender varied by access point: App outputs skewed male (68%), API outputs skewed female (59%). Sora 2 does not invent new visual grammars for depression but compresses and recombines cultural iconographies, while platform-level constraints substantially shape which narratives reach users. Clinicians should be aware that AI-generated mental health video content reflects training data and platform design rather than clinical knowledge, and that patients may encounter such content during vulnerable periods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。