arXiv:2603.27295cs.HCcs.AI2026-03中稿 · CHI 2026

用生成音频让视障者感知远方风景的美感。

Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista Landscapes

  • 基于心理声学设计生成式音频,还原视觉场景的听觉体验。
  • 用户测试显示,语音+音效组合比纯语音更易想象场景。
  • 适合需要沉浸式户外感知的视障人群,提升生活美学体验。

当前面向视障及低视力(BLV)人群的场景感知工具依赖口语描述,但缺乏对优美远距离环境景观(Vista空间)的生动呈现。本文提出的Scene2Audio框架,利用受心理声学与场景音频构成原理指导的生成模型,生成可理解且具欣赏性的非语言音频。11名BLV参与者参与的用户研究发现,将Scene2Audio音效与语音结合,相比纯语音更能增强场景想象。另一项持续超过一周的移动端“真实世界”研究,涉及7名BLV用户,进一步验证了该技术在提升户外场景体验方面的潜力。本工作通过超越单纯描述性辅助,弥合了视觉与听觉场景感知之间的鸿沟,回应了BLV用户对美学体验的需求。

原文摘要 · Abstract (English)

Current scene perception tools for Blind and Low Vision (BLV) individuals rely on spoken descriptions but lack engaging representations of visually pleasing distant environmental landscapes (Vista spaces). Our proposed Scene2Audio framework generates comprehensible and enjoyable nonverbal audio using generative models informed by psychoacoustics, and principles of scene audio composition. Through a user study with 11 BLV participants, we found that combining the Scene2Audio sounds with speech creates a better experience than speech alone, as the sound effects complement the speech making the scene easier to imagine. A mobile app "in-the-wild" study with 7 BLV users for more than a week further showed the potential of Scene2Audio in enhancing outdoor scene experiences. Our work bridges the gap between visual and auditory scene perception by moving beyond purely descriptive aids, addressing the aesthetic needs of BLV users.

无障碍音频生成视障辅助生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。