arXiv:2508.09385cs.LGcs.AI2025-08中稿 · Interspeech 2025被引 1

用扩散模型将痴呆语音转为图像,竟可凭图识病,准确率达75%。

Understanding Dementia Speech Alignment with Diffusion-Based Image Generation

  • 将痴呆语音输入扩散图像生成模型,生成对应图像。
  • 仅从生成图像中即可实现75%的痴呆检测准确率。
  • 通过可解释性方法定位语音中影响判断的关键语义片段。

文本到图像模型基于自然语言描述生成高度逼真的图像,数百万用户在线使用此类模型创作和分享图像。尽管预期这些模型能在潜在空间中对齐输入文本与生成图像,但很少研究探讨病理语音与生成图像之间的对齐可能性。本文研究了此类模型在对齐痴呆相关语音信息与生成图像方面的能力,并开发了相应的解释方法。令人惊讶的是,仅从生成图像中即可实现75%的痴呆检测准确率,该结果在ADReSS数据集上验证。随后,我们利用可解释性方法分析了语言中哪些部分贡献于检测结果。

原文摘要 · Abstract (English)

Text-to-image models generate highly realistic images based on natural language descriptions and millions of users use them to create and share images online. While it is expected that such models can align input text and generated image in the same latent space little has been done to understand whether this alignment is possible between pathological speech and generated images. In this work, we examine the ability of such models to align dementia-related speech information with the generated images and develop methods to explain this alignment. Surprisingly, we found that dementia detection is possible from generated images alone achieving 75% accuracy on the ADReSS dataset. We then leverage explainability methods to show which parts of the language contribute to the detection.

痴呆检测图像生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。