用音频大模型让老人在家安全使用智能助手,说话不清也能懂。
DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality
- 直接处理原始音频,不用先转文字
- 能听懂老人模糊说话,还能识别摔倒等紧急情况
- 全程本地运行,保护隐私适合老年人用
我们提出 DESAMO,一种基于嵌入式音频大模型的设备端智能家居系统,专为老年人友好设计,支持自然且私密的交互。传统语音助手依赖 ASR 管道或 ASR-LLM 级联,常因老年用户发音不清而失效,且无法处理非语音音频。DESAMO 采用音频大模型直接处理原始音频输入,实现对用户意图及关键事件(如跌倒、求助)的鲁棒理解。
原文摘要 · Abstract (English)
We present DESAMO, an on-device smart home system for elder-friendly use powered by Audio LLM, that supports natural and private interactions. While conventional voice assistants rely on ASR-based pipelines or ASR-LLM cascades, often struggling with the unclear speech common among elderly users and unable to handle non-speech audio, DESAMO leverages an Audio LLM to process raw audio input directly, enabling a robust understanding of user intent and critical events, such as falls or calls for help.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。