语音助手助养老院精准管理,安全评估框架验证高准确率
Evaluating a Multi-Agent Voice-Enabled Smart Speaker for Care Homes: A Safety-Focused Framework
- 用语音识别+检索增强生成技术,实现对老人和照护类别的精准识别
- 提醒识别率达89.09%,任务调度准确率达84.65%,无遗漏提醒
- 适合关注医疗语音系统安全性的研究者与养老机构管理者
人工智能正被探索用于减轻健康与社会护理中的行政负担,使工作人员有更多时间专注于患者照护。本文评估了一款面向养老院日常活动的语音智能音箱,支持通过语音访问居民档案、设置提醒及安排任务。提出一个以安全为核心的端到端评估框架,结合Whisper语音识别与检索增强生成(RAG)方法(混合、稀疏、密集)。通过监督式养老院试验和受控测试,评估了330条语音转录内容,覆盖11个照护类别,其中包含184次含提醒的交互。重点考察:(i)居民与照护类别的正确识别;(ii)提醒的识别与提取;(iii)在不确定情况下的端到端任务调度准确性(包括安全延后/澄清)。鉴于养老院的安全敏感性,特别关注嘈杂环境与多样口音下的可靠性,采用置信度评分、澄清提示与人工介入监督。最佳配置(GPT-5.2)下,居民身份与照护类别匹配率达100%(95%置信区间:98.86–100),提醒识别率为89.09%(95%置信区间:83.81–92.80),召回率100%(无漏检),但存在部分误报。通过日历集成实现端到端调度,提醒数量精确匹配率达84.65%(95%置信区间:78.00–89.56),表明口语指令转为可执行事件仍存边缘案例。结果表明,经严格评估并合理保障的语音系统,可支持精准记录、有效任务管理,并实现养老场景中可信的AI应用。
原文摘要 · Abstract (English)
Artificial intelligence (AI) is increasingly being explored in health and social care to reduce administrative workload and allow staff to spend more time on patient care. This paper evaluates a voice-enabled Care Home Smart Speaker designed to support everyday activities in residential care homes, including spoken access to resident records, reminders, and scheduling tasks. A safety-focused evaluation framework is presented that examines the system end-to-end, combining Whisper-based speech recognition with retrieval-augmented generation (RAG) approaches (hybrid, sparse, and dense). Using supervised care-home trials and controlled testing, we evaluated 330 spoken transcripts across 11 care categories, including 184 reminder-containing interactions. These evaluations focus on (i) correct identification of residents and care categories, (ii) reminder recognition and extraction, and (iii) end-to-end scheduling correctness under uncertainty (including safe deferral/clarification). Given the safety-critical nature of care homes, particular attention is also paid to reliability in noisy environments and across diverse accents, supported by confidence scoring, clarification prompts, and human-in-the-loop oversight. In the best-performing configuration (GPT-5.2), resident ID and care category matching reached 100% (95% CI: 98.86-100), while reminder recognition reached 89.09\% (95% CI: 83.81-92.80) with zero missed reminders (100% recall) but some false positives. End-to-end scheduling via calendar integration achieved 84.65% exact reminder-count agreement (95% CI: 78.00-89.56), indicating remaining edge cases in converting informal spoken instructions into actionable events. The findings suggest that voice-enabled systems, when carefully evaluated and appropriately safeguarded, can support accurate documentation, effective task management, and trustworthy use of AI in care home settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。