用1.15亿次真实医患对话优化AI,让医疗AI更安全、更人性。
Perfecting Human-AI Interaction at Clinical Scale. Turning Production Signals into Safer, More Human Conversations
- 基于真实患者交互信号构建实时优化框架
- 临床安全分达99.9%,语音识别错误率降低50%
- 适合关注医疗AI落地安全与体验的从业者
医疗对话AI不应仅追求基准测试准确率,而应面向真实患者对话场景:语音不清晰、意图间接、语言中途变化、合规性取决于沟通方式。本文基于1.15亿+真实患者-AI互动及7000+持证医生参与的50万+测试通话,提炼出实时信号——声调特征、对话轮次动态、澄清触发、升级标记、多语言连贯性与流程确认等,揭示了数据集未覆盖的失效模式,并提供可操作的训练与评估信号。研究显示,医疗级安全无法依赖单一大模型,需通过受控编排、独立校验与验证实现冗余;许多“推理”错误源于上游环节,需在上下文语音识别、澄清修复、环境语音处理及低延迟模型/硬件选择上实现垂直整合。将交互智能(语气、节奏、共情、澄清、轮次)作为首要安全变量,显著提升安全性、文档质量、任务完成率与公平性。部署于超千万真实患者通话中,Polaris系统取得99.9%临床安全分,平均患者评分8.95,较企业级语音识别错误率降低50%。结果证实,真实交互智能是患者导向医疗AI系统安全与可靠性的关键决定因素。
原文摘要 · Abstract (English)
Healthcare conversational AI agents shouldn't be optimized only for clean benchmark accuracy in production-first regime; they must be optimized for the lived reality of patient conversations, where audio is imperfect, intent is indirect, language shifts mid-call, and compliance hinges on how guidance is delivered. We present a production-validated framework grounded in real-time signals from 115M+ live patient-AI interactions and clinician-led testing (7K+ licensed clinicians; 500K+ test calls). These in-the-wild cues -- paralinguistics, turn-taking dynamics, clarification triggers, escalation markers, multilingual continuity, and workflow confirmations -- reveal failure modes that curated data misses and provide actionable training and evaluation signals for safety and reliability. We further show why healthcare-grade safety cannot rely on a single LLM: long-horizon dialogue and limited attention demand redundancy via governed orchestration, independent checks, and verification. Many apparent "reasoning" errors originate upstream, motivating vertical integration across contextual ASR, clarification/repair, ambient speech handling, and latency-aware model/hardware choices. Treating interaction intelligence (tone, pacing, empathy, clarification, turn-taking) as first-class safety variables, we drive measurable gains in safety, documentation, task completion, and equity in building the safest generative AI solution for autonomous patient-facing care. Deployed across more than 10 million real patient calls, Polaris attains a clinical safety score of 99.9%, while significantly improving patient experience with average patient rating of 8.95 and reducing ASR errors by 50% over enterprise ASR. These results establish real-world interaction intelligence as a critical -- and previously underexplored -- determinant of safety and reliability in patient-facing clinical AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。