用声音和临床信号评估哮喘风险,全程可审计且安全可靠。
AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

- 结合声谱分析与轻量级机器学习,自动筛查呼吸音异常。
- 五阶段LLM提示链实现临床推理与分级预警,准确率达77.4%。
- 适合医疗研究者、急症辅助诊断系统开发者参考。
急性哮喘风险评估需快速解读呼吸音、血氧、通气限制、言语能力、呼吸努力程度、意识状态及对缓解药物的反应。传统仅依赖音频的分类器虽能检测喘息模式,但缺乏透明的临床推理和安全升级逻辑。本文提出AeroSpectra Sentinel,一个客户端研究原型与决策支持工作流,整合短时傅里叶变换(STFT)呼吸音分析、轻量级机器学习筛查、临床特征融合及五阶段大语言模型(LLM)提示链流程。该流程分离了信号采集、预处理、声学特征提取、机器学习筛查、临床护栏与FHIR就绪报告。在包含1,211个WAV录音的公开呼吸音数据集上评估音频筛查组件,使用584个分层子集,随机森林在哮喘vs非哮喘二分类中达到91.10%准确率和78.69% F1分数;基于特征的多层感知机达89.73%准确率和78.26% F1分数;紧凑对数谱图卷积网络达73.29%准确率和55.17% F1分数。多分类任务准确率为77.40%,宏平均F1为77.23%。通过40个模拟临床案例的场景审计,对比单次提示、提示链、带护栏提示链、带护栏加FHIR schema验证提示链,结果显示带护栏与schema验证的变体在模拟安全性与文档一致性上表现最佳。AeroSpectra Sentinel为研究原型,非诊断医疗器械或临床验证产品。
原文摘要 · Abstract (English)
Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy. Conventional audio-only classifiers can detect wheeze-like patterns but often lack transparent clinical reasoning and safe escalation logic. This paper presents AeroSpectra Sentinel, a client-side research prototype and decision-support workflow that combines short-time Fourier transform (STFT) respiratory sound analysis, lightweight machine-learning screening, clinical feature fusion, and a five-stage large language model (LLM) prompt-chaining process. The workflow separates signal acquisition, preprocessing, acoustic feature extraction, ML screening, clinical guardrails, and FHIR-ready reporting. We evaluated the audio screening component on a public respiratory sound dataset containing 1,211 WAV recordings from five labels. Using a stratified subset of 584 recordings, a random forest achieved 91.10% binary accuracy and 78.69% F1-score for asthma-vs-non-asthma screening, while a feature-based multilayer perceptron achieved 89.73% accuracy and 78.26% F1-score. A compact log-spectrogram CNN achieved 73.29% accuracy and 55.17% F1-score. Multiclass classification achieved 77.40% accuracy and 77.23% macro-F1. To evaluate the LLM workflow, we conducted a scenario-based audit on 40 simulated clinical vignettes comparing one-shot prompting, prompt chaining, prompt chaining with guardrails, and prompt chaining with guardrails plus FHIR schema validation. The guardrail-plus-schema variant achieved the strongest simulated safety and documentation consistency. AeroSpectra Sentinel is intended as a research prototype, not as a diagnostic medical device or clinically validated risk-assessment product.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。