AI辅助睡眠分期可提升评分准确率,且无偏向性。
World of ScoreCraft: Novel Multi Scorer Experiment on the Impact of a Decision Support System in Sleep Staging
- 用在线平台测试医生在AI或人工建议下评分
- 正确建议提升准确率,错误建议则降低准确率
- AI建议与人工建议无明显偏好差异
手动睡眠多导图(PSG)评分耗时且易受评分者间差异影响,可能降低诊断可靠性。本研究探讨决策支持系统(DSS)融入PSG评分流程的影响,重点关注其对准确性、评分时间及对人工智能(AI)推荐相对于人类推荐的潜在偏倚。通过新型在线评分平台,开展重复测量研究,由睡眠技师对传统与自采式PSG进行评分,并偶尔接收标注为人工或AI生成的建议。结果发现,传统PSG评分略高于自采式,但差异不显著;正确建议显著提升两类PSG的评分准确率,错误建议则降低准确率;未观察到对AI建议或人工建议的显著偏好。研究结果表明,AI可有效提升PSG评分可靠性,但确保其输出准确性至关重要。未来研究应探索DSS对评分流程的长期影响及临床整合策略。
原文摘要 · Abstract (English)
Manual scoring of polysomnography (PSG) is a time intensive task, prone to inter scorer variability that can impact diagnostic reliability. This study investigates the integration of decision support systems (DSS) into PSG scoring workflows, focusing on their effects on accuracy, scoring time, and potential biases toward recommendations from artificial intelligence (AI) compared to human generated recommendations. Using a novel online scoring platform, we conducted a repeated measures study with sleep technologists, who scored traditional and self applied PSGs. Participants were occasionally presented with recommendations labeled as either human or AI generated. We found that traditional PSGs tended to be scored slightly more accurately than self applied PSGs, but this difference was not statistically significant. Correct recommendations significantly improved scoring accuracy for both PSG types, while incorrect recommendations reduced accuracy. No significant bias was observed toward or against AI generated recommendations compared to human generated recommendations. These findings highlight the potential of AI to enhance PSG scoring reliability. However, ensuring the accuracy of AI outputs is critical to maximizing its benefits. Future research should explore the long term impacts of DSS on scoring workflows and strategies for integrating AI in clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。