用AI和智能抽样,让医疗数据验证快40%且省下77%工作量。
A chart review process aided by natural language processing and multi-wave adaptive sampling to expedite validation of code-based algorithms for large database studies
- 用自然语言处理减少人工审阅每份病历时间
- 多轮自适应采样使77%病历无需审查,精度损失小
- 适合做大规模医疗数据库研究的算法验证
为提升大型医保数据库研究中基于编码算法的测量可靠性,需通过人工审阅电子病历文本创建金标准标签,但耗时费力。本文提出一种加速验证流程:一是利用自然语言处理(NLP)降低人工审阅每份病历的时间;二是采用多轮自适应采样并设定预定义停止规则,在性能指标达到足够精度后终止验证。在肥胖患者自伤事件算法验证案例中,实验表明,NLP辅助使每份病历审阅时间减少40%,多轮采样与停止规则可避免77%的病历被审查,对测量特征精度影响有限。该方法有助于更频繁地验证代码算法,提升数据库研究结果的可信度。
原文摘要 · Abstract (English)
Background: One of the ways to enhance analyses conducted with large claims databases is by validating the measurement characteristics of code-based algorithms used to identify health outcomes or other key study parameters of interest. These metrics can be used in quantitative bias analyses to assess the robustness of results for an inferential study given potential bias from outcome misclassification. However, extensive time and resource allocation are typically re-quired to create reference-standard labels through manual chart review of free-text notes from linked electronic health records. Methods: We describe an expedited process that introduces efficiency in a validation study us-ing two distinct mechanisms: 1) use of natural language processing (NLP) to reduce time spent by human reviewers to review each chart, and 2) a multi-wave adaptive sampling approach with pre-defined criteria to stop the validation study once performance characteristics are identified with sufficient precision. We illustrate this process in a case study that validates the performance of a claims-based outcome algorithm for intentional self-harm in patients with obesity. Results: We empirically demonstrate that the NLP-assisted annotation process reduced the time spent on review per chart by 40% and use of the pre-defined stopping rule with multi-wave samples would have prevented review of 77% of patient charts with limited compromise to precision in derived measurement characteristics. Conclusion: This approach could facilitate more routine validation of code-based algorithms used to define key study parameters, ultimately enhancing understanding of the reliability of find-ings derived from database studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。