arXiv:2508.03723eess.IVcs.CV2025-08

构建可自动更新的临床影像数据采集框架,保障AI训练数据时效性与代表性。

Technical specification of a framework for the collection of clinical images and data

  • 设计自动化持续采集系统,确保数据反映当前临床实践。
  • 强调需包含长期随访的旧病例,提升模型验证准确性。
  • 适合医疗AI研发团队及数据治理机构参考落地。

本报告描述了一个用于训练和验证人工智能(AI)工具的临床影像与数据采集框架。不仅涵盖图像与临床数据的收集方法,还包含伦理审查与信息治理流程,确保数据安全采集,并明确了数据共享所需的基础设施与协议。该框架的核心特点是支持自动化、持续的数据采集,以保证数据的实时性与代表性。这对于训练和验证AI工具至关重要——数据需包含具有长期随访的旧病例,以准确评估临床结局;同时,验证数据应反映当前实际诊疗情况,避免使用过时数据导致误判。报告还介绍了非全自动采集方案,适用于初期或资源有限场景,提供逐步过渡的方法指导。

原文摘要 · Abstract (English)

In this report a framework for the collection of clinical images and data for use when training and validating artificial intelligence (AI) tools is described. The report contains not only information about the collection of the images and clinical data, but the ethics and information governance processes to consider ensuring the data is collected safely, and the infrastructure and agreements required to allow for the sharing of data with other groups. A key characteristic of the main collection framework described here is that it can enable automated and ongoing collection of datasets to ensure that the data is up-to-date and representative of current practice. This is important in the context of training and validating AI tools as it is vital that datasets have a mix of older cases with long term follow-up such that the clinical outcome is as accurate as possible, and current data. Validations run on old data will provide findings and conclusions relative to the status of the imaging units when that data was generated. It is important that a validation dataset can assess the AI tools with data that it would see if deployed and active now. Other types of collection frameworks, which do not follow a fully automated approach, are also described. Whilst the fully automated method is recommended for large scale, long-term image collection, there may be reasons to start data collection using semi-automated methods and indications of how to do that are provided.

医疗AI数据采集自动化临床数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。