构建可解释的多模态医学问答数据集,支持可控推理复杂度
Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets
- 基于医疗教学视频生成结构化知识图谱
- 实现答案与视觉时间证据的精准对齐
- 支持推理难度可控,适合医学AI研究者使用
数据密集型人工智能应用越来越依赖大规模、高质量、可解释且可复现的数据集,但此类数据集的构建往往仍存在人工成本高、溯源能力弱、配置困难等问题。这一挑战在多模态医学场景中尤为突出,因为每个问答样本需语义一致,并以视觉和时间证据为依据,同时推理复杂度需可调控。为此,我们提出 Med-CRAFT,一个用于从教学视频构建可解释、可配置的多模态医学问答数据集的信息系统。Med-CRAFT 将数据集构建组织为可溯源的流水线,将原始医疗教学视频转化为结构化的操作知识图谱、基于证据的推理路径以及自然语言问答对。
原文摘要 · Abstract (English)
Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets often remains labor-intensive, weakly traceable, and difficult to configure. This problem is particularly critical in multimodal medical scenarios, where each question-answer sample should be semantically consistent, grounded in visual and temporal evidence, and controllable in terms of reasoning complexity. To address these challenges, we propose Med-CRAFT, an information system for explainable and configurable construction of multimodal medical question answering datasets from instructional videos. Med-CRAFT organizes dataset construction as a provenance-aware pipeline that transforms raw medical instructional videos into structured operation knowledge graphs, evidence-grounded reasoning paths, and natural-language question-answer pairs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。