首个医学视频生成数据集+模型,解决医疗内容不准问题
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
- 构建5.5万条标注医学视频数据集,覆盖真实临床场景
- 模型在视觉质量与医学准确性上超越开源模型,媲美商业系统
- 适合医学教育、模拟训练等需要高精度视频的场景
近期视频生成技术在开放领域取得显著进展,但医学视频生成仍鲜有探索。医学视频对临床培训、教育和模拟至关重要,不仅需高视觉保真度,更要求严格的医学准确性。当前模型在处理医学提示时常生成不真实或错误内容,主要因缺乏大规模、高质量的医学专用数据集。为此,我们提出MedVideoCap-55K,首个大规模、多样化且带丰富字幕的医学视频数据集,包含超过55,000个精心筛选的片段,覆盖真实世界医疗场景,为通用医学视频生成模型训练提供坚实基础。基于此数据集,我们开发了MedGen,其在多个基准测试中表现领先于开源模型,视觉质量与医学准确性均达到与商业系统相当水平。我们希望该数据集与模型能成为推动医学视频生成研究的重要资源。代码与数据已公开于https://github.com/FreedomIntelligence/MedGen。
原文摘要 · Abstract (English)
Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and simulation, requiring not only high visual fidelity but also strict medical accuracy. However, current models often produce unrealistic or erroneous content when applied to medical prompts, largely due to the lack of large-scale, high-quality datasets tailored to the medical domain. To address this gap, we introduce MedVideoCap-55K, the first large-scale, diverse, and caption-rich dataset for medical video generation. It comprises over 55,000 curated clips spanning real-world medical scenarios, providing a strong foundation for training generalist medical video generation models. Built upon this dataset, we develop MedGen, which achieves leading performance among open-source models and rivals commercial systems across multiple benchmarks in both visual quality and medical accuracy. We hope our dataset and model can serve as a valuable resource and help catalyze further research in medical video generation. Our code and data is available at https://github.com/FreedomIntelligence/MedGen
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。