arXiv:2510.23008cs.AI2025-10

构建多维可信度评估框架,提升中文大模型生成肝癌影像报告的可靠性

From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer

  • 提出多维度可信度评估框架,指导不同机构优化提示词设计
  • 在硅流平台测试多个大模型,验证框架对报告可信度的量化评估能力
  • 适用于医学影像报告生成、临床质量控制及医学生培训场景

大型语言模型(LLMs)在根据影像发现生成诊断结论方面展现出良好性能,有助于放射科报告撰写、实习生教学和质量控制。然而,如何在不同临床情境下优化提示词设计仍缺乏系统性指导。此外,尚未建立评估大模型生成放射科报告可信度的全面标准化框架。本研究旨在通过引入多维度可信度评估(MDCA)框架,并提供机构定制化的提示词优化建议,提升大模型生成肝磁共振报告的可信度。该框架被应用于评估和比较多个先进大模型的表现,包括Kimi-K2-Instruct-0905、Qwen3-235B-A22B-Instruct-2507、DeepSeek-V3和ByteDance-Seed-OSS-36B-Instruct,实验基于硅流(SiliconFlow)平台进行。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated promising performance in generating diagnostic conclusions from imaging findings, thereby supporting radiology reporting, trainee education, and quality control. However, systematic guidance on how to optimize prompt design across different clinical contexts remains underexplored. Moreover, a comprehensive and standardized framework for assessing the trustworthiness of LLM-generated radiology reports is yet to be established. This study aims to enhance the trustworthiness of LLM-generated liver MRI reports by introducing a Multi-Dimensional Credibility Assessment (MDCA) framework and providing guidance on institution-specific prompt optimization. The proposed framework is applied to evaluate and compare the performance of several advanced LLMs, including Kimi-K2-Instruct-0905, Qwen3-235B-A22B-Instruct-2507, DeepSeek-V3, and ByteDance-Seed-OSS-36B-Instruct, using the SiliconFlow platform.

医学影像大模型可信度肝癌诊断提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。