arXiv:2512.16530cs.CLcs.AI2025-12被引 1

用大模型简化医学文本,评估不同方法的可读性效果

Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics

  • 对比提示模板、双智能体和微调三种简化策略
  • gpt-4o-mini在可读性指标上表现最佳,微调方法效果较差
  • 基于大模型的G-Eval评分与人工评价结果高度一致

本研究探讨了使用大语言模型(LLMs)简化生物医学文本以提升健康素养的可行性。基于公开数据集,包含生物医学摘要的通俗化版本,我们开发并评估了三种方法:基于提示模板的基线方法、双智能体协作方法以及微调方法。选用OpenAI的gpt-4o和gpt-4o mini作为基准模型。采用定量指标如Flesch-Kincaid年级水平、SMOG指数、SARI、BERTScore、G-Eval,以及定性指标(5点李克特量表)评估简洁性、准确性、完整性和简洁性。结果显示,gpt-4o-mini表现最优,微调方法表现较差。基于大模型的G-Eval指标在排序上与人工评价高度一致,显示出良好潜力。

原文摘要 · Abstract (English)

This study investigated the application of Large Language Models (LLMs) for simplifying biomedical texts to enhance health literacy. Using a public dataset, which included plain language adaptations of biomedical abstracts, we developed and evaluated several approaches, specifically a baseline approach using a prompt template, a two AI agent approach, and a fine-tuning approach. We selected OpenAI gpt-4o and gpt-4o mini models as baselines for further research. We evaluated our approaches with quantitative metrics, such as Flesch-Kincaid grade level, SMOG Index, SARI, and BERTScore, G-Eval, as well as with qualitative metric, more precisely 5-point Likert scales for simplicity, accuracy, completeness, brevity. Results showed a superior performance of gpt-4o-mini and an underperformance of FT approaches. G-Eval, a LLM based quantitative metric, showed promising results, ranking the approaches similarly as the qualitative metric.

文本简化大模型评估健康素养

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。