用大模型简化医学文本,评估不同方法的可读性效果
Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
- 对比提示模板、双智能体和微调三种简化策略
- gpt-4o-mini在可读性指标上表现最佳,微调方法效果较差
- 基于大模型的G-Eval评分与人工评价结果高度一致
本研究探讨了使用大语言模型(LLMs)简化生物医学文本以提升健康素养的可行性。基于公开数据集,包含生物医学摘要的通俗化版本,我们开发并评估了三种方法:基于提示模板的基线方法、双智能体协作方法以及微调方法。选用OpenAI的gpt-4o和gpt-4o mini作为基准模型。采用定量指标如Flesch-Kincaid年级水平、SMOG指数、SARI、BERTScore、G-Eval,以及定性指标(5点李克特量表)评估简洁性、准确性、完整性和简洁性。结果显示,gpt-4o-mini表现最优,微调方法表现较差。基于大模型的G-Eval指标在排序上与人工评价高度一致,显示出良好潜力。
原文摘要 · Abstract (English)
This study investigated the application of Large Language Models (LLMs) for simplifying biomedical texts to enhance health literacy. Using a public dataset, which included plain language adaptations of biomedical abstracts, we developed and evaluated several approaches, specifically a baseline approach using a prompt template, a two AI agent approach, and a fine-tuning approach. We selected OpenAI gpt-4o and gpt-4o mini models as baselines for further research. We evaluated our approaches with quantitative metrics, such as Flesch-Kincaid grade level, SMOG Index, SARI, and BERTScore, G-Eval, as well as with qualitative metric, more precisely 5-point Likert scales for simplicity, accuracy, completeness, brevity. Results showed a superior performance of gpt-4o-mini and an underperformance of FT approaches. G-Eval, a LLM based quantitative metric, showed promising results, ranking the approaches similarly as the qualitative metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。