对比Mistral与QWen在医学文本简化中的可读性与准确性策略差异。
Making Knowledge Accessible: Divergent Readability-Accuracy Strategies of Mistral and QWen in Biomedical Text Simplification
- Mistral采用温和词汇简化,提升可读性同时保持语义连贯性。
- QWen可读性提升但准确率略降,存在可读性与准确性的失衡。
- 21项指标相关性分析揭示度量冗余,指导后续评估优化。
公众对可及的生物医学信息需求日益增长,亟需可扩展的文本简化方法。尽管大语言模型(LLMs)提供了潜在解决方案,但在提升可读性与保持语义准确性之间仍面临挑战。本报告实证比较了两种指令微调模型——Mistral-Small 3 24B与增强推理能力的QWen2.5 32B——在生物医学文本简化任务中的表现,并与人类水平进行基准对比。分析显示,两模型采用不同策略:Mistral采取稳健的词汇简化路径,在多项指标上持续提升可读性,同时保持较高语篇忠实度(BERTScore: 0.91,与人类水平无统计显著差异);相比之下,QWen虽也实现可读性提升和合理准确率(BERTScore: 0.89),但在可读性与准确性间存在权衡失衡。进一步对21项度量的综合相关性分析揭示了度量间的强功能冗余,为评估体系的适应性改进提供依据。
原文摘要 · Abstract (English)
The growing public demand for accessible biomedical information calls for scalable text simplification. While large language models (LLMs) offer solutions, they too struggle with balancing improved readability against preservation of meaning. This report empirically compares how two LLMs - instruction-tuned Mistral-Small 3 24B and the reasoning-augmented QWen2.5 32B- navigate this trade-off in biomedical text simplification, benchmarked against human performance. Our analysis highlights how each model applies distinct operational strategies when simplifying biomedical text. Mistral exhibits a tempered lexical simplification approach that consistently enhances readability across multiple metrics while preserving discourse fidelity (BERTScore: 0.91, statistically comparable to that of humans). In comparison, QWen also attains enhanced readability performance and a reasonable BERTScore of 0.89, but presents a disconnect in balancing between readability and accuracy. Additionally, a comprehensive correlation analysis of a suite of 21 metrics confirms strong functional redundancies in metrics and informs adaptation requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。