arXiv:2509.19329cs.CLstat.ME2025-09中稿 · NCME AIME 2025被引 2
研究大模型大小、温度与提示风格如何影响评分一致性
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
- 对比不同模型规模、温度和提示风格的评分表现
- 模型大小对人机评分一致性影响最显著
- 强调需多维度评估模型评分对齐效果
我们研究了模型规模、温度和提示风格如何影响大型语言模型(LLMs)在评估临床推理能力时,自身内部、模型之间以及与人类之间的评分对齐情况。结果显示,模型规模是影响LLM与人类评分一致性的关键因素。研究强调了在多个层面检查对齐性的重要性。
原文摘要 · Abstract (English)
We examined how model size, temperature, and prompt style affect Large Language Models' (LLMs) alignment within itself, between models, and with human in assessing clinical reasoning skills. Model size emerged as a key factor in LLM-human score alignment. Study highlights the importance of checking alignments across multiple levels.
大模型评估评分对齐提示工程
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。