评测大模型生成问题的质量,发现其答问更长且覆盖更均等。
Can LLMs Ask Good Questions?
- 对比大模型与人类提问在六维上的表现差异
- 生成问题需更长答案且上下文覆盖更均匀
- 适合研究问答质量或构建自动评测系统的人看
我们评估了大语言模型(LLMs)从上下文中生成的问题,将其与人工撰写的问题在六个维度上进行比较:问题类型、长度、上下文覆盖度、可回答性、罕见性以及所需答案长度。研究涵盖两个开源和两个专有状态领先模型。结果表明,大模型生成的问题通常需要更长的描述性回答,并表现出更均匀的上下文关注分布,与传统问答任务中常见的位置偏倚形成对比。这些发现揭示了大模型生成问题的独特特征,为未来问题质量研究及下游应用提供了参考。
原文摘要 · Abstract (English)
We evaluate questions generated by large language models (LLMs) from context, comparing them to human-authored questions across six dimensions: question type, question length, context coverage, answerability, uncommonness, and required answer length. Our study spans two open-source and two proprietary state-of-the-art models. Results reveal that LLM-generated questions tend to demand longer descriptive answers and exhibit more evenly distributed context focus, in contrast to the positional bias often seen in QA tasks. These findings provide insights into the distinctive characteristics of LLM-generated questions and inform future work on question quality and downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。