大模型能自己出常识推理题,且答得好就写得棒
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
- 用大模型生成类COPA风格的常识测验题
- 能答对COPA的模型更擅长自主出题
- 适合研究大模型认知能力与创作能力
大语言模型如今已能完成多种复杂写作任务,并在自然语言推理与常识推理问答中表现优异。本文探讨将大模型视为常识测评题作者的可能性。我们通过提示大模型生成类似著名常识推理基准COPA(Choice of Plausible Alternatives)风格的题目,并借助大模型自身分析与人工标注进行评估。结果表明,能在原始COPA测试中表现良好的模型,在生成自己的测评题时也更为成功。
原文摘要 · Abstract (English)
LLMs can now perform a variety of complex writing tasks. They also excel in answering questions pertaining to natural language inference and commonsense reasoning. Composing these questions is itself a skilled writing task, so in this paper we consider LLMs as authors of commonsense assessment items. We prompt LLMs to generate items in the style of a prominent benchmark for commonsense reasoning, the Choice of Plausible Alternatives (COPA). We examine the outcome according to analyses facilitated by the LLMs and human annotation. We find that LLMs that succeed in answering the original COPA benchmark are also more successful in authoring their own items.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。