arXiv:2604.17574cs.CL2026-04

用上下文学习让大模型生成有推理过程的干扰项。

Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation

论文配图:Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
图 1 · 摘自论文原文
  • 通过少样本示例和语义检索,引导大模型生成带推理过程的干扰项。
  • 在六个数据集上表现超越现有方法,平均提升超过15%。
  • 适合需要高质量多选题的教育评测与自动命题场景。

干扰项生成(DG)仍依赖领域专家,是一项劳动密集型任务。其目标是为多选题生成看似合理但错误的选项。可靠的干扰项需与问题上下文相关,并能通过隐含推理误导答题者。尽管已有方法结合预训练编码器-解码器模型与对比学习生成语义相关的干扰项,但仍难以捕捉专家在选择干扰项时的深层推理过程。本文探索大语言模型(LLMs)通过上下文学习生成具有推理过程的干扰项,采用无监督语义检索选取少样本示例。我们设计了一个理性增强的干扰项生成框架,可联合生成干扰项及其推理理由。在六个不同领域、平均干扰项长度各异的基准数据集上进行的大量实验表明,使用少样本提示显著提升了性能,优于近期方法,在生成符合人类标注基准的有理据干扰项方面达到当前最优水平。

原文摘要 · Abstract (English)

Distractor generation (DG) remains a labor-intensive task that still significantly depends on domain experts. The task focuses on generating plausible yet incorrect options, known as distractors, for multiple-choice questions. A reliable distractor must be contextually relevant to the question and able to mislead examinees through implicit reasoning when identifying the correct answer. While a recent method integrates fine-tuning pre-trained encoder-decoder models with contrastive learning to generate semantically relevant distractors for a given question-answer, it often fails to capture the underlying reasoning process that experts utilize when selecting distractors in benchmarks. In this paper, we explore large language models (LLMs) reasoning for DG through in-context learning with unsupervised semantic retrieval for selecting few-shot examples. We design a rationale-augmented DG framework that jointly generates distractors and their rationales for a given question-answer. Extensive experiments on six benchmarks, with varying average distractor lengths and domains, demonstrate that prompting LLMs with few-shot examples substantially improves the performance compared to recent DG models. It outperforms recent approaches and achieves state-of-the-art results in generating reasoned distractors that align with human-labeled benchmarks.

干扰项生成大模型推理少样本学习多选题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。