用对抗生成问题提升专业领域大模型的推理能力。
Agentic Adversarial QA for Improving Domain-Specific LLMs
- 通过对比待调优模型与专家模型输出,迭代生成挑战性问题。
- 仅用少量合成样本就实现更高准确率,样本效率显著提升。
- 适合需要高效适配专业领域的大模型研究者使用。
大语言模型虽在广泛互联网语料上预训练,但在专业化领域适应能力仍有限。现有微调方法受限于高质量、任务相关数据的稀缺与覆盖不足。当前合成数据生成方法如改写或知识抽取虽能提升事实记忆和概念理解,却难以支持领域内解释性推理,且常生成冗余庞大的数据集,导致样本效率低下。为此,我们提出一种对抗式问题生成框架,通过迭代反馈机制比较待适配模型与基于参考文档的稳健专家模型输出,生成紧凑而语义挑战性强的问题,以揭示并填补理解差距。在LegalBench专业子集上的评估表明,该方法在显著减少合成样本数量的同时,实现了更高的准确性。
原文摘要 · Abstract (English)
Large Language Models (LLMs), despite extensive pretraining on broad internet corpora, often struggle to adapt effectively to specialized domains. There is growing interest in fine-tuning these models for such domains; however, progress is constrained by the scarcity and limited coverage of high-quality, task-relevant data. To address this, synthetic data generation methods such as paraphrasing or knowledge extraction are commonly applied. Although these approaches excel at factual recall and conceptual knowledge, they suffer from two critical shortcomings: (i) they provide minimal support for interpretive reasoning capabilities in these specialized domains, and (ii) they often produce synthetic corpora that are excessively large and redundant, resulting in poor sample efficiency. To overcome these gaps, we propose an adversarial question-generation framework that produces a compact set of semantically challenging questions. These questions are constructed by comparing the outputs of the model to be adapted and a robust expert model grounded in reference documents, using an iterative, feedback-driven process designed to reveal and address comprehension gaps. Evaluation on specialized subsets of the LegalBench corpus demonstrates that our method achieves greater accuracy with substantially fewer synthetic samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。