用少样本+分类框架生成专业领域高质量问答对,效率翻倍且覆盖全面。
ExpertGenQA: Open-ended QA generation in Specialized Domains
- 结合少样本学习与主题风格分类生成问答对
- 效率提升一倍,主题覆盖率94.4%,优于基线方法
- 生成的问答对更贴近专家思维层次,适合技术场景应用
在专业领域生成高质量问答对仍具挑战,现有方法难以兼顾专家样例利用与主题多样性。本文提出ExpertGenQA,通过少样本学习结合结构化主题与风格分类,生成全面的领域特定问答对。以美国联邦铁路管理局文档为测试集,结果表明该方法效率是基线少样本方法的两倍,同时保持94.4%的主题覆盖率。系统评估显示,当前基于LLM的评判模型和奖励模型对表面写作风格存在显著偏倚。基于布卢姆认知分类学的分析表明,ExpertGenQA比模板方法更有效保留专家问题的认知复杂性分布。用于训练检索模型时,生成查询使顶1准确率相比基线提升13.02%,证明其在技术领域下游任务中的有效性。
原文摘要 · Abstract (English)
Generating high-quality question-answer pairs for specialized technical domains remains challenging, with existing approaches facing a tradeoff between leveraging expert examples and achieving topical diversity. We present ExpertGenQA, a protocol that combines few-shot learning with structured topic and style categorization to generate comprehensive domain-specific QA pairs. Using U.S. Federal Railroad Administration documents as a test bed, we demonstrate that ExpertGenQA achieves twice the efficiency of baseline few-shot approaches while maintaining $94.4\%$ topic coverage. Through systematic evaluation, we show that current LLM-based judges and reward models exhibit strong bias toward superficial writing styles rather than content quality. Our analysis using Bloom's Taxonomy reveals that ExpertGenQA better preserves the cognitive complexity distribution of expert-written questions compared to template-based approaches. When used to train retrieval models, our generated queries improve top-1 accuracy by $13.02\%$ over baseline performance, demonstrating their effectiveness for downstream applications in technical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。