arXiv:2601.11581cs.CLcs.AI2026-01

通过多领域去偏框架提升QA模型在复杂查询下的准确性

Enhancing the QA Model through a Multi-domain Debiasing Framework

  • 结合知识蒸馏与去偏技术,扩展至多领域数据增强模型鲁棒性
  • 在SQuAD和对抗数据集上,精确匹配和F1分数最高提升2.6个百分点
  • 适合关注模型公平性与真实场景泛化能力的研究者

问答模型在机器阅读理解方面已取得显著进展,但在复杂查询的对抗性条件下常表现出偏差,影响性能。本研究在SQuAD v1.1及对抗数据集AddSent和AddOneSent上评估ELECTRA-small模型,识别出词汇偏见、数值推理和实体识别相关错误,并提出融合知识蒸馏、去偏技术和领域扩展的多领域去偏框架。实验结果表明,所有测试集上的精确匹配(EM)和F1分数最高提升2.6个百分点,尤其在对抗情境下表现更优。研究证明针对性的去偏策略能有效提升自然语言理解系统的鲁棒性与可靠性。

原文摘要 · Abstract (English)

Question-answering (QA) models have advanced significantly in machine reading comprehension but often exhibit biases that hinder their performance, particularly with complex queries in adversarial conditions. This study evaluates the ELECTRA-small model on the Stanford Question Answering Dataset (SQuAD) v1.1 and adversarial datasets AddSent and AddOneSent. By identifying errors related to lexical bias, numerical reasoning, and entity recognition, we develop a multi-domain debiasing framework incorporating knowledge distillation, debiasing techniques, and domain expansion. Our results demonstrate up to 2.6 percentage point improvements in Exact Match (EM) and F1 scores across all test sets, with gains in adversarial contexts. These findings highlight the potential of targeted bias mitigation strategies to enhance the robustness and reliability of natural language understanding systems.

问答系统去偏鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。