arXiv:2501.03670cs.CLcs.AI2025-01被引 18

提升数学应用题解法多样性,让模型更灵活应对不同题型。

A Diversity-Enhanced Knowledge Distillation Model for Practical Math Word Problem Solving

  • 用自适应知识蒸馏让学生模型学出多种解题路径。
  • 在四个数据集上准确率超越强基线,且推理效率高。
  • 适合需要多样化解法的教育类AI系统使用。

数学应用题求解是自然语言处理中的关键任务,近年来受到广泛关注。现有研究多依赖序列到序列及其扩展模型(如Seq2Tree、Graph2Tree)生成数学方程,虽有效但难以生成多样且合理的解题方程,限制了其在不同题型中的泛化能力。本文提出一种新型多样性增强知识蒸馏模型(DivKD),通过自适应多样性蒸馏方法,使学生模型有选择地从教师模型中学习高质量知识,从而生成多样化的方程。此外,设计了带多样性先验增强的学生模型,结合条件变分自编码器以更好捕捉方程的分布特性。在四个数学应用题基准数据集上的实验表明,该方法在保持高效率的同时,显著提升了答案准确率,优于多个强基线模型。

原文摘要 · Abstract (English)

Math Word Problem (MWP) solving is a critical task in natural language processing, has garnered significant research interest in recent years. Various recent studies heavily rely on Seq2Seq models and their extensions (e.g., Seq2Tree and Graph2Tree) to generate mathematical equations. While effective, these models struggle to generate diverse but counterpart solution equations, limiting their generalization across various math problem scenarios. In this paper, we introduce a novel Diversity-enhanced Knowledge Distillation (DivKD) model for practical MWP solving. Our approach proposes an adaptive diversity distillation method, in which a student model learns diverse equations by selectively transferring high-quality knowledge from a teacher model. Additionally, we design a diversity prior-enhanced student model to better capture the diversity distribution of equations by incorporating a conditional variational auto-encoder. Extensive experiments on {four} MWP benchmark datasets demonstrate that our approach achieves higher answer accuracy than strong baselines while maintaining high efficiency for practical applications.

数学题求解知识蒸馏多样性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。