arXiv:2505.17565cs.CL2025-05中稿 · AACL-IJCNLP被引 1

通过自动化构建过程偏好数据,提升表格问答模型性能与效率。

p2-TQA: A Process-based Preference Learning Framework for Self-Improving Table Question Answering Models

  • 基于表格特性的自动化管道生成偏好数据,无需人工标注。
  • 仅用8000样本使模型在域内提升5%、域外提升2.4%。
  • 相比大型模型更高效,适合资源受限场景下的模型优化。

表格问答(TQA)旨在基于表格数据回答问题,涉及单元格检索与数据分析等任务。尽管近期研究通过微调改进了TQA系统,但现有方法常未能充分利用可用数据,且忽视了后训练阶段的提升潜力。本文提出p2-TQA,一种面向TQA后训练的过程型偏好学习框架。该框架通过表格特定的自动化流程生成过程偏好数据,避免了手动或高成本的数据收集。随后在收集的数据上采用对比学习优化模型。实验表明,p2-TQA在域内数据集上可使模型性能提升最高5%,在域外数据集上提升2.4%,仅需8000个训练实例。此外,经p2-TQA增强的模型在性能上媲美更大更复杂的先进TQA系统,同时效率最高可达其五倍。

原文摘要 · Abstract (English)

Table question answering (TQA) focuses on answering questions based on tabular data. Developing TQA systems targets effective interaction with tabular data for tasks such as cell retrieval and data analysis. While recent work has leveraged fine-tuning to improve TQA systems, existing approaches often under-utilize available data and neglect the potential of post-training for further gains. In this work, we introduce p2-TQA, a process-based preference learning framework for TQA post-training. p2-TQA automatically constructs process-based preference data via a table-specific pipeline, eliminating the need for manual or costly data collection. It then optimizes models through contrastive learning on the collected data. Experiments show that p2-TQA effectively improves TQA models by up to 5% on in-domain datasets and 2.4% on out-of-domain datasets with only 8,000 training instances. Furthermore, models enhanced with p2-TQA achieve competitive results against larger, more complex state-of-the-art TQA systems, while maintaining up to five times higher efficiency.

表格问答偏好学习后训练高效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。