arXiv:2503.02695cs.IR2025-03被引 2

零样本框架让社科研究者一键问答长篇论文中的复杂问题。

Zero-Shot Complex Question-Answering on Long Scientific Documents

  • 用预训练模型融合抽取与生成,处理多跨度、多跳推理等难题。
  • 在新数据集MLPsych上表现强劲,无需微调即可应对复杂问答。
  • 适合无机器学习背景的社科研究者快速解析长篇文献。

随着基于Transformer的语言模型快速发展,针对短文档和简单问题的阅读理解任务已基本解决。然而,知识密集型的长篇科学文档(尤其是社会科学领域)及其复杂的、更接近真实场景的问题仍鲜有研究。本文提出一个零样本管道框架,使社会科学研究者无需机器学习知识,即可对完整研究论文中预设格式的复杂问题进行问答。该方法结合预训练语言模型,应对多跨度提取、多跳推理和长答案生成等挑战。在社会心理学论文构建的新数据集MLPsych上评估,通过抽取与生成模型的组合,框架展现出强大性能。本工作提升了社会科学领域的文档理解能力,同时为研究者提供了实用工具。

原文摘要 · Abstract (English)

With the rapid development in Transformer-based language models, the reading comprehension tasks on short documents and simple questions have been largely addressed. Long documents, specifically the scientific documents that are densely packed with knowledge discovered and developed by humans, remain relatively unexplored. These documents often come with a set of complex and more realistic questions, adding to their complexity. We present a zero-shot pipeline framework that enables social science researchers to perform question-answering tasks that are complex yet of predetermined question formats on full-length research papers without requiring machine learning expertise. Our approach integrates pre-trained language models to handle challenging scenarios including multi-span extraction, multi-hop reasoning, and long-answer generation. Evaluating on MLPsych, a novel dataset of social psychology papers with annotated complex questions, we demonstrate that our framework achieves strong performance through combination of extractive and generative models. This work advances document understanding capabilities for social sciences while providing practical tools for researchers.

问答系统长文档零样本社会科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。