让论文表格生成更懂研究意图,提升对比分析效率
Intent-Aware Schema Generation And Refinement For Literature Review Tables
- 通过合成研究意图构建新数据集,解决评估模糊问题
- 小模型微调后性能媲美顶级大模型,降低使用成本
- 引入大模型辅助优化,显著提升表格结构质量
学术文献数量激增,研究人员亟需高效组织、比较和对比文献。大型语言模型(LLMs)可生成定义共性维度的表格结构(schema),以支持论文间比较。然而,现有研究在schema生成方面进展缓慢,主要受限于:(i) 基于参考的评估存在歧义,(ii) 缺乏有效的编辑与优化方法。本文首次同时解决这两项挑战。首先,我们提出一种方法,为未标注的表格语料库添加合成的研究意图(synthesized intents),并基于此构建一个面向特定信息需求的schema生成研究数据集,有效降低评估歧义。在此数据集上,我们证明引入表格意图能显著提升基线方法重建参考schema的性能。接着,我们系统评测多种单次生成的schema生成方法,包括提示工程的LLM流程与微调模型,发现小型开放权重模型经微调后可达到与顶尖提示式LLM相当的性能。最后,我们提出若干基于LLM的schema优化技术,并验证其能进一步提升生成结果的质量。
原文摘要 · Abstract (English)
The increasing volume of academic literature makes it essential for researchers to organize, compare, and contrast collections of documents. Large language models (LLMs) can support this process by generating schemas defining shared aspects along which to compare papers. However, progress on schema generation has been slow due to: (i) ambiguity in reference-based evaluations, and (ii) lack of editing/refinement methods. Our work is the first to address both issues. First, we present an approach for augmenting unannotated table corpora with \emph{synthesized intents}, and apply it to create a dataset for studying schema generation conditioned on a given information need, thus reducing ambiguity. With this dataset, we show how incorporating table intents significantly improves baseline performance in reconstructing reference schemas. We start by comprehensively benchmarking several single-shot schema generation methods, including prompted LLM workflows and fine-tuned models, showing that smaller, open-weight models can be fine-tuned to be competitive with state-of-the-art prompted LLMs. Next, we propose several LLM-based schema refinement techniques and show that these can further improve schemas generated by these methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。