arXiv:2409.06883cs.CLcs.AI2024-09

构建首个研究问题提取的人类评估数据集,验证现有大模型评估函数效果不佳。

A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task

  • 构建包含论文、GPT-4提取的RQ及多维度人类评价的数据集
  • 发现现有大模型评估函数与人工评价相关性不足
  • 为改进研究问题提取评估方法提供基准数据,适合评测研究者使用

文本摘要技术进展显著,但从高度专业化的文档(如科研论文)中准确提取并总结必要信息的任务仍缺乏深入研究。本文聚焦于从科研论文中提取研究问题(RQ)的任务,构建了一个新数据集,包含机器学习论文、GPT-4提取的RQ,以及从多个角度进行的人类对提取结果的评估。利用该数据集,我们系统比较了近期提出的基于大模型的摘要评估函数,发现它们与人工评价的相关性均未达到足够高的水平。我们期望该数据集能为开发更适配研究问题提取任务的评估函数奠定基础,从而提升该任务的性能。数据集已公开:https://github.com/auto-res/PaperRQ-HumanAnno-Dataset。

原文摘要 · Abstract (English)

The progress in text summarization techniques has been remarkable. However the task of accurately extracting and summarizing necessary information from highly specialized documents such as research papers has not been sufficiently investigated. We are focusing on the task of extracting research questions (RQ) from research papers and construct a new dataset consisting of machine learning papers, RQ extracted from these papers by GPT-4, and human evaluations of the extracted RQ from multiple perspectives. Using this dataset, we systematically compared recently proposed LLM-based evaluation functions for summarizations, and found that none of the functions showed sufficiently high correlations with human evaluations. We expect our dataset provides a foundation for further research on developing better evaluation functions tailored to the RQ extraction task, and contribute to enhance the performance of the task. The dataset is available at https://github.com/auto-res/PaperRQ-HumanAnno-Dataset.

研究问题提取大模型评估数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。