自动提取法语新闻中的5W1H信息,效果媲美GPT-4o。
Automated Journalistic Questions: A New Method for Extracting 5W1H in French
- 构建首个法语新闻5W1H自动化提取流程
- 在250篇魁北克新闻上达到GPT-4o水平准确率
- 适合法语信息抽取与新闻自动化处理研究者
5W1H问题(谁、什么、何时、何地、为何、如何)是新闻报道中确保事件描述清晰系统化的常用工具。回答这些问题对于摘要生成、聚类和新闻聚合等任务至关重要。本文设计了首个从法语新闻文章中自动提取5W1H信息的流程。为评估算法性能,我们创建了一个包含250篇魁北克新闻的文章语料库,并由四位人工标注者标记了5W1H答案。结果表明,该流程在该任务上的表现与大型语言模型GPT-4o相当。
原文摘要 · Abstract (English)
The 5W1H questions -- who, what, when, where, why and how -- are commonly used in journalism to ensure that an article describes events clearly and systematically. Answering them is a crucial prerequisites for tasks such as summarization, clustering, and news aggregation. In this paper, we design the first automated extraction pipeline to get 5W1H information from French news articles. To evaluate the performance of our algorithm, we also create a corpus of 250 Quebec news articles with 5W1H answers marked by four human annotators. Our results demonstrate that our pipeline performs as well in this task as the large language model GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。