构建首个巴西葡萄牙语听证会长文摘要数据集
PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents
- 收集巴西众议院听证会原文与配套新闻、结构化摘要
- 提供标注的自然语言推理数据,支持幻觉检测评估
- 适合研究葡语长文档摘要与大模型幻觉问题的学者
本文提出PublicHearingBR,一个用于葡萄牙语长文档摘要的巴西葡萄牙语数据集。该数据集包含巴西众议院听证会的原始记录,以及对应的新闻报道和结构化摘要,涵盖参会人员及其陈述或观点。数据集旨在支持葡萄牙语长文档摘要系统的开发与评估。研究贡献包括数据集本身、一个混合摘要系统作为未来研究基线,以及针对大语言模型生成摘要中幻觉问题的评估指标讨论。基于此讨论,数据集还包含用于评估自然语言推断任务的标注数据。
原文摘要 · Abstract (English)
This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Brazilian Chamber of Deputies, paired with news articles and structured summaries containing the individuals participating in the hearing and their statements or opinions. The dataset supports the development and evaluation of long document summarization systems in Portuguese. Our contributions include the dataset, a hybrid summarization system to establish a baseline for future studies, and a discussion of evaluation metrics for summarization involving large language models, addressing the challenge of hallucination in the generated summaries. As a result of this discussion, the dataset also includes annotated data to evaluate natural language inference tasks in Portuguese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。