arXiv:2411.19579cs.CL2024-11

识别社交媒体中言论片段,助力假信息核查

ICPR 2024 Competition on Multilingual Claim-Span Identification

  • 基于双语标注数据集,定位文本中的主张片段
  • 使用8000条英/印文帖子训练,提升多语言识别能力
  • 适合关注虚假信息检测与NLP多语言应用的研究者

社交媒体中存在大量主张,可能包含错误信息或假新闻。因此,识别主张是开展验证的第一步。面对海量社交内容,自动化识别主张至关重要。本次竞赛聚焦‘主张片段识别’任务:给定一段文本,需定位其中属于主张的文本片段。该任务比传统的二分类(是否为主张)更具挑战性,需融合模式识别、自然语言处理与机器学习的前沿方法。竞赛采用新构建的数据集HECSI,包含约8000条英文和8000条印地文帖子,均由人工标注主张片段。本文综述了竞赛设置及各参赛团队提出的技术方案。

原文摘要 · Abstract (English)

A lot of claims are made in social media posts, which may contain misinformation or fake news. Hence, it is crucial to identify claims as a first step towards claim verification. Given the huge number of social media posts, the task of identifying claims needs to be automated. This competition deals with the task of 'Claim Span Identification' in which, given a text, parts / spans that correspond to claims are to be identified. This task is more challenging than the traditional binary classification of text into claim or not-claim, and requires state-of-the-art methods in Pattern Recognition, Natural Language Processing and Machine Learning. For this competition, we used a newly developed dataset called HECSI containing about 8K posts in English and about 8K posts in Hindi with claim-spans marked by human annotators. This paper gives an overview of the competition, and the solutions developed by the participating teams.

主张识别多语言虚假信息文本标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。