跨语言分析维基百科对LGBT人物的描述,发现事实缺失与负面信息突出差异。
Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on Wikipedia
- 提出InfoGap方法,定位多语言文章中的事实空白与不一致
- 在2700个词条中发现英、俄、法三语覆盖差异显著
- 俄语版更倾向强调带有负面含义的个人事实
为解释社会现象并识别系统性偏见,计算社会科学中的大量研究聚焦于比较文本分析。这些研究通常依赖粗粒度的语料库统计或局部词汇级分析,且主要集中在英文。我们引入InfoGap方法——一种高效可靠的跨语言事实级信息缺口与不一致定位技术。通过分析英、俄、法三语维基百科中2700个关于LGBT人物的传记页面,我们发现不同语言间事实覆盖存在显著差异。此外,分析显示带有负面含义的事实在俄语维基中更易被突出。关键的是,InfoGap既支持大规模分析,又能精确定位到文档和具体事实层面的信息缺口,为大规模、细致的跨语言比较分析奠定了新基础。
原文摘要 · Abstract (English)
To explain social phenomena and identify systematic biases, much research in computational social science focuses on comparative text analyses. These studies often rely on coarse corpus-level statistics or local word-level analyses, mainly in English. We introduce the InfoGap method -- an efficient and reliable approach to locating information gaps and inconsistencies in articles at the fact level, across languages. We evaluate InfoGap by analyzing LGBT people's portrayals, across 2.7K biography pages on English, Russian, and French Wikipedias. We find large discrepancies in factual coverage across the languages. Moreover, our analysis reveals that biographical facts carrying negative connotations are more likely to be highlighted in Russian Wikipedia. Crucially, InfoGap both facilitates large scale analyses, and pinpoints local document- and fact-level information gaps, laying a new foundation for targeted and nuanced comparative language analysis at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。