arXiv:2411.10955cs.CL2024-11

构建两岸中文社交媒体可比语料库,追踪现代语言使用差异。

A Topic-aware Comparable Corpus of Chinese Variations

  • 基于两岸社交平台数据,构建主题感知的可比语料库
  • 涵盖大陆微博与台湾Dcard,定期更新反映最新用语
  • 适合研究汉语方言差异与社会语言学的学者

本研究旨在填补空白,从中国大陆和台湾的社会媒体中分别收集普通话数据,构建一个主题感知的可比语料库。使用台湾的Dcard和中国大陆的新浪微博作为数据源,创建了一个定期更新、反映现代社交媒体语言使用的可比语料库。

原文摘要 · Abstract (English)

This study aims to fill the gap by constructing a topic-aware comparable corpus of Mainland Chinese Mandarin and Taiwanese Mandarin from the social media in Mainland China and Taiwan, respectively. Using Dcard for Taiwanese Mandarin and Sina Weibo for Mainland Chinese, we create a comparable corpus that updates regularly and reflects modern language use on social media.

语料库两岸语言社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。