通过计算风格学实现英、僧伽罗罗马字文本的作者识别
Stylomech: Unveiling Authorship via Computational Stylometry in English and Romanized Sinhala
- 仅需对比嫌疑文本与匿名文本,无需大规模语料库
- 将文本转化为数值特征进行模型训练,适应多作者场景
- 首次系统研究僧伽罗罗马字文本的作者归属,助力数字内容可信
随着网络2.0时代的发展,社交媒体技术与全球通信的结合为社会带来正负面影响。版权争议与作者识别日益重要,因社会伦理缺失导致的内容侵权现象显著增加。近年来,英语及僧伽罗罗马字文本的作者归属成为关键需求。该研究在僧伽罗罗马字这一长期未被充分探索的语言背景下,提出一种新颖的作者归属方法,仅需比较嫌疑作者文本与匿名文本两组数据,不同于传统依赖大规模语料库的方法。研究通过将同一作者与不同作者的文本对转换为数值表示,使模型基于这些表示进行训练,而非原始文本,从而可在合理质量的文本条件下适用于多种作者和语境。该工作拓展了作者归属在多样化语言环境中的应用,有助于提升数字交流中的信任与责任意识,尤其在斯里兰卡具有实际意义。本研究在英语与僧伽罗罗马字中均实现了作者归属的开创性方法,满足数字时代内容验证与知识产权保护的关键需求。
原文摘要 · Abstract (English)
With the advent of Web 2.0, the development in social technology coupled with global communication systematically brought positive and negative impacts to society. Copyright claims and Author identification are deemed crucial as there has been a considerable amount of increase in content violation owing to the lack of proper ethics in society. The Author's attribution in both English and Romanized Sinhala became a major requirement in the last few decades. As an area largely unexplored, particularly within the context of Romanized Sinhala, the research contributes significantly to the field of computational linguistics. The proposed author attribution system offers a unique approach, allowing for the comparison of only two sets of text: suspect author and anonymous text, a departure from traditional methodologies which often rely on larger corpora. This work focuses on using the numerical representation of various pairs of the same and different authors allowing for, the model to train on these representations as opposed to text, this allows for it to apply to a multitude of authors and contexts, given that the suspected author text, and the anonymous text are of reasonable quality. By expanding the scope of authorship attribution to encompass diverse linguistic contexts, the work contributes to fostering trust and accountability in digital communication, especially in Sri Lanka. This research presents a pioneering approach to author attribution in both English and Romanized Sinhala, addressing a critical need for content verification and intellectual property rights enforcement in the digital age.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。