arXiv:2501.10072cs.CL2025-01被引 1

用深度学习分析作者的词性分布,发现词序模式更显个人风格。

Author-Specific Linguistic Patterns Unveiled: A Deep Learning Study on Word Class Distributions

  • 基于词性标签和二元组构建作者特征向量
  • 二元组模型分类准确率显著高于单字模型
  • 适合对文学风格分析和作者识别感兴趣的读者

深度学习方法在计算语言学中日益广泛应用,用于挖掘文本数据中的模式。本研究通过词性(POS)标注和二元组分析,探讨作者特有的词性分布。利用深度神经网络,我们基于作者作品的词性向量和二元组频率矩阵进行文学作者分类。采用全连接与卷积神经网络架构,评估基于一元组和二元组的表示效果。结果表明,虽然一元组特征达到中等分类准确率,但二元组模型性能显著提升,说明词序中的词性模式更能体现作者风格。多维缩放(MDS)可视化显示作者作品呈现有意义的聚类,支持通过计算方法捕捉风格细微差别的假设。这些发现突显了深度学习与语言特征分析在作者画像和文学研究中的潜力。

原文摘要 · Abstract (English)

Deep learning methods have been increasingly applied to computational linguistics to uncover patterns in text data. This study investigates author-specific word class distributions using part-of-speech (POS) tagging and bigram analysis. By leveraging deep neural networks, we classify literary authors based on POS tag vectors and bigram frequency matrices derived from their works. We employ fully connected and convolutional neural network architectures to explore the efficacy of unigram and bigram-based representations. Our results demonstrate that while unigram features achieve moderate classification accuracy, bigram-based models significantly improve performance, suggesting that sequential word class patterns are more distinctive of authorial style. Multi-dimensional scaling (MDS) visualizations reveal meaningful clustering of authors' works, supporting the hypothesis that stylistic nuances can be captured through computational methods. These findings highlight the potential of deep learning and linguistic feature analysis for author profiling and literary studies.

作者识别深度学习词性标注文学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。