arXiv:2502.18590cs.CL2025-02被引 2

快速提取96种可解释的文本风格特征,提升语篇分析效率

Neurobiber: Fast and Interpretable Stylistic Feature Extraction

  • 基于Transformer构建,从文本中自动预测96个风格特征
  • 速度比现有开源系统快56倍,且在权威数据集上表现良好
  • 适合需要实时分析的场景,如内容监控与司法取证

语言风格对理解文本意义和沟通目的至关重要,但大规模提取详细风格特征仍具挑战。我们提出Neurobiber,一个基于Transformer的快速、可解释风格分析系统,建立在Biber的多维分析(MDA)基础上。Neurobiber从开源的BiberPlus库(一个计算风格特征并提供主成分分析与因子分析等集成分析的Python工具包)中预测96个Biber风格特征。尽管速度比现有开源系统快达56倍,Neurobiber在CORE语料库上复现了经典的MDA洞察,并在PAN 2020作者身份验证任务中取得具有竞争力的表现,无需大量重新训练。其高效且可解释的表示可直接融入下游NLP流程,推动大规模风格计量研究、司法分析与实时文本监控。所有组件均公开可用。

原文摘要 · Abstract (English)

Linguistic style is pivotal for understanding how texts convey meaning and fulfill communicative purposes, yet extracting detailed stylistic features at scale remains challenging. We present Neurobiber, a transformer-based system for fast, interpretable style profiling built on Biber's Multidimensional Analysis (MDA). Neurobiber predicts 96 Biber-style features from our open-source BiberPlus library (a Python toolkit that computes stylistic features and provides integrated analytics, e.g., PCA and factor analysis). Despite being up to 56 times faster than existing open source systems, Neurobiber replicates classic MDA insights on the CORE corpus and achieves competitive performance on the PAN 2020 authorship verification task without extensive retraining. Its efficient and interpretable representations readily integrate into downstream NLP pipelines, facilitating large-scale stylometric research, forensic analysis, and real-time text monitoring. All components are made publicly available.

风格分析可解释性TransformerNLP工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。