arXiv:2507.13115cs.CL2025-07ACL

构建自我的文本识别框架,助力心理健康与哲学研究。

A Computational Framework to Identify Self-Aspects in Text

  • 设计自我特质本体与标注数据集,奠定分析基础。
  • 对比判别模型、生成大模型与嵌入检索,评估四项性能指标。
  • 应用于心理与现象学案例,推动跨领域应用。

本博士提案提出构建一个计算框架,用于识别文本中的自我特质。自我是多维度概念,在语言中有所体现。尽管在认知科学和现象学等领域已有描述,但在自然语言处理(NLP)中仍属未充分探索领域。自我诸多方面与心理健康的多种现象高度相关,凸显了系统性基于NLP分析的必要性。为此,我们计划构建自我特质本体及金标准标注数据集。在此基础上,将开发并评估传统判别模型、生成式大语言模型以及基于嵌入的检索方法,依据可解释性、真实标签符合度、准确率与计算效率四项标准进行比较。表现最佳的模型将应用于心理健康与经验现象学的案例研究。

原文摘要 · Abstract (English)

This Ph.D. proposal introduces a plan to develop a computational framework to identify Self-aspects in text. The Self is a multifaceted construct and it is reflected in language. While it is described across disciplines like cognitive science and phenomenology, it remains underexplored in natural language processing (NLP). Many of the aspects of the Self align with psychological and other well-researched phenomena (e.g., those related to mental health), highlighting the need for systematic NLP-based analysis. In line with this, we plan to introduce an ontology of Self-aspects and a gold-standard annotated dataset. Using this foundation, we will develop and evaluate conventional discriminative models, generative large language models, and embedding-based retrieval approaches against four main criteria: interpretability, ground-truth adherence, accuracy, and computational efficiency. Top-performing models will be applied in case studies in mental health and empirical phenomenology.

自我识别心理分析本体构建大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。