用嵌入+概率+降噪,把新闻变量化语义信号
Text-as-Signal: Quantitative Semantic Scoring with Embeddings, Logprobs, and Noise Reduction

- 用文档嵌入和位置词典计算语义得分
- 在1.19万篇葡萄牙AI新闻中识别出6维语义空间
- 可配置框架适合多种文本分析任务
本文提出一种将文本语料转化为定量语义信号的实用流程。每篇新闻以完整文档嵌入表示,通过可配置的位置词典进行基于对数概率的评分,并投影到降噪后的低维流形以支持结构化解读。本案例研究中,词典设定为六个语义维度,应用于11,922篇关于人工智能的葡萄牙语新闻。所得身份空间支持文档级语义定位与语料库级特征描述,通过聚合概览实现。实验表明,Qwen嵌入、UMAP降维、模型输出空间直接提取的语义指标,以及三阶段异常检测流程,共同构成面向AI工程任务(如语料审查、监控与下游分析支持)的可操作文本转信号工作流。由于身份层可配置,该框架可适配不同分析需求,而非绑定于通用模式。
原文摘要 · Abstract (English)
This paper presents a practical pipeline for turning text corpora into quantitative semantic signals. Each news item is represented as a full-document embedding, scored through logprob-based evaluation over a configurable positional dictionary, and projected onto a noise-reduced low-dimensional manifold for structural interpretation. In the present case study, the dictionary is instantiated as six semantic dimensions and applied to a corpus of 11,922 Portuguese news articles about Artificial Intelligence. The resulting identity space supports both document-level semantic positioning and corpus-level characterization through aggregated profiles. We show how Qwen embeddings, UMAP, semantic indicators derived directly from the model output space, and a three-stage anomaly-detection procedure combine into an operational text-as-signal workflow for AI engineering tasks such as corpus inspection, monitoring, and downstream analytical support. Because the identity layer is configurable, the same framework can be adapted to the requirements of different analytical streams rather than fixed to a universal schema.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。