arXiv:2608.26135cs.CL2026-08

用数据科学方法分析网络言论,评估个人公众声誉风险。

Data Science Approaches to Evaluating Honours Candidates

论文配图:Data Science Approaches to Evaluating Honours Candidates
图 1 · 摘自论文原文
  • 构建端到端NLP流水线,从碎片化网络信息提取可审计的个体情感分布。
  • 自研的MINOS算法在区分正负面声誉上效果最优,准确率高于AFINN和VADER。
  • 适用于高风险决策场景,如英国荣誉体系候选人资格审查。

我们提出一种模块化的数据科学流程,用于从零散、非结构化的开源情报(OSINT)中估算公众对个体的情感倾向。该方法整合网页搜索、文本提取、相关性过滤、分词、共指消解与情感分析,将异构网络内容转化为可审计的个体级情感分布。我们对比了AFINN、VADER与MINOS三种情感分析算法,其中MINOS是专为识别声誉风险、不当行为及正面公众贡献语言而设计的领域感知算法。在已知声誉结果的公众人物上的应用表明,MINOS在区分正面、模糊与负面案例方面表现最清晰。结果表明,链式NLP与OSINT方法可支持透明、可复现且人工介入的情感评估,适用于高风险决策支持。我们以英国荣誉制度为例,该制度要求候选人具备高水平公共行为标准以维持荣誉资格。

原文摘要 · Abstract (English)

We present a modular data-science pipeline for estimating public sentiment towards individuals from fragmented, unstructured open-source intelligence (OSINT). The method chains web search, text extraction, relevance filtering, tokenisation, co-reference resolution, and sentiment analysis to convert heterogeneous web material into auditable person-level sentiment distributions. We compare AFINN and VADER with MINOS, a domain-informed sentiment algorithm designed to detect language associated with reputational risk, misconduct, and positive public contribution. Applied to public figures with known reputational outcomes, MINOS gives the clearest separation between positive, ambiguous, and negative cases. The results show that chained NLP and OSINT methods can support transparent, reproducible, human-in-the-loop sentiment assessment for high-stakes decision support. We demonstrate the approach on the UK Honours system, where individuals are required to display high standards of public conduct to maintain an Honour.

情感分析声誉评估OSINT决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。