arXiv:2601.14683cs.AI2026-01中稿 · and Waiting to be …

用本地大模型实现智能敏感信息匿名化,更准且不破坏文本含义。

Local Language Models for Context-Aware Adaptive Anonymization of Sensitive Text

  • 基于本地LLM构建三步框架,按风险等级动态选择四种匿名策略。
  • Phi模型识别超91%敏感信息,94.8%保持原文情感,准确性高。
  • 适合需要隐私保护的质性研究,如访谈数据处理与合规分析。

定性研究常包含个人、情境及组织细节,若处理不当会带来隐私风险。人工匿名耗时长、不一致,常遗漏关键标识符;现有自动化工具多依赖模式匹配或固定规则,无法捕捉上下文,可能改变数据原意。本研究采用本地大语言模型,构建可重复、上下文感知的匿名化流程,用于检测与匿名化质性转录文本中的敏感信息。提出结构化自适应匿名框架(SFAA),包含检测、分类与自适应匿名三步。SFAA融合四种策略:基于规则替换、上下文感知重写、泛化和抑制,根据标识类型与风险等级动态应用。标识范围依据GDPR、HIPAA、OECD等国际隐私与伦理标准。研究采用双方法评估,结合人工与LLM辅助处理,通过两个案例验证:其一为82场关于组织游戏化的面对面访谈;其二为93场由AI驱动的访谈,测试大模型对工作场所隐私的认知。使用本地模型LLaMA与Phi进行性能评估。结果表明,大模型发现的敏感数据多于人工评审,其中Phi在识别敏感数据上优于LLaMA,虽略有误判,但识别率超过91%,且94.8%的文本保留原始语义情感,不影响后续质性分析。

原文摘要 · Abstract (English)

Qualitative research often contains personal, contextual, and organizational details that pose privacy risks if not handled appropriately. Manual anonymization is time-consuming, inconsistent, and frequently omits critical identifiers. Existing automated tools tend to rely on pattern matching or fixed rules, which fail to capture context and may alter the meaning of the data. This study uses local LLMs to build a reliable, repeatable, and context-aware anonymization process for detecting and anonymizing sensitive data in qualitative transcripts. We introduce a Structured Framework for Adaptive Anonymizer (SFAA) that includes three steps: detection, classification, and adaptive anonymization. The SFAA incorporates four anonymization strategies: rule-based substitution, context-aware rewriting, generalization, and suppression. These strategies are applied based on the identifier type and the risk level. The identifiers handled by the SFAA are guided by major international privacy and research ethics standards, including the GDPR, HIPAA, and OECD guidelines. This study followed a dual-method evaluation that combined manual and LLM-assisted processing. Two case studies were used to support the evaluation. The first includes 82 face-to-face interviews on gamification in organizations. The second involves 93 machine-led interviews using an AI-powered interviewer to test LLM awareness and workplace privacy. Two local models, LLaMA and Phi were used to evaluate the performance of the proposed framework. The results indicate that the LLMs found more sensitive data than a human reviewer. Phi outperformed LLaMA in finding sensitive data, but made slightly more errors. Phi was able to find over 91% of the sensitive data and 94.8% kept the same sentiment as the original text, which means it was very accurate, hence, it does not affect the analysis of the qualitative data.

隐私保护大模型质性研究匿名化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。