arXiv:2603.14756cs.CLcs.AI2026-03中稿 · IEEE Journal of Se…被引 1

提出推理阶段隐私保护翻译新任务,聚焦命名实体隐私防护。

Towards Privacy-Preserving Machine Translation at the Inference Stage: A New Task and Benchmark

  • 定义推理阶段隐私保护翻译任务,仅保护文本中命名实体
  • 构建3个基准测试集,设计专用评估指标
  • 提供起点方法,推动隐私敏感场景下的翻译应用

当前在线翻译服务需将用户文本发送至云端服务器,若文本含敏感信息,存在隐私泄露风险,限制其在隐私敏感场景的应用。为缓解此风险,可引入针对翻译模型推理阶段的隐私保护机制。然而,与文本分类、摘要等NLP子领域相比,机器翻译领域对推理阶段隐私保护的研究仍较薄弱,缺乏明确的任务定义、专用评估数据集、度量标准及基准方法。为此,本文提出新的「隐私保护机器翻译」(PPMT)任务,旨在保护推理阶段文本中的私密信息。研究聚焦于命名实体——因其常包含个人隐私与商业机密——提出保护策略。同时,构建了三个基准测试数据集,设计相应评估指标,并提出一系列基准方法,为该方向研究提供起点。期望本工作能为机器翻译中的隐私保护问题提供新视角与坚实基础。

原文摘要 · Abstract (English)

Current online translation services require sending user text to cloud servers, posing a risk of privacy leakage when the text contains sensitive information. This risk hinders the application of online translation services in privacy-sensitive scenarios. One way to mitigate this risk for online translation services is introducing privacy protection mechanisms targeting the inference stage of translation models. However, compared to subfields of NLP like text classification and summarization, the machine translation research community has limited exploration of privacy protection during the inference stage. There is no clearly defined privacy protection task for the inference stage, dedicated evaluation datasets and metrics, and reference benchmark methods. The absence of these elements has seriously constrained researchers' in-depth exploration of this direction. To bridge this gap, this paper proposes a novel "Privacy-Preserving Machine Translation" (PPMT) task, aiming to protect the private information in text during the model inference stage. For this task, we constructed three benchmark test datasets, designed corresponding evaluation metrics, and proposed a series of benchmark methods as a starting point for this task. The definition of privacy is complex and diverse. Considering that named entities often contain a large amount of personal privacy and commercial secrets, we have focused our research on protecting only the named entity's privacy in the text. We expect this research work will provide a new perspective and a solid foundation for the privacy protection problem in machine translation.

隐私保护机器翻译推理阶段命名实体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。