arXiv:2509.25086cs.CL2025-09被引 1

用小模型实现安全高效的词汇简化,兼顾隐私与正确性。

Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs

  • 用小模型结合合成数据和上下文学习,本地部署更安全。
  • 小模型自动评分高,但知识蒸馏会增加有害简化风险。
  • 通过输出概率识别有害简化,过滤后仍保留有益改写。

尽管大语言模型在词汇简化任务上表现优异,但在隐私敏感和资源受限的场景中应用存在挑战。由于残障人士等弱势群体是该技术的主要受益者,确保输出的安全性和正确性至关重要。为此,我们提出一种基于小语言模型的高效词汇简化框架,可在本地环境中部署。在框架内,我们探索了合成数据的知识蒸馏和上下文学习作为基线方法。我们在五种语言上进行了自动与人工评估。人工分析发现,知识蒸馏虽提升自动指标分数,但增加了有害简化风险。重要的是,我们发现模型输出概率可有效指示有害简化。基于此,提出一种过滤策略,在抑制有害简化的同时,基本保留有益改写。本工作建立了小模型在效率与安全性的基准,揭示了性能、效率与安全之间的关键权衡,并展示了安全落地的可行路径。

原文摘要 · Abstract (English)

Despite their strong performance, large language models (LLMs) face challenges in real-world application of lexical simplification (LS), particularly in privacy-sensitive and resource-constrained environments. Moreover, since vulnerable user groups (e.g., people with disabilities) are one of the key target groups of this technology, it is crucial to ensure the safety and correctness of the output of LS systems. To address these issues, we propose an efficient framework for LS systems that utilizes small LLMs deployable in local environments. Within this framework, we explore knowledge distillation with synthesized data and in-context learning as baselines. Our experiments in five languages evaluate model outputs both automatically and manually. Our manual analysis reveals that while knowledge distillation boosts automatic metric scores, it also introduces a safety trade-off by increasing harmful simplifications. Importantly, we find that the model's output probability is a useful signal for detecting harmful simplifications. Leveraging this, we propose a filtering strategy that suppresses harmful simplifications while largely preserving beneficial ones. This work establishes a benchmark for efficient and safe LS with small LLMs. It highlights the key trade-offs between performance, efficiency, and safety, and demonstrates a promising approach for safe real-world deployment.

词汇简化小模型安全本地部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。