arXiv:2602.07499cs.CL2026-02Conference of the …被引 1

分步引导大模型实现多语言无监督句式简化

Let's Simplify Step by Step: Guiding LLM Towards Multilingual Unsupervised Proficiency-Controlled Sentence Simplification

  • 将复杂简化拆解为可管理步骤,动态规划路径并利用对话历史推理
  • 跨五种语言在两个数据集上提升效果,计算步骤减少22%-42%
  • 适合需要精准控制简化程度的研究者和多语言NLP开发者

大型语言模型在跨大阅读难度层级的可控句式简化上表现有限。本文提出一种框架,通过动态路径规划、语义感知样例选择及结合对话历史的思维链生成,将复杂简化过程分解为可管理步骤。在五个语言的两个基准上评估显示,该方法提升了简化效果,同时减少22%-42%的计算步骤。人工评估确认简化效果与语义保留之间存在根本权衡。值得注意的是,即使人类标注者也难以就语义保留达成一致,凸显该任务的内在复杂性。研究表明,虽然分步简化增强了控制力,但在大幅简化中保持语义保真仍是一个开放挑战。

原文摘要 · Abstract (English)

Large language models demonstrate limited capability in proficiency-controlled sentence simplification, particularly when simplifying across large readability levels. We propose a framework that decomposes complex simplifications into manageable steps through dynamic path planning, semantic-aware exemplar selection, and chain-of-thought generation with conversation history for coherent reasoning. Evaluation on five languages across two benchmarks shows our approach improves simplification effectiveness while reducing computational steps by 22-42%. Human evaluation confirms the fundamental trade-off between simplification effectiveness and meaning preservation. Notably, even human annotators struggle to agree on semantic preservation judgments, highlighting the inherent complexity of this task. Our work shows that while step-by-step simplification improves control, preserving semantic fidelity during extensive simplification remains an open challenge.

句式简化大模型多语言可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。