用西班牙语模型和提示工程,让复杂文本变易懂,效果亮眼但可读性有短板。
Prompt-Based Simplification for Plain Language using Spanish Language Models
- 基于西班牙语模型,用提示词设计实现零样本简化,辅以LoRA微调提升效果。
- 最终系统在语义相似度上表现最佳(SIM=0.75),但可读性仅排第四(FH=69.72)。
- 适合关注语言可读性优化、尤其对西班牙语自然语言处理的研究者参考。
本文介绍HULAT-UC3M团队参与CLEARS 2025任务1:西班牙语文本到通俗语言(Plain Language, PL)的适配。我们探索了基于西班牙语训练模型的策略,包括使用提示工程的零样本配置,以及采用低秩适应(LoRA)微调的版本。在训练数据的代表性内部子集上评估不同策略,使用官方任务指标——余弦相似度(SIM)和Fernández-Huerta可读性指数(FH)来选择最优模型与提示组合。最终系统结合归一化步骤、RigoChat-7B-v2模型及专用的通俗语言提示,取得最高语义相似度(SIM=0.75),但在可读性上位列第四(FH=69.72)。研究还讨论了训练数据异质性带来的挑战,以及现有评估指标在捕捉语言清晰度与内容保留之间的局限。
原文摘要 · Abstract (English)
This paper describes the participation of HULAT-UC3M in CLEARS 2025 Subtask 1: Adaptation of Text to Plain Language (PL) in Spanish. We explored strategies based on models trained on Spanish texts, including a zero-shot configuration using prompt engineering and a fine-tuned version with Low-Rank Adaptation (LoRA). Different strategies were evaluated on representative internal subsets of the training data, using the official task metrics, cosine similarity (SIM) and the Fernández-Huerta readability index (FH) to guide the selection of the optimal model and prompt combination. The final system was selected for its balanced and consistent performance, combining normalization steps, the RigoChat-7B-v2 model, and a dedicated PL-oriented prompt. It ranked first in semantic similarity (SIM = 0.75), however, fourth in readability (FH = 69.72). We also discuss key challenges related to training data heterogeneity and the limitations of current evaluation metrics in capturing both linguistic clarity and content preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。