arXiv:2503.10675cs.CLcs.AI2025-03被引 3

让摘要自动适配不同读者的阅读水平,提升信息可及性。

Beyond One-Size-Fits-All Summarization: Customizing Summaries for Diverse Users

  • 自建数据集+定制模型,精准控制摘要难易度。
  • 在土耳其语上验证,显著优于传统模型的可读性控制。
  • 适合教育、科普等需分层传播的场景使用。

近年来,基于Transformer的自动文本摘要技术取得显著进展,但对生成摘要可读性的控制仍是一个未充分探索的问题,尤其对于土耳其语这类语言结构复杂的语言。可读性控制对面向不同文化程度读者(如从小学到研究生)的信息传播至关重要,能有效提升理解与参与度。当前摘要模型普遍缺乏调整输出复杂度的能力,导致内容要么过于简单,要么过于艰深。为此,我们构建了专属数据集,并设计了定制化模型架构,实现了在保持摘要准确性和连贯性的同时,对可读性进行有效调控。通过与监督微调基线模型的严格对比,证明了本方法在生成可读性感知摘要方面的优越性。

原文摘要 · Abstract (English)

In recent years, automatic text summarization has witnessed significant advancement, particularly with the development of transformer-based models. However, the challenge of controlling the readability level of generated summaries remains an under-explored area, especially for languages with complex linguistic features like Turkish. This gap has the effect of impeding effective communication and also limits the accessibility of information. Controlling readability of textual data is an important element for creating summaries for different audiences with varying literacy and education levels, such as students ranging from primary school to graduate level, as well as individuals with diverse educational backgrounds. Summaries that align with the needs of specific reader groups can improve comprehension and engagement, ensuring that the intended message is effectively communicated. Furthermore, readability adjustment is essential to expand the usability of summarization models in educational and professional domains. Current summarization models often don't have the mechanisms to adjust the complexity of their outputs, resulting in summaries that may be too simplistic or overly complex for certain types of reader groups. Developing adaptive models that can tailor content to specific readability levels is therefore crucial. To address this problem, we create our own custom dataset and train a model with our custom architecture. Our method ensures that readability levels are effectively controlled while maintaining accuracy and coherence. We rigorously compare our model to a supervised fine-tuned baseline, demonstrating its superiority in generating readability-aware summaries.

文本摘要可读性控制多层级传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。