首个评估大模型科学文献多属性可控摘要的基准
CCSBench: Evaluating Compositional Controllability in LLMs for Scientific Document Summarization
- 构建首个科学文档多属性可控摘要评测基准
- 发现大模型在隐含属性控制上存在显著能力瓶颈
- 适合关注可控生成与科学文本理解的研究者
为促进科学知识向多元受众传播,理想的科学文献摘要系统应能同时控制长度、实证焦点等多重属性。然而现有研究多聚焦单一属性控制,对多属性组合控制研究不足。为此,我们提出CCSBench,首个面向科学领域可组合可控摘要的评估基准。该基准支持对显性属性(如长度)和隐性属性(如概念或实证焦点)的细粒度控制。我们在多种大语言模型(LLMs)上进行实验,涵盖上下文学习、参数高效微调及两阶段模块化方法等设置。结果表明,大模型在平衡不同属性控制,尤其是需要深层理解与抽象推理的隐性属性方面存在明显局限。
原文摘要 · Abstract (English)
To broaden the dissemination of scientific knowledge to diverse audiences, it is desirable for scientific document summarization systems to simultaneously control multiple attributes such as length and empirical focus. However, existing research typically focuses on controlling single attributes, leaving the compositional control of multiple attributes underexplored. To address this gap, we introduce CCSBench, the first evaluation benchmark for compositional controllable summarization in the scientific domain. Our benchmark enables fine-grained control over both explicit attributes (e.g., length), which are objective and straightforward, and implicit attributes (e.g., conceptual or empirical focus), which are more subjective and abstract. We conduct extensive experiments using various large language models (LLMs) under various settings, including in-context learning, parameter-efficient fine-tuning, and two-stage modular methods for balancing control over different attributes. Our findings reveal significant limitations in LLMs capabilities in balancing trade-offs between control attributes, especially implicit ones that require deeper understanding and abstract reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。