用主题-解释结构让大模型更懂韩语表格,输出更易读。
Theme-Explanation Structure for Table Summarization using Large Language Models: A Case Study on Korean Tabular Data
- 分步推理+记者角色提示,提升大模型对表格的理解
- 输出分主题和解释两部分,可读性显著提升
- 无需大量标注数据,适合资源少的场景
表格是行政领域传递关键信息的主要载体,但其复杂性限制了大语言模型(LLMs)的有效利用。本文提出基于主题-解释结构的表格摘要方法(Tabular-TX),专门针对韩语行政文档设计,旨在生成高度可解释的摘要。现有方法常忽视人类友好的输出。Tabular-TX首先通过多步推理增强大模型对表格的深层理解,再采用记者角色提示策略生成清晰句子。关键在于将输出结构化为主题部分(副词性短语)与解释部分(谓语性从句),极大提升可读性。该方法依赖上下文学习,无需大规模微调及标注数据或高算力资源。实验表明,Tabular-TX能有效处理复杂表格结构与元数据,在低资源场景下仍具备鲁棒性和高效性,是面向人类的表格摘要新方案。
原文摘要 · Abstract (English)
Tables are a primary medium for conveying critical information in administrative domains, yet their complexity hinders utilization by Large Language Models (LLMs). This paper introduces the Theme-Explanation Structure-based Table Summarization (Tabular-TX) pipeline, a novel approach designed to generate highly interpretable summaries from tabular data, with a specific focus on Korean administrative documents. Current table summarization methods often neglect the crucial aspect of human-friendly output. Tabular-TX addresses this by first employing a multi-step reasoning process to ensure deep table comprehension by LLMs, followed by a journalist persona prompting strategy for clear sentence generation. Crucially, it then structures the output into a Theme Part (an adverbial phrase) and an Explanation Part (a predicative clause), significantly enhancing readability. Our approach leverages in-context learning, obviating the need for extensive fine-tuning and associated labeled data or computational resources. Experimental results show that Tabular-TX effectively processes complex table structures and metadata, offering a robust and efficient solution for generating human-centric table summaries, especially in low-resource scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。