基于最大熵的汉语篇章结构自动分析方法

北京大学学报（自然科学版）

基于最大熵的汉语篇章结构自动分析方法

涂眉,周玉,宗成庆

中国科学院自动化研究所模式识别国家重点实验室, 北京100190;

收稿日期:2013-06-15 出版日期:2014-01-20 发布日期:2014-01-20

Automatically Parsing Chinese Discourse Based on Maximum Entropy

TU Mei, ZHOU Yu, ZONG Chengqing

National Laboratory of Pattern Recongnition NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing 100190;

Received:2013-06-15 Online:2014-01-20 Published:2014-01-20

摘要/Abstract

摘要： 在标有复句逻辑语义关系的清华汉语树库上, 研究汉语篇章语义片段自动切分以及篇章关系的自动标注方法。通过比较不同序列标注模型对汉语篇章语义单元切分的性能, 提出基于最大熵模型的汉语篇章结构分析方法。实验结果表明, 篇章语义单元自动切分的F值能达到89.1%, 当篇章语义结构树的高度不超过6层时, 篇章语义关系标注的F值为63%。

关键词: 语义片段自动切分, 篇章结构分析, 逻辑语义关系, 树库

Abstract: The authors focus on how to segment semantic units in Chinese discourse and how to label relations among semantic units automatically. During the parsing process, several sequence labelling methods are compared for discourse segmentation, while a maximum entropy-based training and decoding algorithm is specially proposed. Experiments are done based on Tsinghua Chinese Treebank, which is annotated with logical and semantic relations at complex-sentence level. Experimental results show that F-score of discourse segmentation reaches 89.1%. When parsing discourses with no more than 6 relations included, the labeling F-score can achieve 63%.

Key words: automatic discourse segmentation, discourse structure parsing, Chinese logical and semantic relation, Tsinghua Chinese Treebank

中图分类号:

TP391

涂眉,周玉,宗成庆. 基于最大熵的汉语篇章结构自动分析方法[J]. 北京大学学报（自然科学版）.

TU Mei,ZHOU Yu,ZONG Chengqing. Automatically Parsing Chinese Discourse Based on Maximum Entropy[J]. Acta Scientiarum Naturalium Universitatis Pekinensis.

导出引用管理器 EndNote|Ris|BibTeX

链接本文: https://xbna.pku.edu.cn/CN/

https://xbna.pku.edu.cn/CN/Y2014/V50/I1/125

[1]	史林林, 邱立坤, 亢世勇. 基于规则的依存树库错误自动检测与分析[J]. 北京大学学报（自然科学版）, 2016, 52(1): 58-64.
[2]	李艳翠,孙静,周国栋,冯文贺. 基于清华汉语树库的复句关系词识别与分类研究[J]. 北京大学学报（自然科学版）, 2014, 50(1): 118-124.
[3]	孙静,李艳翠,周国栋,冯文贺. 汉语隐式篇章关系识别[J]. 北京大学学报（自然科学版）, 2014, 50(1): 111-117.
[4]	王慧兰. 汉语句类依存树库的构建研究[J]. 北京大学学报（自然科学版）, 2013, 49(1): 25-30.