Use este identificador para citar ou linkar para este item: http://repositorio.ufla.br/jspui/handle/1/32557
Título: Subtopic annotation and automatic segmentation for news texts in Brazilian Portuguese
Palavras-chave: Corpus annotation
Newspaper discourse
Subtopics
Text segmentation
Discurso de jornal
Subtópicos
Segmentação de texto
Data do documento: 2017
Editor: Edinburgh University Press
Citação: CARDOSO, P. C. F.; PARDO, T. A. S. TABOADA, M. Subtopic annotation and automatic segmentation for news texts in Brazilian Portuguese. Corpora, v. 12, n. 1, p. 23-54, 2017.
Resumo: Subtopic segmentation aims to break documents into subtopical text passages, which develop a main topic in a text. Being capable of automatically detecting subtopics is very useful for several Natural Language Processing applications. For instance, in automatic summarisation, having the subtopics at hand enables the production of summaries with good subtopic coverage. Given the usefulness of subtopic segmentation, it is common to assemble a reference-annotated corpus that supports the study of the envisioned phenomena and the development and evaluation of systems. In this paper, we describe the subtopic annotation process in a corpus of news texts written in Brazilian Portuguese, following a systematic annotation process and answering the main research questions when performing corpus annotation. Based on this corpus, we propose novel methods for subtopic segmentation following patterns of discourse organisation, specifically using Rhetorical Structure Theory. We show that discourse structures mirror the subtopic changes in news texts. An important outcome of this work is the freely available annotated corpus, which, to the best of our knowledge, is the only one for Portuguese. We demonstrate that some discourse knowledge may significantly help to find boundaries automatically in a text. In particular, the relation type and the level of the tree structure are important features.
URI: https://www.euppublishing.com/doi/10.3366/cor.2017.0108
http://repositorio.ufla.br/jspui/handle/1/32557
Aparece nas coleções:DCC - Artigos publicados em periódicos

Arquivos associados a este item:
Não existem arquivos associados a este item.


Os itens no repositório estão protegidos por copyright, com todos os direitos reservados, salvo quando é indicado o contrário.

Ferramentas do administrador