Subtopic annotation and automatic segmentation for news texts in Brazilian Portuguese

dc.creatorCardoso, Paula C. F.
dc.creatorPardo, Thiago A. S.
dc.creatorTaboada, Maite
dc.date.accessioned2019-01-25T12:46:02Z
dc.date.available2019-01-25T12:46:02Z
dc.date.issued2017
dc.description.abstractSubtopic segmentation aims to break documents into subtopical text passages, which develop a main topic in a text. Being capable of automatically detecting subtopics is very useful for several Natural Language Processing applications. For instance, in automatic summarisation, having the subtopics at hand enables the production of summaries with good subtopic coverage. Given the usefulness of subtopic segmentation, it is common to assemble a reference-annotated corpus that supports the study of the envisioned phenomena and the development and evaluation of systems. In this paper, we describe the subtopic annotation process in a corpus of news texts written in Brazilian Portuguese, following a systematic annotation process and answering the main research questions when performing corpus annotation. Based on this corpus, we propose novel methods for subtopic segmentation following patterns of discourse organisation, specifically using Rhetorical Structure Theory. We show that discourse structures mirror the subtopic changes in news texts. An important outcome of this work is the freely available annotated corpus, which, to the best of our knowledge, is the only one for Portuguese. We demonstrate that some discourse knowledge may significantly help to find boundaries automatically in a text. In particular, the relation type and the level of the tree structure are important features.pt_BR
dc.description.provenanceSubmitted by André Calsavara (andre.calsavara@biblioteca.ufla.br) on 2019-01-10T17:58:02Z No. of bitstreams: 0en
dc.description.provenanceApproved for entry into archive by André Calsavara (andre.calsavara@biblioteca.ufla.br) on 2019-01-25T12:46:02Z (GMT) No. of bitstreams: 0en
dc.description.provenanceMade available in DSpace on 2019-01-25T12:46:02Z (GMT). No. of bitstreams: 0 Previous issue date: 2017en
dc.identifier.citationCARDOSO, P. C. F.; PARDO, T. A. S. TABOADA, M. Subtopic annotation and automatic segmentation for news texts in Brazilian Portuguese. Corpora, v. 12, n. 1, p. 23-54, 2017.pt_BR
dc.identifier.urihttps://repositorio.ufla.br/handle/1/32557
dc.identifier.urihttps://www.euppublishing.com/doi/10.3366/cor.2017.0108pt_BR
dc.languageen_USpt_BR
dc.publisherEdinburgh University Presspt_BR
dc.rightsopenAccesspt_BR
dc.sourceCorporapt_BR
dc.subjectCorpus annotationpt_BR
dc.subjectNewspaper discoursept_BR
dc.subjectSubtopicspt_BR
dc.subjectText segmentationpt_BR
dc.subjectDiscurso de jornalpt_BR
dc.subjectSubtópicospt_BR
dc.subjectSegmentação de textopt_BR
dc.titleSubtopic annotation and automatic segmentation for news texts in Brazilian Portuguesept_BR
dc.typeArtigopt_BR

Arquivos

Licença do pacote

Agora exibindo 1 - 1 de 1
Carregando...
Imagem de Miniatura
Nome:
license.txt
Tamanho:
953 B
Formato:
Item-specific license agreed upon to submission
Descrição: