Search and Browse – PORTULAN CLARIN

Royal inquiries of 1258 (primarily published in the Portugaliae Monumenta Historica).

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

CRPC Discourse Bank v1.0

The CRPC Discourse Bank is labeled for discourse relations (also referred to as rhetorical relations or coher- ence relations), such as cause and condition, that hold between two spans of text and contribute to ensure the overall cohesion and coherence of the text. The scheme follows the principl...

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

Europarl QTLeap WSD/NED corpus

Europarl QTLeap WSD/NED corpus This corpora is part of Deliverable 5.5 of the European Commission project QTLeap FP7-ICT-2013.4.1-610516 (http://qtleap.eu). The texts are sentences from the Europarl parallel corpus (Koehn, 2005). We selected the monolingual sentences from parallel corpora ...

Resource Type:	Corpus
Media Type:	Text
Languages:	Basque
	Bulgarian
	Czech
	English
	Portuguese
	Spanish; Castilian

SIMPLE Portuguese Lexicon

The SIMPLE Portuguese Lexicon is constituted by 10,438 entries semantically encoded, accordingly to the parole common encoding standards.

Resource Type:	Lexical / Conceptual
Media Type:	Text
Language:	Portuguese

Archivo dos Açores, dir. Ernesto do Canto, 1.ª série, Ponta Delgada, Vol. 1-12

A publicação Arquivo dos Açores, consagrada como obra de referência para a investigação histórica sobre o arquipélago dos Açores, conta com duas séries, num total de 20 volumes. A primeira série do Arquivo dos Açores, composta por 15 volumes, decorreu entre 1878 e 1959, com grandes interrupções r...

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

askIT Dataset

Collection of dialogues extracted from subreddits related to Information Technology (IT) and extracted with RDET (Reddit Dataset Extraction Tool). It is composed of 61,842,638 tokens in 179,358 dialogues.

Resource Type:	Corpus
Media Type:	Text
Language:	English

UIMA/U-Compare Apertium POS Tagger

This tool assigns a part-of-speech tag and base form to each token in a text. It operates on text that has previously been tokenised and morphologically analysed. The POS tagger is a module of Apertium machine translation system. The provided tool can currently operate on a subset of the language...

Resource Type:	Tool / Service
Languages:	Basque
	Catalan
	English
	Galician
	Portuguese
	Spanish

UIMA Apertium Translator

This tool translates text from a source language into a target language. It operates on text that has previously been tokenised and morphologically analysed, and POS-tagged. Target language tokens are assigned POS tags and morphological analyses. The Apertium Translator is a module of Apertium ma...

Resource Type:	Tool / Service
Languages:	Basque
	Catalan
	English
	Galician
	Portuguese
	Spanish

Summ-it

The corpus was developed as a linguistic resource for Automatic Summarization research and his relation with different issues to engage studies on the discourse treatment. Summ-it consists of fifty texts from Science domain extracted from Science section of Brazilian daily newspaper Folha de Sã...

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

U-Compare Part-of-Speech Tagging service

Web service created by exporting UIMA-based workflow from the U-Compare text mining system. Functionality: Identifies tokens in plain text and assigns parts-of-speech Tools in workflow: MLRS POS Tagger web service (University of Malta) NOTE: The licence provided covers the web service only. To...

Resource Type:	Tool / Service
Language:	Maltese

Order by:

Filter by: