Search and Browse – PORTULAN CLARIN

UIMA/U-Compare Apertium POS Tagger

This tool assigns a part-of-speech tag and base form to each token in a text. It operates on text that has previously been tokenised and morphologically analysed. The POS tagger is a module of Apertium machine translation system. The provided tool can currently operate on a subset of the language...

Resource Type:	Tool / Service
Languages:	Basque
	Catalan
	English
	Galician
	Portuguese
	Spanish

LX-Rare Word Similarity Dataset

The LX-Rare Word Similarity Data set was created from Stanford Rare Word (RW) Similarity data set (Luong et al., 2013). This list contains 2 034 words (1 017 pairs of words). All the words were extracted from Wikipedia and from WordNet (Miller, 1995), a lexical database where the concepts are gro...

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

UIMA Apertium Translator

This tool translates text from a source language into a target language. It operates on text that has previously been tokenised and morphologically analysed, and POS-tagged. Target language tokens are assigned POS tags and morphological analyses. The Apertium Translator is a module of Apertium ma...

Resource Type:	Tool / Service
Languages:	Basque
	Catalan
	English
	Galician
	Portuguese
	Spanish

COVID-19 EU presscorner v1 dataset. Bilingual (EN-PT)

Bilingual (EN-PT) corpus acquired from website (https://ec.europa.eu/commission/presscorner/) of the EU portal (14th May 2020).

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Portuguese

COVID-19 EUROPARL dataset v1. Bilingual (EN-PT)

Bilingual (EN-PT) corpus acquired from the website (https://www.europarl.europa.eu/) of the European Parliament (25th April 2020)

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Portuguese

COVID-19 EUROPARL v2 dataset. Bilingual (EN-PT)

Bilingual (EN-PT) corpus acquired from the website (https://www.europarl.europa.eu/) of the European Parliament (9th May 2020)

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Portuguese

COVID-19 Parallel Global Voices dataset. Bilingual (EN-PT)

EN-PT Bilingual COVID-19-related corpus acquired from the website (https://globalvoices.org/) of GlobalVoices (28th April 2020)

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Portuguese

COVID-19 EU presscorner v2 dataset. Bilingual (EN-PT)

Bilingual (EN-PT) corpus acquired from website (https://ec.europa.eu/commission/presscorner/) of the EU portal (8th July 2020).

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Portuguese

LX-UTagger

LX-UTagger is a POS tagger for Portuguese that adopts the Universal Part-of-Speech tagset (UPOS), related to the Universal Dependency framework, with an initial performance of 99.06% under a ten-fold cross validation scheme. It is described in this article: António Branco, João Ricardo Silv...

Resource Type:	Tool / Service
Language:	Portuguese

CRPC-Named Entity Recognizer

A NER-classifier based on memory-based learning, trained on the CINTIL dataset, a corpus that contains part of the Corpus de Referência do Português Contemporâneo - CRPC (Reference Corpus of Contemporary Portuguese). https://portulanclarin.net/repository/browse/cintil-corpus-internacional-do-por...

Resource Type:	Tool / Service
Language:	Portuguese

Order by:

Filter by: