Search and Browse – PORTULAN CLARIN

U-Compare Discourse Parsing Service

Web service created by exporting UIMA-based workflow from the U-Compare text mining system. Functionality: Performs discourse parsing on plain text. Also identifies sentences, tokens, parts of speech, lemmas, clauses and coreference chains Tools in workflow: UAIC-POSTagger, UAIC-NPChunker, UAI...

Resource Type:	Tool / Service
Language:	Romanian

U-Compare syntactic chunking service

Web service created by exporting UIMA-based workflow from the U-Compare text mining system. Functionality: Identifies and categorises syntactic chunks in plain text Tools in workflow: Freeling shallow parser web service (service provided by the PANACEA project) NOTE: The licence provided cove...

Resource Type:	Tool / Service
Language:	Galician

Polish-English parallel corpus from the website of the Citizens Information Board (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Polish-English parallel corpus from the website of the C...

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Polish

Dicionário de Gentílicos e Topónimos

Dicionário de Gentílicos e Topónimos is a list of pairs of toponyms and demonyms. The toponyms and demonyms included have a morphologically compositional relation between each other. The list contains around 1500 such pairs and additionally provides information on the toponym referent (upper unit...

Resource Type:	Lexical / Conceptual
Media Type:	Text
Language:	Portuguese

U-Compare Syntactic Parsing Service

Web service created by exporting UIMA-based workflow from the U-Compare text mining system. Functionality: Carries out syntactic parsing on plain text Tools in workflow: Cafetiere Sentence Splitter (University of Manchester), OpenNLP Tokenizer (Apache), STEPP Tagger (University of Manchester), ...

Resource Type:	Tool / Service
Language:	English

Bulgarian-English Wikipedia WSD/NED corpus

Bulgarian-English Wikipedia WSD/NED corpus is composed of articles from the Bulgarian version of Wikipedia and their English counterparts.

Resource Type:	Corpus
Media Type:	Text
Languages:	Bulgarian
Languages:	English

FEUP Tweets

Tweet corpus

Resource Type:	Corpus
Media Type:	Text
Language:	English

Georeferenced Tweets

Tweets annotated with geographic coordinates

Resource Type:	Corpus
Media Type:	Text
Language:	English

Portuguese Parliamentary Corpus 4.0

The Portuguese Parliamentary Corpus is part of the Mutlilingual ParlaMint Corpus, a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions. The Portuguese corpus (ParlaMint-PT) comprehends transcripts of sessions in the time pe...

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

Biographies of Portuguese People

This is a set of 11.361 biographies of Portuguese people. The compilation of the data involved the biography collection from wikipedia and data conversion. Several filters were applied to remove entries that were mostly empty or non applicable content. Format: JSON (conversion from HTML) ...

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

Order by:

Filter by: