Search and Browse – PORTULAN CLARIN

LX-Abbreviations

LX-Abbreviations resource is a collection of abbreviations of different types from European Portuguese composed by 208 words. Each type of abbreviation is manually divided and annotated with grammatical categories, gender and number, and, finally, with the respective abbreviations.

Resource Type:	Lexical / Conceptual
Media Type:	Text
Language:	Portuguese

Parallel texts from Swedish Social Security Authority (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts, email templates and forms in pdf file fo...

Resource Type:	Corpus
Media Type:	Text
Languages:	Croatian
	English
	Finnish
	French
	German
	Italian
	Polish
	Romanian
	Spanish; Castilian
	Swedish

GENIA POS & Term Corpus

A corpus of 2,000 MEDLINE abstracts, collected using the three MeSH terms human, blood cells and transcription factors. The corpus is available in three formats: 1) A text file containing part-of-speech (POS) annotation, based on the Penn Treebank format, 2) An XML file containing inline POS anno...

Resource Type:	Corpus
Media Type:	Text
Language:	English

Termcat Fairs and Congresses

Terms for Fairs and Congresses

Resource Type:	Lexical / Conceptual
Media Type:	Text
Languages:	Catalan; Valencian
	English
	French
	German
	Italian
	Portuguese
	Spanish; Castilian

PS corpus (Post-Scriptum) - treebank

PS corpus (Post-Scriptum) - treebank is a treebank corpus of 586 informal mail letters written in Portuguese and Spanish during the Modern Ages (from the XVIth century to the beginning of the XIXth century). This treebank is a syntactically annotated subset of the Portuguese "PS corpus (Post-S...

Resource Type:	Corpus
Media Type:	Text
Languages:	Portuguese
Languages:	Spanish; Castilian

Maltese Wiktionary

This lexicon is part of the collection of the Wikimedia Dumps which was retrieved as an XML file from http://dumps.wikimedia.org/mtwiktionary/20121105/ on November 5, 2012. In the Wikimedia dump, it is accompanied by a text file mtwiktionary-20121105-pages-articles-multistream-index.txt which li...

Resource Type:	Lexical / Conceptual
Media Type:	Text
Language:	Maltese

DVPM-SynSem

DVPM-SynSem is a lexical database with syntactic and semantic information in Medieval Portuguese. It contains around 3000 verbs.

Resource Type:	Lexical / Conceptual
Media Type:	Text
Language:	Portuguese

CIPM-POS

CIPM-POS is a set of historical, religious, notarial, literary texts in prose and verse, written is medieval portuguese. It contains around 88000 words.

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

Parallel texts from Swedish Work environment Authority (Processed)

Resource Type:	Corpus
Media Type:	Text
Languages:	Bulgarian
	Czech
	English
	Estonian
	Finnish
	French
	German
	Greek, Modern (1453-)
	Hungarian
	Italian
	Latvian
	Lithuanian
	Polish
	Romanian
	Spanish; Castilian
	Swedish

U-Compare NP Chunking Service

Web service created by exporting UIMA-based workflow from the U-Compare text mining system. Functionality: Identifies NP chunks in plain text. Also carries out sentence splitting, tokenisation and POS tagging Tools in workflow: MLRS Sentence Splitter (University of Malta), UAIC-POSTagger, UAIC-...

Resource Type:	Tool / Service
Language:	Romanian

Order by:

Filter by: