Search and Browse – PORTULAN CLARIN

International Agreements (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. International Agreements have been translated into natio...

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Latvian

Parallel corpus from Parliament of Estonia (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel corpus compiled from contents of website of Par...

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Estonian

EUIPO - IP case law French-English (Processed)

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	French

Natolin European Centre Dataset (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. The Polish-English parallel corpus is composed of three ...

Resource Type:	Corpus
Media Type:	Text
Language:	Polish

Dundee GCG-Bank

Dundee GCG-Bank contains hand-corrected deep syntactic annotations for the Dundee eye-tracking corpus (Kennedy et al., 2003). The annotations are designed to support psycholinguistic investigation into the structural determinants of sentence processing effort. Dundee GCG-Bank is distributed as a ...

Resource Type:	Corpus
Media Type:	Text
Language:	English

Datasets for classification experiments IS-pros

Datasets is arff format (for Weka machine learning software) are made available to reproduce the validation experiments presented in the paper.

Resource Type:	Corpus
Media Type:	Text
Language:	English

LX-WordSim-353

The LX-WordSim-353 was created from WordSim-353 (Agirre et al., 2009). As the name suggests, this data set contains 353 pairs of words. Both words in each pair can have different morphosyntactic categories. The data set is made of nouns, adjectives, verbs and named entities, and has no multiwords...

Resource Type:	Corpus
Media Type:	Text
Language:	Portuguese

COVID-19 EU presscorner v1 dataset. Multilingual (CEF languages)

Multilingual (CEF languages) corpus acquired from website (https://ec.europa.eu/commission/presscorner/) of the EU portal (14th May 2020). It contains 23 TMX files (EN-X, where X is a CEF language) with 83217 TUs in total.

Resource Type:	Corpus
Media Type:	Text
Languages:	Bulgarian
	Croatian
	Czech
	Danish
	Dutch; Flemish
	English
	Estonian
	Finnish
	French
	German
	Greek, Modern (1453-)
	Hungarian
	Irish
	Italian
	Latvian
	Lithuanian
	Maltese
	Moldavian; Moldovan
	Polish
	Portuguese
	Romanian
	Slovak
	Slovenian
	Spanish; Castilian
	Swedish

Code-switched English-Spanish Tweets

This package contains the collection of tweets described in the LREC 2018 paper: "Collecting Code-Switched Data from Social Media", Gideon Mendels, Victor Soto, Aaron Jaech and Julia Hirschberg, LREC 2018. Please remember to cite this paper if you use this resource. The tagged_tweets_ids file con...

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Spanish; Castilian

Memorias de traducción Portal oficial de turismo de España www.spain.info

Memoria de traducción Portal oficial de turismo de España www.spain.info

Resource Type:	Corpus
Media Type:	Text
Languages:	English
	French
	German
	Italian
	Portuguese
	Spanish; Castilian

Order by:

Filter by: