N3-Collection

We publish three novel datasets called N3. N3 will be published using NIF ensuring a greater interoperability to overcome the need for corpus-specific parsers. The data can be downloaded from our project homepage.

Resource Type:Corpus
Media Type:Text
Languages:English
German
PKN Orlen Dataset (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Dataset of the Polish public sector company PKN Orlen, a...

Resource Type:Corpus
Media Type:Text
Languages:English
Polish
Datasets for classification experiments IS-pros

Datasets is arff format (for Weka machine learning software) are made available to reproduce the validation experiments presented in the paper.

Resource Type:Corpus
Media Type:Text
Language:English
CW Corpus

The Complex Word (CW) Corpus contains 731 sentences each with one annotated CW. These simplifications were mined from Simple Wikipedia edit histories. Each entry gives an example of a sentence requiring simplification by means of a single lexical edit. This resource is primarily designed for t...

Resource Type:Corpus
Media Type:Text
Language:English
COVID-19 - HEALTH Wikipedia dataset. Bilingual (EN-PT)

Bilingual (EN-PT) corpus acquired from Wikipedia on health and COVID-19 domain (2nd May 2020)

Resource Type:Corpus
Media Type:Text
Languages:English
Portuguese
Parallel texts from Swedish Work environment Authority (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts from the Swedish Work Environment authori...

Resource Type:Corpus
Media Type:Text
Languages:Bulgarian
Czech
English
Estonian
Finnish
French
German
Greek, Modern (1453-)
Hungarian
Italian
Latvian
Lithuanian
Polish
Romanian
Spanish; Castilian
Swedish
Trilingual Documents related to International Judicial Cooperation in Civil Matters (Greek-English-French) (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Trilingual (Greek-English-French) documents - standard f...

Resource Type:Corpus
Media Type:Text
Languages:English
French
Greek, Modern (1453-)
Letter of rights for persons arrested on the basis of a European Arrest Warrant (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Letter of rights for persons arrested on the basis of a ...

Resource Type:Corpus
Media Type:Text
Languages:Bulgarian
Dutch; Flemish
English
French
German
Greek, Modern (1453-)
Italian
Latvian
Polish
Romanian
Letter of rights for persons arrested and or detained (Processed)  

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Collection of transaltion units (1906 in total) in 21 la...

Resource Type:Corpus
Media Type:Text
Languages:Bulgarian
English
French
Greek, Modern (1453-)
Latvian
Polish
Romanian
Hallituskausi 2007-2011 -- Finnish-English Translation Memory (Processed)

ID: ELRA-W0220 This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. The "Hallituskausi 2007–2011" translat...

Resource Type:Corpus
Media Type:Text
Languages:English
Finnish

Order by:

Filter by: