N3-Collection

We publish three novel datasets called N3. N3 will be published using NIF ensuring a greater interoperability to overcome the need for corpus-specific parsers. The data can be downloaded from our project homepage.

Resource Type:Corpus
Media Type:Text
Languages:English
German
National Health Fund Dataset (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. The dataset is a 274K-token Polish-English parallel reso...

Resource Type:Corpus
Media Type:Text
Languages:English
Polish
SETimes.HR

We present SETimes.HR ― the first linguistically annotated corpus of Croatian that is freely available for all purposes. The corpus is built on top of the SETimes parallel corpus of nine Southeast European languages and English. It is manually annotated for lemmas, morphosyntactic tags, named ent...

Resource Type:Corpus
Media Type:Text
Language:Czech
Luxembourg Museum Websites (de-en) (Processed)

Luxembourg Museum Websites (de-en) (Processed)

Resource Type:Corpus
Media Type:Text
Languages:English
French
German
Elhuyar-QTLeap WSD/NED corpus

This corpus is created from documents from translation memorios of Elhuyar Fundation (obtained via Eleka, member of the Advisory Board of Potential Users).

Resource Type:Corpus
Media Type:Text
Languages:Basque
English
LX-Rare Word Similarity Dataset

The LX-Rare Word Similarity Data set was created from Stanford Rare Word (RW) Similarity data set (Luong et al., 2013). This list contains 2 034 words (1 017 pairs of words). All the words were extracted from Wikipedia and from WordNet (Miller, 1995), a lexical database where the concepts are gro...

Resource Type:Corpus
Media Type:Text
Language:Portuguese
RedditPT Dataset

This dataset is a collection of dialogues extracted from the Portugal subreddit with RDET (Reddit Dataset Extraction Tool). It is composed of around 58,964,715 tokens in 218,550 dialogues.

Resource Type:Corpus
Media Type:Text
Language:Portuguese
Hesita-POS

Hesita-POS is an annotaded corpus. Tv News.

Resource Type:Corpus
Media Type:Text
Language:Portuguese
Czech Banking Association Terminology (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Terms in Czech - English relating to finance

Resource Type:Corpus
Media Type:Text
Languages:Czech
English
GREC

GREC is a semantically annotated corpus of 240 MEDLINE abstracts (167 on the subject of E. coli species and 73 on the subject of the Human species) which is intended for training IE systems and/or resources which are used to extract events from biomedical literature.

Resource Type:Corpus
Media Type:Text
Language:English

Order by:

Filter by:

Text (440)
Audio (18)
Image (1)