Search and Browse – PORTULAN CLARIN

The full editions of ILLUM from 12/11/2006 to 30/05/2010 (185 issues).

Resource Type:	Corpus
Media Type:	Text
Language:	Maltese

MalToBi/SPAN Corpus

Audio corpus: 8 subfolders with .wav files Each containing : • 2 sound files containing a read story (“The sun and the wind”, each by speaker A and speaker B) • 2 sound files containing each 30 read sentences (each by speaker A and speaker B) • 2 x each of the 30 sentences as a single sound f...

Resource Type:	Corpus
Media Type:	Audio
Language:	Maltese

F_Mona_1/ Spoken Newspaper

108 WAV files of spoken Maltese newspaper texts, subdivided into 12 directories with a variable number of sentences (sometimes: clauses) each. They come together with transcriptions and tables of phoneme durations.

Resource Type:	Corpus
Media Type:	Audio
Language:	Maltese

Monolingual documents from the Government of Lithuania (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Monolingual documents received from the Government of th...

Resource Type:	Corpus
Media Type:	Text
Language:	Lithuanian

NoSta-D: German NER Dataset Train/Dev

Freely available large dataset, manually annotated for German NER. Includes nested span annotations. Source text from German Wikipedia and news. This data set does not contain the test data, which is used for the GermEval 2014 NER task at KONVENS. Test data will be available from September 2014.

Resource Type:	Corpus
Media Type:	Text
Language:	German

Multilingual corpora with coreferential annotation of person entities

Multilingual corpora with coreferential annotation of person entities ===================================================================== In-progress corpora with coreferent annotation of person entities. Sources: journals and Wikipedia. Languages: * Portuguese: varieties from Portugal, Brazi...

Resource Type:	Corpus
Media Type:	Text
Languages:	Galician
	Portuguese
	Spanish; Castilian

U-Compare syntactic chunking service

Web service created by exporting UIMA-based workflow from the U-Compare text mining system. Functionality: Identifies and categorises syntactic chunks in plain text Tools in workflow: Freeling shallow parser web service (service provided by the PANACEA project) NOTE: The licence provided cove...

Resource Type:	Tool / Service
Language:	Galician

LuxId

Corpus of mixed language (French, German,Luxemburguish) sentences from {sc Chamber} (House of Parliament) debate reports manually annotated at segment level with 6 labels : Lux, Fre, Ger, Lux + Fre, Lux + Ger, Lux + Fre + Ger

Resource Type:	Corpus
Media Type:	Text
Languages:	French
	German
	Luxembourgish; Letzeburgesch

Estonian resource grammar for Grammatical Framework

A GF resource grammar for Estonian, implementing the language-neutral API of the GF Resource Grammar Library as well as a morphological synthesizer.

Resource Type:	Language Description
Media Type:	Text
Language:	Estonian

English-Swedish parallel corpus from the web site of the Swedish Migration Board - Migrationsverket (Processed)

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. All texts have been collected from their website of the ...

Resource Type:	Corpus
Media Type:	Text
Languages:	English
Languages:	Swedish

Order by:

Filter by: