Common Voice dataset

E405893

The Common Voice dataset is a large, open-source multilingual speech corpus created by Mozilla to support and democratize voice recognition research and technology.

All labels observed (1)

Label Occurrences
Common Voice dataset canonical 1

How this entity was disambiguated

Statements (102)

Predicate Object
instanceOf multilingual dataset
open-source dataset
speech corpus
collectionMethod crowdsourcing
creator Mozilla
dataFormat JSON
MP3
WAV
developer Mozilla
downloadURL https://commonvoice.mozilla.org/datasets
goal democratize voice recognition technology
support under-resourced languages
hasFeature accent metadata
speaker age metadata
speaker gender metadata
hasLanguage Afrikaans
Albanian
linked to: Albanian language

Amharic
Arabic
Basque
Belarusian
linked to: Belarusian language

Bengali
Bosnian
Breton
Cantonese
Catalan
Chinese
Croatian
Czech
Danish
Dutch
English
Esperanto
Finnish
linked to: Finnish language

French
Georgian
linked to: Georgian language

German
Greek
Gujarati
Hausa
Hindi
Hungarian
linked to: Hungarian language

Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kabyle
Kannada
Kinyarwanda
Korean
Kurdish
linked to: Kurdish language

Latvian
Lithuanian
Macedonian
Malay
Malayalam
Marathi
linked to: Marathi language

Norwegian
linked to: Norwegian language

Persian
Polish
linked to: Polish language

Portuguese
Punjabi
linked to: Punjabi language

Romanian
linked to: Romanian language

Russian
Scottish Gaelic
Serbian
linked to: Serbian language

Slovak
linked to: Slovak language

Slovenian
linked to: Slovene

Spanish
Swahili
linked to: Swahili language

Swedish
linked to: Swedish language

Tagalog
Tamil
Tatar
linked to: Tatar language

Telugu
Thai
Turkish
Turkmen
linked to: Turkmen language

Ukrainian
Urdu
linked to: Urdu language

Vietnamese
Welsh
Xhosa
Yoruba
Zulu
hasLicense CC0
Creative Commons Zero
linked to: CC0
hasPart audio recordings
speaker metadata
text transcripts
hostedBy Mozilla
inception 2017
isAccessibleForFree true
sponsor Mozilla Foundation
use automatic speech recognition training
language identification research
speaker diarization research
speech technology research
website https://commonvoice.mozilla.org

How these facts were elicited

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

Mozilla operates Common Voice dataset