Yorùbá Text and Speech Corpora
A comprehensive scholarly survey of documented Yorùbá lexical databases, parallel machine translation corpora, named entity benchmarks, and speech datasets.
A comprehensive scholarly survey of documented Yorùbá lexical databases, parallel machine translation corpora, named entity benchmarks, and speech datasets.
The Yorùbá wey dem quote, the proverbs, oríkì, ẹsẹ Ifá, word list headwords, Odù names and citations dey exactly as the corpus record dem, for every language.
Yorùbá computational linguistics and corpus development comprise one structured collection of lexical resources, parallel translation bitexts, sequence-labelling benchmarks, and acoustic voice recordings. Dem resources dey document both Standard Yorùbá (èdè Yorùbá gbogbogbò) and diaspora varieties across specific domains, recording modalities, and licensing frameworks. The ecosystem spread across curated lexical databases, crowdsourced acoustic voice sets, studio-recorded single-speaker read speech, multi-domain machine translation corpora, and text wey dem mine from web at large scale.
How Yorùbá language resources take develop show say e shift from early proprietary lexical compilations and religious translations wey heavily balance one side go community-driven, open-access benchmarks wey dem design to handle diacritical marks preservation, tonal confusion, and domain imbalance. Dis document provide descriptive catalogue and analytical overview of all verified Yorùbá language corpora, dia quantitative parameters, legal licensing, and methodological limitations.
Lexical databases dey provide the baseline inventory of lemmas, phonetic properties, glosses, and morphological patterns wey dey necessary for computational analysis. The single lexical database wey big pass wey dem compile for the language na the Global Yorùbá Lexical Database.
As Yíwọlá Awóyalé compile am and Linguistic Data Consortium (LDC) publish am for 2008, the Global Yorùbá Lexical Database na the biggest structured electronic dictionary and lexical repository for Yorùbá and related varieties .
Database Metric Comparison:
+------------------------------------+--------------------------+
| Sub-Variety | Lexical Entries (approx) |
+------------------------------------+--------------------------+
| Standard Yorùbá | 368,000+ |
| Lucumí (Cuba) | 8,000 |
| Gullah (Sea Islands, US) | 3,600 |
| Trinidadian Yorùbá | 1,000 |
| Total Global Yorùbá Lexicon | 450,000+ |
+------------------------------------+--------------------------+
Until targeted grassroots initiatives come up, sequence labelling and information extraction tasks for Yorùbá suffer because human gold standards wey dem verify by hand no dey.
As David Ifeoluwa Adelani, Jesujoba O. Alabi, Angela Fan, Julia Kreutzer, and the Masakhane research collective publish am for 2021, MasakhaNER establish the standard human-annotated named entity recognition (NER) benchmark for African languages, with Yorùbá as major focus .
Machine translation (MT) models need aligned parallel text across source and target languages. Dem divide Yorùbá parallel resources into curated multi-domain benchmarks, religious open corpora, professional multi-way evaluations, and web-mined text wey dem no curate.
Corpus Comparison: Machine Translation Bitexts
+--------------+------------------+--------------------------+-----------------------+
| Corpus | Sentence Pairs | Primary Domain | Licence |
+--------------+------------------+--------------------------+-----------------------+
| MENYO-20k | 20,100 | News, TED, Proverbs, TV | CC BY-NC 4.0 |
| MAFAND-MT | Multi-thousand | Human-translated News | CC BY / Open |
| JW300 | Large-scale (MT) | Religious Publications | Distribution Restrict |
| FLORES-101 | 1,012 (dev/test) | Multilingual Wikipedia | CC BY-SA 4.0 |
| CCAligned | Multi-million | Web Crawled (CommonCrawl)| Web Terms (Uncleaned) |
+--------------+------------------+--------------------------+-----------------------+
David Ifeoluwa Adelani, Dana Ruiter, Jesujoba O. Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, Cristina España-Bonet, and Dietrich Klakow na dem create MENYO-20k to solve di serious domain over-specialization wahala wey affect Yorùbá translation research before .
David Ifeoluwa Adelani and di Masakhane network publish MAFAND-MT (Masakhane A Few Thousand Translations) for 2022, and e bring high-quality, professional human translations wey focus specifically on news domain across African languages .
Željko Agić and Ivan Vulić extract JW300 automatically for 2019, through scraping of parallel articles, magazines, and leaflets from jw.org across hundreds of low-resource languages, including Yorùbá .
Ahmed El-Kishky et al. (2020) and Holger Schwenk et al. (2021) generate CCAligned and CCMatrix through automated extraction pipelines, where dem mine parallel text pairs from Common Crawl web archives using cross-lingual document and sentence embeddings .
Nsikak Goyal et al. (2022) construct di FLORES evaluation benchmark and NLLB Team (2022) extend am; e get professional human translations of 101 to 200+ languages based on English Wikipedia articles .
Speech datasets need audio recordings wey dem pair with correct orthographic transcriptions. Yorùbá speech technology datasets get differences for speaker diversity, acoustic sample rates, recording environments, and total hours.
Acoustic Speech Datasets: Parameters and Licensing
+---------------------+-------------+-------------+------------+--------------------+
| Dataset | Audio Hours | Sample Rate | Speakers | Licence |
+---------------------+-------------+-------------+------------+--------------------+
| BibleTTS (SLR129) | ~33.3 hrs | 48 kHz | 1 (studio) | CC BY-SA 4.0 |
| Mozilla Common Voice| ~6–10+ hrs | Variable | Mult./Crowd| CC0 1.0 (Public) |
| OpenSLR SLR86 | ~4.1 hrs | 48 kHz | 36 | CC BY-SA 4.0 |
| Lagos-NWU | ~2.75 hrs | 16 kHz | 33 | CC BY 2.5 ZA |
| ALFFA | 0 hrs (N/A) | N/A | N/A | MIT (No Yorùbá) |
+---------------------+-------------+-------------+------------+--------------------+
Alexander Gutkin, Işın Demirşahin, Oddur Kjartansson, Clara Rivera, and Kọ́lá release SLR86 for 2020 Túbọ̀sún, as open-source speech corpus wey dem design specifically to support speech recognition (ASR) and text-to-speech (TTS) development .
BibleTTS, wey Josh Meyer, David Ifeoluwa Adelani, Edresson Casanova, Alp Öktem, Daniel Whitenack, Julian Weber, plus people wey follow dem work bring come out for 2022, na one big studio-recorded voice resource for African language speech synthesis .
Daniel R. van Niekerk, Etienne Barnard, Oluwapelumi Giwa, and Azeez Sosimi work together develop this corpus for 2012 through partnership between the Centre for Text Technology (CTexT) for North-West University and University of Lagos, and e provide early acoustic standard for speech processing .
Mozilla Common Voice na one massive, crowdsourced, publicly accessible multilingual platform wey dem publish under public domain terms .
People wey dey do research for broader African NLP literature dey often cite the African Languages in the Field: speech Fundamentals and Automation (ALFFA) project (Besacier et al., 2016) as one major open ASR work .
The Lacuna Fund, wey dem dey manage through the Meridian Institute, the International Development Research Centre (IDRC), and other partner organizations, dey work as targeted funding way to bridge data gaps for low-resource environments . E no be just one single standalone dataset, but e dey fund different open-access Yorùbá NLP resources :
The work to build and manage Yorùbá corpora dey bring linguistic, orthographic, and structural challenges wey make tonal language processing dey different from non-tonal language processing.
Tone and Vowel Marking in Standard Yorùbá Orthography:
+--------+------------------+------------------+-----------------------+
| Vowel | High Tone (Ó) | Mid Tone (O) | Low Tone (Ò) |
+--------+------------------+------------------+-----------------------+
| e | é [e˥] | e [e˧] | è [e˩] |
| ẹ | ẹ́ [ɛ˥] | ẹ [ɛ˧] | ẹ̀ [ɛ˩] |
| o | ó [o˥] | o [o˧] | ò [o˩] |
| ọ | ọ́ [ɔ˥] | ọ [ɔ˧] | ọ̀ [ɔ˩] |
+--------+------------------+------------------+-----------------------+
Standard Yorùbá dey use three contrastive register tones: High (marked with acute accent: ´), Mid (no mark), and Low (marked with grave accent: `) . Aside from that, e dey use sub-dots or vertical lines under specific vowels and consonants to distinguish open-mid vowels from close-mid vowels, and the postalveolar fricative from the alveolar sibilant:
Wen web scrapers or uncurated pipelines extract text from digital platforms, dis diacritical marks dey often commot because of keyboard encoding limitations or wen dem misinterpret characters . Dis removal of diacritics dey create serious semantic ambiguity. For example, di bare string owo fit represent four different lexical items depending totally on di tone and sub-dot configurations wey e get:
Machine translation models wey dem train on undiacritized bitexts dey produce degraded outputs because as di tone and sub-dots commot, e dey collapse different semantic categories into di same surface forms .
One major tension for Yorùbá corpus design na domain skew. Historically, di biggest parallel resources wey dey publicly available come from religious institutions (like di 1884 Bible translation and di JW300 dataset) . As a result, early statistical and neural models show high performance on theological and old literary syntax, but dia performance drop wen dem test dem on modern legal, biomedical, computational, or political texts. Di introduction of MENYO-20k, MasakhaNER, and MAFAND-MT don partly reduce dis gap as dem establish news, dialogue, and web benchmarks .
Practical division dey between proprietary corpora and open community benchmarks:
A scholarly analysis of major reference grammars, foundational syntactic debates, and dialectological classifications in Yorùbá linguistics.
How Yorùbá came to be written in Arabic and then Roman script, who decided the spelling rules, and why the subdots and tone marks are not optional.
A comprehensive scholarly survey of Yorùbá computational linguistics, covering automatic diacritic restoration, machine translation, speech recognition, speech synthesis, and community benchmarks.
The historical development of Yoruba dictionary-making from nineteenth-century missionary foundations through modern computational lexical databases.
An analysis of Yorùbá language pedagogy, instructional materials, second-language tone acquisition research, and mother-tongue education policy in Nigeria.
The Yorùbá press, D. O. Fágúnwà and the Yorùbá novel, Tutuola's contested place, the travelling theatre, Yorùbá film, and where the language stands in education and online.