Yorùbá Text and Speech Corpora
A comprehensive scholarly survey of documented Yorùbá lexical databases, parallel machine translation corpora, named entity benchmarks, and speech datasets.
Yorùbá computational linguistics and corpus development comprise a structured body of lexical resources, parallel translation bitexts, sequence-labelling benchmarks, and acoustic speech recordings. These resources document both Standard Yorùbá (èdè Yorùbá gbogbogbò) and diaspora varieties across specific domains, recording modalities, and licensing frameworks. The ecosystem spans curated lexical databases, crowdsourced acoustic speech sets, studio-recorded single-speaker read speech, multi-domain machine translation corpora, and web-scale mined texts.
The development of Yorùbá language resources is characterized by a transition from early proprietary lexical compilations and heavily skewed religious translations to community-driven, open-access benchmarks designed to address diacritical preservation, tonal ambiguity, and domain imbalance. This document provides a descriptive catalogue and analytical overview of all verified Yorùbá language corpora, their quantitative parameters, legal licensing, and methodological limitations.
Lexical Databases and Dictionaries
Lexical databases provide the baseline inventory of lemmas, phonetic properties, glosses, and morphological patterns required for computational analysis. The most extensive single lexical database compiled for the language is the Global Yorùbá Lexical Database.
Global Yorùbá Lexical Database (v. 1.0)
Compiled by Yíwọlá Awóyalé and published through the Linguistic Data Consortium (LDC) in 2008, the Global Yorùbá Lexical Database is the largest structured electronic dictionary and lexical repository for Yorùbá and related varieties .
- Total Volume: Over 450,000 lexical entries and headwords .
- Sub-corpora Distribution: Over 368,000 entries of Standard Yorùbá; approximately 8,000 entries of Lucumí (the liturgical variety preserved in Cuba); approximately 3,600 entries of Gullah; and approximately 1,000 entries of Trinidadian Yorùbá .
- Licensing and Distribution: Proprietary. Access is managed strictly through standard LDC User Agreements for members and non-members under catalog identifier LDC2008L03 .
- Linguistic Architecture: The database systematically incorporates tone markings and diacritics across standard and diaspora variants, documenting morphological variations, compound formations, reduplications, and dialectal divergences. The inclusion of Atlantic diaspora varieties provides lexical preservation for comparative historical linguistics and creolistics .
Database Metric Comparison:
+------------------------------------+--------------------------+
| Sub-Variety | Lexical Entries (approx) |
+------------------------------------+--------------------------+
| Standard Yorùbá | 368,000+ |
| Lucumí (Cuba) | 8,000 |
| Gullah (Sea Islands, US) | 3,600 |
| Trinidadian Yorùbá | 1,000 |
| Total Global Yorùbá Lexicon | 450,000+ |
+------------------------------------+--------------------------+
Named Entity Recognition and Sequence Tagging Benchmarks
Until the development of targeted grassroots initiatives, sequence labelling and information extraction tasks in Yorùbá suffered from an absence of manually verified human gold standards.
MasakhaNER
Published in 2021 by David Ifeoluwa Adelani, Jesujoba O. Alabi, Angela Fan, Julia Kreutzer, and the Masakhane research collective, MasakhaNER established the standard human-annotated named entity recognition (NER) benchmark for African languages, with Yorùbá as a primary focus .
- Dataset Size and Splits: The Yorùbá sub-corpus comprises 9,824 manually annotated sentences drawn from contemporary local news publications, containing approximately 165,000 tokens .
- Training Split: 6,877 sentences.
- Development Split: 983 sentences.
- Test Split: 1,964 sentences.
- Tagset: Standard four-class flat named entity categories, encompassing Personal Names (PER), Locations (LOC), Organizations (ORG), and Dates (DATE) .
- Licensing: Released under an open framework where the annotations are licensed under Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), with the aggregate dataset distributed under Creative Commons Attribution 4.0 International (CC BY 4.0), subject to host domain terms for underlying news texts .
- Methodological Significance: MasakhaNER addressed the persistent failure of multilingual zero-shot transfer from non-tonal Indo-European models by demonstrating that in-language, human-annotated tokens carrying correct tone and sub-dot marks produce superior boundary detection and classification accuracy .
Parallel Bitext Corpora for Machine Translation
Machine translation (MT) models require aligned parallel text across source and target languages. Yorùbá parallel resources are divided into curated multi-domain benchmarks, religious open corpora, professional multi-way evaluations, and uncurated web-mined text.
Corpus Comparison: Machine Translation Bitexts
+--------------+------------------+--------------------------+-----------------------+
| Corpus | Sentence Pairs | Primary Domain | Licence |
+--------------+------------------+--------------------------+-----------------------+
| MENYO-20k | 20,100 | News, TED, Proverbs, TV | CC BY-NC 4.0 |
| MAFAND-MT | Multi-thousand | Human-translated News | CC BY / Open |
| JW300 | Large-scale (MT) | Religious Publications | Distribution Restrict |
| FLORES-101 | 1,012 (dev/test) | Multilingual Wikipedia | CC BY-SA 4.0 |
| CCAligned | Multi-million | Web Crawled (CommonCrawl)| Web Terms (Uncleaned) |
+--------------+------------------+--------------------------+-----------------------+
MENYO-20k
Created by David Ifeoluwa Adelani, Dana Ruiter, Jesujoba O. Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, Cristina España-Bonet, and Dietrich Klakow, MENYO-20k was designed to resolve the extreme domain over-specialization that previously affected Yorùbá translation research .
- Size and Structure: 20,100 parallel English–Yorùbá sentence pairs partitioned into 10,070 training, 3,397 development, and 6,633 test sentences .
- Domain Coverage: A clean, multi-domain balance drawn from news articles (Global Voices, Voice of Nigeria), TED talks, movie transcripts, radio broadcasts, digital literature, and traditional Yorùbá proverbs (òwe) .
- Orthographic Normalization: The dataset manually standardizes tone marks (high, mid, low) and sub-dots (ẹ, ọ, ṣ) to counteract the non-standard diacritization prevalent in raw digital media .
- Licensing: Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), governed by downstream terms of multi-source redistribution .
MAFAND-MT
Published in 2022 by David Ifeoluwa Adelani and the Masakhane network, MAFAND-MT (Masakhane A Few Thousand Translations) introduced high-quality, professional human translations specifically focused on news domains across African languages .
- Architecture: Explicitly pairs Yorùbá with both English and French to test direct vehicular-to-indigenous language translation pathways rather than routing exclusively through English .
- Quality Benchmark: Evaluated to show that even a few thousand high-precision, human-verified, diacritized parallel sentences outperform hundreds of thousands of noisy web-scraped sentences in neural machine translation (NMT) fine-tuning .
JW300
Extracted automatically by Željko Agić and Ivan Vulić in 2019, JW300 was generated by scraping parallel articles, magazines, and leaflets from jw.org across hundreds of low-resource languages, including Yorùbá .
- Volume: One of the largest single parallel corpora for Yorùbá in terms of sentence counts, enabling early deep neural MT baseline experiments .
- Distribution and Legal Status: The distribution of JW300 was subsequently halted and restricted due to formal copyright infringement challenges brought by the source copyright holders .
- Domain Limitation: The language register is exclusively religious and moral didactic prose. The syntax and vocabulary mirror translated Christian literature, which skews statistical models away from natural everyday conversational and colloquial speech patterns .
CCAligned and CCMatrix
Generated via automated extraction pipelines by Ahmed El-Kishky et al. (2020) and Holger Schwenk et al. (2021), CCAligned and CCMatrix mined parallel text pairs from Common Crawl web archives using cross-lingual document and sentence embeddings .
- Characteristics: Massive token counts, providing raw web text for broad exploratory mining .
- Documented Deficiencies: Scholarly evaluation reveals high rates of sentence misalignment, language misclassification (such as mixing Yorùbá with other West African languages), structural truncation, and the systematic stripping of tone marks and diacritical sub-dots .
FLORES-101 and FLORES-200
Constructed by Nsikak Goyal et al. (2022) and extended by the NLLB Team (2022), the FLORES evaluation benchmark consists of professional human translations of 101 to 200+ languages based on English Wikipedia articles .
- Role: Serves as a standard, multi-way parallel evaluation split allowing strictly identical content to be compared across diverse language families under normalized orthographic standards .
Speech and Acoustic Corpora
Speech datasets require audio recordings paired with precise orthographic transcriptions. Yorùbá speech technology datasets vary in speaker diversity, acoustic sample rates, recording environments, and total hours.
Acoustic Speech Datasets: Parameters and Licensing
+---------------------+-------------+-------------+------------+--------------------+
| Dataset | Audio Hours | Sample Rate | Speakers | Licence |
+---------------------+-------------+-------------+------------+--------------------+
| BibleTTS (SLR129) | ~33.3 hrs | 48 kHz | 1 (studio) | CC BY-SA 4.0 |
| Mozilla Common Voice| ~6–10+ hrs | Variable | Mult./Crowd| CC0 1.0 (Public) |
| OpenSLR SLR86 | ~4.1 hrs | 48 kHz | 36 | CC BY-SA 4.0 |
| Lagos-NWU | ~2.75 hrs | 16 kHz | 33 | CC BY 2.5 ZA |
| ALFFA | 0 hrs (N/A) | N/A | N/A | MIT (No Yorùbá) |
+---------------------+-------------+-------------+------------+--------------------+
OpenSLR SLR86 (Crowdsourced Yorùbá Speech Dataset)
Released in 2020 by Alexander Gutkin, Işın Demirşahin, Oddur Kjartansson, Clara Rivera, and Kọ́lá Túbọ̀sún, SLR86 is an open-source speech corpus specifically designed to support speech recognition (ASR) and text-to-speech (TTS) development .
- Recorded Audio: Approximately 4.1 hours of high-fidelity, crowdsourced audio .
- Acoustic Details: Recorded at a native sampling rate of 48 kHz from 36 native Yorùbá speakers reading prompt texts .
- Licensing: Distributed openly under the Creative Commons Attribution-ShareAlike 4.0 International license (CC BY-SA 4.0) via the OpenSLR repository .
- Quality Assurance: Prompts were systematically rendered with verified tone marks and diacritics to ensure that pitch contours in the acoustic recording accurately matched the orthographic transcription .
BibleTTS (Yorùbá Subset / SLR129)
Introduced in 2022 by Josh Meyer, David Ifeoluwa Adelani, Edresson Casanova, Alp Öktem, Daniel Whitenack, Julian Weber, and collaborators, BibleTTS represents a large studio-recorded speech resource for African language synthesis .
- Recorded Audio: Approximately 33.3 hours of clean, single-speaker studio audio derived and segmented from over 80 hours of raw high-quality Scripture readings .
- Acoustic Details: High-fidelity 48 kHz audio, precisely segmented and aligned at the biblical verse level .
- Licensing: Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) hosted as OpenSLR dataset SLR129 .
- Limitations: The dataset is restricted to a single speaker, limiting acoustic variance for multi-speaker recognition, and uses specialized ecclesiastical language registers .
Lagos-NWU Yorùbá Speech Corpus
Developed collaboratively in 2012 by Daniel R. van Niekerk, Etienne Barnard, Oluwapelumi Giwa, and Azeez Sosimi through a partnership between the Centre for Text Technology (CTexT) at North-West University and the University of Lagos, this corpus provided an early acoustic benchmark for speech processing .
- Recorded Audio: Approximately 2.75 hours of speech .
- Acoustic Details: Recorded at 16 kHz across 33 speakers (16 female, 17 male), with each participant reading approximately 130 balanced phonetically designed utterances .
- Licensing: Creative Commons Attribution 2.5 South Africa (CC BY 2.5 ZA) .
- Phonetic Analysis: Accompanied by analytical studies documenting pitch track realization, fundamental frequency ($F_0$) tonal alignment, and vowel quality variations across tonal registers .
Mozilla Common Voice (Yorùbá)
Mozilla Common Voice is a massive, crowdsourced, publicly accessible multilingual platform published under public domain terms .
- Volume: Accumulates progressively through public recording and validation interfaces. Initial integration provided approximately 6 validated hours, with newer corpus snapshots expanding to 6 to 10+ validated hours .
- Licensing: Creative Commons Zero Universal 1.0 (CC0 1.0, Public Domain Dedication) .
- Acoustic Diversity: Features broad demographic diversity, capturing varied microphones, background noise environments, accents, and age profiles, though varying transcription fidelity requires automated filtering .
ALFFA Project Record Clarification
The African Languages in the Field: speech Fundamentals and Automation (ALFFA) project (Besacier et al., 2016) is frequently cited in broader African NLP literature as a seminal open ASR effort .
- Record Note: The primary ALFFA repository contains no Yorùbá speech dataset .
- Documented Coverage: The ALFFA project officially compiled, documented, and released speech corpora for Amharic, Swahili, Wolof, Fongbe, and Hausa (such as OpenSLR SLR25) .
- Scholarly Clarification: While secondary research reviews occasionally associate ALFFA generally with sub-Saharan speech initiatives, no primary Yorùbá audio corpus was produced under the project .
The Lacuna Fund and Infrastructural Grants
The Lacuna Fund, managed via the Meridian Institute, the International Development Research Centre (IDRC), and associated organizations, operates as a targeted funding mechanism to bridge data gaps in low-resource contexts . It does not constitute a single standalone dataset, but rather funds a series of open-access Yorùbá NLP resources :
- NaijaSenti: A large-scale Twitter sentiment analysis corpus comprising approximately 30,000 human-annotated Yorùbá tweets .
- MasakhaNER 2.0 / African POS: Expanded named entity recognition datasets and part-of-speech (POS) annotations spanning diverse contemporary news and digital domains .
- AFRIDOC-MT: A curated benchmark of 605 parallel English–African documents providing document-level parallel context including Yorùbá .
- NaijaVoices: A large-scale speech initiative targeting approximately 1,000 hours of multi-speaker speech data across Yorùbá, Igbo, and Hausa .
- Licensing Policy: Mandates open licensing frameworks, primarily utilizing Creative Commons Attribution 4.0 International (CC BY 4.0) with non-commercial designations (CC BY-NC 4.0) applied where source data restrictions require it .
Linguistic Challenges and Open Issues in Digital Corpora
The curation of Yorùbá corpora presents linguistic, orthographic, and structural challenges that distinguish tonal language processing from non-tonal language processing.
Tone and Vowel Marking in Standard Yorùbá Orthography:
+--------+------------------+------------------+-----------------------+
| Vowel | High Tone (Ó) | Mid Tone (O) | Low Tone (Ò) |
+--------+------------------+------------------+-----------------------+
| e | é [e˥] | e [e˧] | è [e˩] |
| ẹ | ẹ́ [ɛ˥] | ẹ [ɛ˧] | ẹ̀ [ɛ˩] |
| o | ó [o˥] | o [o˧] | ò [o˩] |
| ọ | ọ́ [ɔ˥] | ọ [ɔ˧] | ọ̀ [ɔ˩] |
+--------+------------------+------------------+-----------------------+
Diacritization and Tonal Ambiguity
Standard Yorùbá employs three contrastive register tones: High (marked with an acute accent: ´), Mid (unmarked), and Low (marked with a grave accent: `) . In addition, it utilizes sub-dots or vertical lines beneath specific vowels and consonants to distinguish open-mid vowels from close-mid vowels, and the postalveolar fricative from the alveolar sibilant:
- ẹ ([ɛ]) vs e ([e])
- ọ ([ɔ]) vs o ([o])
- ṣ ([ʃ]) vs s ([s])
When web scrapers or uncurated pipelines extract text from digital platforms, these diacritical marks are frequently stripped due to keyboard encoding limitations or character misinterpretations . The stripping of diacritics generates severe semantic ambiguity. For example, the bare string owo can represent four distinct lexical items depending entirely on its tone and sub-dot configurations:
- owó (money / currency)
- ọwọ́ (hand / arm)
- ọ̀wọ̀ (respect / honor / broom)
- òwò (trade / business / commerce)
Machine translation models trained on undiacritized bitexts produce degraded outputs because the loss of tone and sub-dots collapses distinct semantic categories into identical surface forms .
Domain Imbalance
A central tension in Yorùbá corpus design is domain skew. The largest publicly available parallel resources historically derived from religious institutions (such as the 1884 Bible translation and the JW300 dataset) . Consequently, early statistical and neural models exhibited high performance on theological and archaic literary syntax, but degraded when evaluated on contemporary legal, biomedical, computational, or political texts. The introduction of MENYO-20k, MasakhaNER, and MAFAND-MT has partially mitigated this gap by establishing news, dialogue, and web benchmarks .
Licensing Barriers and Open Access
A practical division exists between proprietary corpora and open community benchmarks:
- Proprietary datasets (such as the Global Yorùbá Lexical Database under LDC licensing) require institutional subscriptions, restricting access for independent or local Nigerian developers .
- Copyright disputes (such as the takedown and distribution restrictions of JW300) have underscored the precariousness of scraping commercial or organizational domains without explicit open licenses .
- Open initiatives (such as OpenSLR, Common Voice, and Masakhane projects under CC BY, CC BY-SA, and CC0 licenses) have established accessible repositories that allow unencumbered redistribution and retraining of speech and text models .