Yorùbá Natural Language Processing and Speech Technology
A comprehensive scholarly survey of Yorùbá computational linguistics, covering automatic diacritic restoration, machine translation, speech recognition, speech synthesis, and community benchmarks.
Yorùbá Natural Language Processing (NLP) and speech technology cover computational methods, datasets, and machine learning architectures wey dem design to process, translate, transcribe, and synthesize Yorùbá language. As Yorùbá na tonal language wey dem dey write with modified Latin orthography wey get underdots and diacritical tone marks, computational systems dey face special structural obstacles wey normal text processing pipelines no fit handle. For inside di past twenty years, di field don grow from early rule-based and neural acoustic modeling for university laboratories reach large-scale participatory benchmarks, multi-domain translation corpora, and commercial voice systems .
Dis document explain di state of Yorùbá computational linguistics across all major functional areas. E dey check di technical connection between orthographic marks and semantic fidelity, review established datasets and algorithmic benchmarks, assess how e shift from closed pipelines enter open participatory frameworks, and document di structural limitations wey still dey challenge researchers.
Di Orthographic and Phonological Landscape for Computation
Standard Yorùbá orthography depend on diacritics to mark both segmental vowel quality and suprasegmental pitch contours. Vowels wey get underdots or vertical subdots, specifically ẹ ([ɛ]), ọ ([ɔ]), and di consonant ṣ ([ʃ]), dey represent different phonemes from di ones wey no get mark: e ([e]), o ([o]), and s ([s]). Suprasegmental tone get three level pitch registers: high (wey dem mark with acute accent, like á), mid (wey no get mark, like a), and low (wey dem mark with grave accent, like à). Orthography dey show nasalization through n wey dey follow at di back (e.g., an, ẹn, in, ọn, un).
Owo (mid-mid): Hand / Broom
Owó (mid-high): Money
Ọwọ̀ (subdot low - subdot low): Respect / Honor
Ọ̀wọ́ (subdot low - subdot high): Group / Broom cluster
For computational settings, when dem leave out dis marks, e dey create serious lexical ambiguity. Single token wey dem write without diacritics like owo fit mean pass four different words wey no relate in meaning at all . As standard web scraping dey collect Yorùbá text wey dem type on top keyboards wey no get native character support, majority of text wey dem mine from web no get diacritics at all, and dis one dey create big bottleneck for downstream language processing .
Original: Ọwọ́ tọ̀ ọ́.Literal gloss: Ọwọ́ (hand / group) tọ̀ (reach / visit) ọ́ (him / her / it).
Idiomatic English: A hand reached it / It was touched. (Rendered by corpus annotator).
Notes on the translation: Without tone marks, "Owo to o" could alternatively signify "Money is enough for him" (Owó tó o) or "Respect visits him" (Ọwọ̀ tọ̀ ọ́). The tone marks alone carry the grammatical and semantic distinctions.
Subword tokenization algorithms wey dem dey commonly use for large multilingual language models, like Byte-Pair Encoding (BPE) and WordPiece, come make dis wahala worse. Standard tokenizers dey often split accented vowels into base characters and separate Unicode combining diacritics, and dis dey scatter Yorùbá words into multiple sub-tokens. Dis fragmentation dey make sequence length long pass normal, dey reduce quality of positional representations for transformer models, and dey increase error rate during downstream generation .
Automatic Diacritic Restoration (ADR)
Because machine translation engines or text-to-speech synthesizers no fit accurately process text wey no get diacritics, Automatic Diacritic Restoration (ADR) dey serve as primary preprocessing step inside Yorùbá text pipelines .
+-------------------------------------------------------------+
| Automatic Diacritic Restoration |
| |
| Raw Input: owo -> [Sequence Processor] |
| Resolved Output: owó / ọwọ̀ / ọ̀wọ́ (Context Dependent) |
+-------------------------------------------------------------+
For past, diacritic restoration depend on dictionary lookups, rule-based morpho-syntactic heuristics, and word-level or character-level n-gram language models. Dis methods fail when dem encounter out-of-vocabulary terms, inflectional variations, or complex syntactic contexts.
One major change inside methodology happen when Iroro Orife show say dem fit frame Yorùbá ADR as character-level sequence-to-sequence translation task using recurrent neural networks with attention mechanisms . Di attentive sequence-to-sequence model dey take sentence wey no get diacritics as input characters and generate full diacritized orthographic sequences, and e successfully capture long-distance contextual dependencies . Orife show say dis neural architecture perform far better pass traditional dictionary lookup and n-gram baselines across standard evaluation corpora .
Open Debates and Domain Fragility
Scholars for Yorùbá NLP still dey divided on top di best way to frame ADR:
- Sequence Generation (Seq2Seq): To treat ADR as generation task dey allow di model change string lengths, insert missing subdots, and handle character reordering. But, e dey bring risk of hallucinating extra characters or changing base consonants .
- Sequence Tagging (Token/Character Classification): To frame ADR as sequence tagging problem dey restrict di model make e predict diacritic label for each static character position, wey dey prevent hallucinations. People wey support am argue say dis dey guarantee character alignment, while critics note say e dey struggle with non-standard transcriptions where vowels don elide or expand .
One limitation wey dem document for ADR literature na domain transfer degradation. Models wey dem train on clean religious texts wey get diacritics (like Bible translations) dey drop well well for accuracy when dem evaluate dem on top modern social media text or informal news reports wey get code-switching and slang . As well as dat, di published literature never talk anyting about how to standardize diacritic variants across regional Yorùbá dialects, bikos all existing ADR systems dey run only on Standard Written Yorùbá .
Neural Machine Translation (NMT) and di Diacritic Dependency
Historically, machine translation for Yorùbá suffer serious lack of parallel data and heavy domain bias. Early multilingual datasets depend too much on religious corpora like JW300, wey make models dey produce archaic, Bible-style translations when dem apply dem to general domain text .
+-------------------------------------------------------------+
| The Yorùbá Translation Dilemma (MENYO-20k) |
| |
| Undiacritized Input ==> Significant BLEU Drop |
| Diacritized Input ==> High Semantic Fidelity |
| Fine-Tuned Multilingual (M2M-100) ==> +8 to +9 BLEU Gain |
+-------------------------------------------------------------+
Di Participatory Paradigm (Masakhane)
To address di limitations of centralized scraping, di Masakhane research collective bring come participatory research framework for African NLP . Wilhelmina Nekoto and im colleagues lead dis framework, wey bring native speakers, linguists, and computer scientists togeda to curate open datasets, validate outputs, and build machine translation benchmarks for pass 30 African languages . Di study show say standard massive multilingual models dey perform very poorly on African languages unless local researchers and native speakers dey directly involved for dataset curation and output validation .
Di MENYO-20k Benchmark
As dem build on top dis framework, David Ifeoluwa Adelani and im collaborators develop MENYO-20k, wey be curated multi-domain parallel corpus wey contain 20,100 sentence pairs for Yorùbá and English . Di benchmark cover news, digital communication, literature, and general web domains.
Di MENYO-20k experiments bring out two main findings:
- Di Diacritic Impact: To remove diacritics from Yorùbá text dey cause heavy drop in translation quality across all automated metrics. When models receive input wey get complete diacritics, translation fidelity dey improve well well bikos tonal and phonemic ambiguities don komot from di source .
- Domain Adaptation: To fine-tune general pre-trained multilingual sequence-to-sequence models (like Facebook M2M-100) on top clean, multi-domain in-domain data pass commercial baselines, wey bring improvement of pass +8 to +9 BLEU points compared to baseline off-the-shelf multilingual models .
Metric Limitations
One well-known point of contestation for literature concern evaluation metrics. Standard string-overlap metrics like BLEU and chrF dey calculate token or character overlap without to give weight to tone or subdot errors according to how dem affect meaning . Just one missed underdot or wrong acute accent fit change whole meaning of sentence, but string-overlap metrics dey treat di error like small character discrepancy. Di research community never agree on one machine translation metric wey dey specific to Yorùbá wey everybody dey use, wey go penalize tone errors according to di damage wey dem cause to meaning .
Named Entity Recognition and Representation Learning
Named Entity Recognition (NER) for Yorùbá dey present serious structural challenges. Proper names for Yorùbá na compound phrases most times wey get clear lexical meanings wey carry tone patterns, verbs, and nouns (e.g., Ayọ̀délé, wey mean "Joy has arrived home"). Capitalization for digital text no consistent at all, and dem dey frequently omit tonal markers, wey dey make NER systems dey confuse personal names with common nouns or verbs .
Original: Ayọ̀délé lọ sí Èkó.Literal gloss: Ayọ̀délé (Joy-arrives-home) lọ (went) sí (to) Èkó (Lagos).
Idiomatic English: Ayodele went to Lagos. (Rendered by corpus annotator).
Notes on the translation: Without capitalization or tonal marks, "ayodele" can be read as a descriptive verb phrase rather than a proper personal name, confounding standard entity extractors.
MasakhaNER and Language-Adaptive Pre-training
To provide solid benchmarks, di Masakhane community create MasakhaNER, di first human-annotated NER dataset wey cover African languages, wey dem later expand into MasakhaNER 2.0 . Speaker-linguists annotate di dataset natively to capture four named entity categories: Personal Names (PER), Locations (LOC), Organizations (ORG), and Dates (DATE) .
Di accompanying research, wey David Ifeoluwa Adelani and im co-researchers write, show key representation learning dynamics:
- Inadequacy of High-Resource Transfer: Standard cross-lingual transfer from high-resource European languages wey dey use generic multilingual models (like mBERT) dey give poor generalization on Yorùbá text .
- Typological Transfer: To transfer representations between African languages wey relate typologically and geographically dey perform significantly better pass transfer from English or French .
- AfroXLMR: Continual language-adaptive pre-training on top African corpora (wey produce specialized models like AfroXLMR) achieve state-of-the-art F1 performance across Yorùbá NER tasks .
Annotation Strategies
Scholars dey argue weda synthetic or distantly supervised datasets fit replace human annotators wey scarce. Even though distant supervision dey make dataset scaling cheap, Adelani et al. show say distant supervision wey no get tone-aware filtering dey bring heavy noise, dey wrongly classify common words wey depend on tone as proper nouns . General agreement still be say high-quality gold standards wey humans annotate dey very necessary for reliable evaluation .
Sentiment Analysis and Social Corpora
To process sentiment inside Yorùbá digital text need make person handle informal language, code-switching between Yorùbá and Nigerian English (or Nigerian Pidgin), plus spelling wey no follow standard orthography.
To set computational baselines, the community build AfriSenti, one large-scale sentiment analysis benchmark wey get over 110,000 annotated tweets across 14 African languages, including Yorùbá, Hausa, and Igbo . Shamsuddeen Hassan Muhammad and him colleagues lead the work, and the dataset serve as the foundation for the SemEval-2023 Task 12 shared task .
+-------------------------------------------------------------+
| AfriSenti Benchmark Architecture |
| |
| 110,000+ Annotated Tweets across 14 African Languages |
| Includes Yorùbá, Hausa, Igbo |
| Gold standard for multi-class sentiment classification |
+-------------------------------------------------------------+
AfriSenti show how code-switching and social media expressions wey no get tone marks dey disturb standard monolingual classifiers . Evaluating models on top this corpus prove say language-specific fine-tuning on informal corpora dey very necessary for practical social listening and opinion mining systems .
Speech Processing: Foundations, Acoustic Modeling, and Corpora
Speech technologies for Yorùbá must model three distinct linguistic tone registers (high, mid, low) alongside standard phonemic vowel and consonant inventories. Pitch contours dey carry lexical and grammatical meaning; Automatic Speech Recognition (ASR) system wey transcribe segments without pitch no go fit know the correct lexical meaning, while Text-to-Speech (TTS) synthesizer wey produce flat intonation go sound unnatural and e no go dey clear to understand .
+-------------------------------------------------------------+
| Yorùbá Speech Architecture |
| |
| [F0 Fundamental Frequency] ---> [Pitch Contour Tracking] |
| [Acoustic Phonemes] ---> [Neural ASR / TTS Engine] |
| Outputs: Full Lexical Tone & Segmental Resolution |
+-------------------------------------------------------------+
Early Foundations: Tone Recognition and Syllable Segmentation
Early computational work on Yorùbá speech modeling start for Obafemi Awolowo University (OAU) under Ọdétúnjí Àjàdí Ọdẹ́jọbí . Ọdẹ́jọbí's research focus on pitch-contour fundamental frequencies ($F_0$) wey artificial neural networks process to achieve computational tone recognition and syllable segmentation .
Ọdẹ́jọbí show say to separate pitch contours and track fundamental frequency transitions allow neural architectures to classify the three tone registers of Yorùbá correctly from spoken speech samples, wey come lay the acoustic foundation for modern neural voice processing .
The ÌròyìnSpeech Corpus
For pass ten years after Ọdẹ́jọbí's early experiments, progress inside Yorùbá speech technology slow down because open speech corpora wey get diacritics no dey. In 2024, one collaborative team wey include Tolúlọpẹ́ Ógúnrẹ̀mí, Kọ́lá Túbọ̀sún, Anuoluwapo Aremu, Iroro Orife, and David Ifeoluwa Adelani release ÌròyìnSpeech .
+-------------------------------------------------------------+
| ÌròyìnSpeech Characteristics |
| |
| Total Duration: ~42 hours of diacritized speech |
| Speakers: 80 volunteer native speakers |
| Tasks Supported: ASR, Single-Speaker TTS, Multi-Speaker |
| Baseline ASR WER: 23.8% |
| TTS Data Req.: High-fidelity output at ~5 hours |
+-------------------------------------------------------------+
The ÌròyìnSpeech dataset get around 42 hours of high-quality, contemporary news speech wey get complete diacritics, recorded by 80 volunteer speakers . The authors set baseline benchmarks for both ASR and TTS:
- ASR Benchmark: Fine-tuning modern self-supervised speech models (like wav2vec 2.0 and Whisper) on top ÌròyìnSpeech set baseline Word Error Rate (WER) of 23.8% on standard test splits .
- TTS Benchmark: The study show say person fit train high-fidelity single-speaker TTS synthesizers with only 5 hours of clean audio wey get complete diacritics, and e go produce natural pitch transitions across tone registers .
Dialectal Divergence and the Standard Written Constraint
One major structural blind spot inside Yorùbá NLP na the near-exclusive focus on Standard Written Yorùbá (SWY), wey dem historically base on Oyo and Lagos dialects (North-West Yorùbá) .
Regional varieties across the Yorùbá-speaking region, like Ìjẹ̀bú, Èkìtì, Òndó, and Ọ̀wọ̀, show clear phonological, lexical, and tonal differences. For example, specific vowel shifts, elisions, and different lexical choices dey common inside everyday speech across regional communities.
+-------------------------------------------------------------+
| Dialectal Performance Disparity |
| |
| Standard NW Yorùbá Training Data (Standard Models) |
| | |
| +---> High Accuracy on Standard News / Audio |
| | |
| +---> Severe Zero-Shot Performance Drop on |
| Regional Dialects (Ìjẹ̀bú, Èkìtì, Òndó) |
+-------------------------------------------------------------+
To measure dis performance gap, Tolúlọpẹ́ Ógúnrẹ̀mí and dia colleagues develop YorùLect (wey dem publish for di Voices Unheard initiative) . Dia research evaluate mainstream pre-trained speech models, like Meta's Massively Multilingual Speech (MMS) and OpenAI's Whisper, across regional Yorùbá dialects.
Di results show say speech models wey dem train only on Standard North-West Yorùbá dey suffer heavy zero-shot performance drop when dem face regional dialect audio . Di authors argue say to build speech technologies wey carry everybody follow go require deliberate collection of multi-dialect corpus, instead of assuming say standard models go fit work across dialect boundaries .
Industry Ecosystem and Open Infrastructure
Along with academic research, industrial and open-source infrastructure for Yorùbá language technology don come up, through collaborations between African startups, government bodies, and international research organizations.
+------------------------------------------------------------------+
| Yorùbá Language Tech Ecosystem |
| |
| Commercial & Startups: Spitch (STT/TTS via Cencori) |
| Open Models & Governance: Awarri & FMCIDE (Yoruba-ASR/N-ATLAS) |
| Academic / Grassroots: Masakhane, OAU Labs, ÌròyìnSpeech |
+------------------------------------------------------------------+
Commercial Implementations: Spitch
Spitch na one African voice-AI technology company wey dey provide production-grade speech infrastructure . Dem fit deploy am through application programming interfaces like di Cencori AI Gateway, and Spitch dey offer:
- Speech-to-Text (STT) for Yorùbá, Hausa, Igbo, and Amharic.
- Text-to-Speech (TTS) synthesis wey fit produce native tonal intonation.
- Automatic Diacritics Restoration wey dem integrate into voice-note processing and Interactive Voice Response (IVR) enterprise applications .
Proprietary training sets and exact out-of-domain error rates for closed commercial engines like Spitch never dey documented for peer-reviewed literature, and dis leave empirical evaluation gap between commercial claims and public benchmarks.
Open-Source Government and Enterprise Partnerships: Awarri
Awarri, an AI research organization wey Silas Adekunle start, don advance open-source foundational models for Nigerian languages in partnership with public institutions .
In collaboration with Nigeria's Federal Ministry of Communications, Innovation, and Digital Economy (FMCIDE) and di National Centre for Artificial Intelligence and Robotics (NCAIR), Awarri develop and release:
Yoruba-ASR: One open fine-tuned Whisper-Small model wey dem optimize for Nigerian speech contexts .N-ATLAS: One open-source, voice-first multilingual model wey cover Yorùbá, Hausa, Igbo, and Nigerian Pidgin, wey dem release to provide foundational linguistic infrastructure for civic and commercial developers .
Gaps, Methodological Contradictions, and Open Problems
Even with rapid progress since 2020, Yorùbá NLP still dey face theoretical and methodological trade-offs wey dem never resolve:
- Strict Diacritic Enforcement vs. Robustness to Unmarked Text: Researchers no agree whether text generation pipelines must strictly mandate automatic diacritic restoration as a compulsory pre-processing step or make dem develop downstream encoders wey naturally get power to handle missing diacritics. To force ADR dey bring pipeline delay and dey multiply upstream restoration errors, but to bypass ADR dey leave serious semantic confusion inside di latent space .
- Subword Tokenizer Inefficiencies: Standard subword algorithms still dey break Yorùbá vowels wey get diacritic mark into scattered characters and Unicode combiners. Dis dey increase sequence length and dey blow up memory use during training. Major model registries never standardize any dedicated, universally adopted Yorùbá subword tokenizer wey naturally represent combined vowel-diacritic units .
- Absence of Dialect Orthographies: Outside Standard Written Yorùbá, non-standard regional dialects no get standardized spelling rules. As a result, computational work on regional dialects mostly remain limited to speech audio, while text-based NLP remain strictly monolingual inside standard North-West Yorùbá .
- Scarcity of Peer-Reviewed Commercial Benchmarks: Even though commercial API platforms dey claim high operational accuracy on interactive voice applications, peer-reviewed comparative studies never systematically verify dia training data distribution, how dem dey handle code-switching, and dia resistance to background noise .
Summary of Key Benchmarks and Corpora
| Resource | Main Domain | Main Task | Reference |
|---|---|---|---|
| MENYO-20k | Multi-domain text (News, Web, Literature) | Neural Machine Translation | Adelani et al. (2021) |
| MasakhaNER / 2.0 | News and Local Article dem | Named Entity Recognition | Adelani et al. (2021) |
| AfriSenti | Social Media (Twitter) | Sentiment Classification | Muhammad et al. (2023) |
| ÌròyìnSpeech | Broadcast News Speech (~42 hrs) | ASR, TTS Synthesis | Ógúnrẹ̀mí et al. (2024) |
| YorùLect (Voices Unheard) | Regional Dialect Speech | Multi-dialect ASR Evaluation | Ógúnrẹ̀mí et al. (2024) |
| Yoruba-ASR / N-ATLAS | Multi-domain Voice / Language Models | Speech Recognition and Multimodal NLP | Awarri & FMCIDE (2024) |