Yorùbá Natural Language Processing and Speech Technology
A comprehensive scholarly survey of Yorùbá computational linguistics, covering automatic diacritic restoration, machine translation, speech recognition, speech synthesis, and community benchmarks.
Yorùbá Natural Language Processing (NLP) and speech technology encompass the computational methods, datasets, and machine learning architectures developed to process, translate, transcribe, and synthesize the Yorùbá language. Because Yorùbá is a tonal language written in a modified Latin orthography with underdots and diacritical tone marks, computational systems face unique structural hurdles that standard text processing pipelines do not accommodate. Over the past two decades, the field has evolved from early rule-based and neural acoustic modeling in university laboratories into large-scale participatory benchmarks, multi-domain translation corpora, and commercial voice systems .
This document details the state of Yorùbá computational linguistics across all major functional areas. It examines the technical dependencies between orthographic marks and semantic fidelity, reviews established datasets and algorithmic benchmarks, assesses the shift from closed pipelines to open participatory frameworks, and documents the structural limitations that continue to challenge researchers.
The Orthographic and Phonological Landscape in Computation
Standard Yorùbá orthography relies on diacritics to mark both segmental vowel quality and suprasegmental pitch contours. Vowels with underdots or vertical subdots, specifically ẹ ([ɛ]), ọ ([ɔ]), and the consonant ṣ ([ʃ]), represent distinct phonemes from their unmarked counterparts e ([e]), o ([o]), and s ([s]). Suprasegmental tone comprises three level pitch registers: high (marked with an acute accent, as in á), mid (unmarked, as in a), and low (marked with a grave accent, as in à). Nasalization is indicated orthographically by a trailing n (e.g., an, ẹn, in, ọn, un).
Owo (mid-mid): Hand / Broom
Owó (mid-high): Money
Ọwọ̀ (subdot low - subdot low): Respect / Honor
Ọ̀wọ́ (subdot low - subdot high): Group / Broom cluster
In computational settings, the omission of these marks creates massive lexical ambiguity. A single token written without diacritics as owo can map to more than four distinct lexical items with completely unrelated semantics . Because standard web scraping gathers Yorùbá text written on keyboards lacking native character support, web-mined text is overwhelmingly undiacritized, creating a severe bottleneck for downstream language processing .
Original:
Ọwọ́ tọ̀ ọ́.
Literal gloss:
Ọwọ́ (hand / group) tọ̀ (reach / visit) ọ́ (him / her / it).
Idiomatic English:
A hand reached it / It was touched. (Rendered by corpus annotator).
Notes on the translation:
Without tone marks, "Owo to o" could alternatively signify "Money is enough for him" (Owó tó o) or "Respect visits him" (Ọwọ̀ tọ̀ ọ́). The tone marks alone carry the grammatical and semantic distinctions.
Subword tokenization algorithms commonly used in large multilingual language models, such as Byte-Pair Encoding (BPE) and WordPiece, compound this challenge. Standard tokenizers frequently split accented vowels into base characters and separate Unicode combining diacritics, fragmenting Yorùbá words into multiple sub-tokens. This fragmentation inflates sequence length, degrades positional representations in transformer models, and increases the error rate during downstream generation .
Automatic Diacritic Restoration (ADR)
Because un-diacritized text cannot be reliably parsed by machine translation engines or text-to-speech synthesizers, Automatic Diacritic Restoration (ADR) serves as a primary preprocessing step in Yorùbá text pipelines .
+-------------------------------------------------------------+
| Automatic Diacritic Restoration |
| |
| Raw Input: owo -> [Sequence Processor] |
| Resolved Output: owó / ọwọ̀ / ọ̀wọ́ (Context Dependent) |
+-------------------------------------------------------------+
Historically, diacritic restoration relied on dictionary lookups, rule-based morpho-syntactic heuristics, and word-level or character-level n-gram language models. These approaches collapsed when encountering out-of-vocabulary terms, inflectional variations, or complex syntactic contexts.
A major methodological shift occurred when Iroro Orife demonstrated that Yorùbá ADR could be framed as a character-level sequence-to-sequence translation task using recurrent neural networks with attention mechanisms . The attentive sequence-to-sequence model takes an un-diacritized sentence as input characters and generates fully diacritized orthographic sequences, successfully capturing long-distance contextual dependencies . Orife showed that this neural architecture significantly outperformed traditional dictionary lookup and n-gram baselines across standard evaluation corpora .
Open Debates and Domain Fragility
Scholars in Yorùbá NLP remain divided on the optimal framing of ADR:
- Sequence Generation (Seq2Seq): Treating ADR as a generation task allows the model to alter string lengths, insert missing subdots, and handle character reordering. However, it introduces the risk of hallucinating extra characters or altering base consonants .
- Sequence Tagging (Token/Character Classification): Framing ADR as a sequence tagging problem constrains the model to predict a diacritic label for each static character position, preventing hallucinations. Proponents argue this guarantees character alignment, while critics note it struggles with non-standard transcriptions where vowels have been elided or expanded .
A documented limitation in ADR literature is domain transfer degradation. Models trained on clean, diacritized religious texts (such as biblical translations) experience severe accuracy drops when evaluated on modern social media text or informal news reports containing code-switching and slang . Furthermore, the published literature remains silent on standardizing diacritic variants across regional Yorùbá dialects, as all existing ADR systems operate exclusively on Standard Written Yorùbá .
Neural Machine Translation (NMT) and the Diacritic Dependency
Machine translation for Yorùbá historically suffered from an acute shortage of parallel data and severe domain bias. Early multilingual datasets relied disproportionately on religious corpora like JW300, which led models to generate archaic, biblically skewed translations when applied to general domain text .
+-------------------------------------------------------------+
| The Yorùbá Translation Dilemma (MENYO-20k) |
| |
| Undiacritized Input ==> Significant BLEU Drop |
| Diacritized Input ==> High Semantic Fidelity |
| Fine-Tuned Multilingual (M2M-100) ==> +8 to +9 BLEU Gain |
+-------------------------------------------------------------+
The Participatory Paradigm (Masakhane)
To address the limitations of centralized scraping, the Masakhane research collective introduced a participatory research framework for African NLP . Spearheaded by Wilhelmina Nekoto and colleagues, this framework united native speakers, linguists, and computer scientists to curate open datasets, validate outputs, and build machine translation benchmarks for over 30 African languages . The study established that standard massive multilingual models severely underperform on African languages unless local researchers and native speakers are directly involved in dataset curation and output validation .
The MENYO-20k Benchmark
Building on this framework, David Ifeoluwa Adelani and collaborators developed MENYO-20k, a curated multi-domain parallel corpus comprising 20,100 sentence pairs for Yorùbá and English . The benchmark spans news, digital communication, literature, and general web domains.
The MENYO-20k experiments yielded two critical findings:
- The Diacritic Impact: Stripping diacritics from Yorùbá text causes a steep decline in translation quality across all automated metrics. When models receive fully diacritized input, translation fidelity improves substantially because tonal and phonemic ambiguities are eliminated at the source .
- Domain Adaptation: Fine-tuning general pre-trained multilingual sequence-to-sequence models (such as Facebook M2M-100) on clean, multi-domain in-domain data surpassed commercial baselines, yielding improvements of over +8 to +9 BLEU points compared to baseline off-the-shelf multilingual models .
Metric Limitations
A recognized point of contestation in the literature concerns evaluation metrics. Standard string-overlap metrics such as BLEU and chrF calculate token or character overlap without weighting tone or subdot errors according to their semantic severity . A single missed underdot or incorrect acute accent can invert sentence meaning, yet string-overlap metrics treat the error as a minor character discrepancy. The research community has not yet converged on a universally adopted Yorùbá-specific machine translation metric that penalizes tone errors proportionally to semantic damage .
Named Entity Recognition and Representation Learning
Named Entity Recognition (NER) in Yorùbá presents severe structural challenges. Proper names in Yorùbá are frequently compound phrases with explicit lexical meanings that incorporate tone patterns, verbs, and nouns (e.g., Ayọ̀délé, meaning "Joy has arrived home"). Capitalization in digital text is highly inconsistent, and tonal markers are frequently omitted, causing NER systems to confuse personal names with common nouns or verbs .
Original:
Ayọ̀délé lọ sí Èkó.
Literal gloss:
Ayọ̀délé (Joy-arrives-home) lọ (went) sí (to) Èkó (Lagos).
Idiomatic English:
Ayodele went to Lagos. (Rendered by corpus annotator).
Notes on the translation:
Without capitalization or tonal marks, "ayodele" can be read as a descriptive verb phrase rather than a proper personal name, confounding standard entity extractors.
MasakhaNER and Language-Adaptive Pre-training
To provide rigorous benchmarks, the Masakhane community created MasakhaNER, the first human-annotated NER dataset covering African languages, later expanded into MasakhaNER 2.0 . The dataset was annotated natively by speaker-linguists to capture four named entity categories: Personal Names (PER), Locations (LOC), Organizations (ORG), and Dates (DATE) .
The accompanying research, authored by David Ifeoluwa Adelani and co-researchers, demonstrated key representation learning dynamics:
- Inadequacy of High-Resource Transfer: Standard cross-lingual transfer from high-resource European languages using generic multilingual models (such as mBERT) yields poor generalization on Yorùbá text .
- Typological Transfer: Transferring representations between typologically and geographically related African languages significantly outperforms transfer from English or French .
- AfroXLMR: Continual language-adaptive pre-training on African corpora (yielding specialized models such as AfroXLMR) achieved state-of-the-art F1 performance across Yorùbá NER tasks .
Annotation Strategies
Scholars debate whether synthetic or distantly supervised datasets can substitute for scarce human annotators. While distant supervision enables low-cost dataset scaling, Adelani et al. showed that distant supervision without tone-aware filtering introduces severe noise, misclassifying tone-dependent common words as proper nouns . The consensus remains that high-quality, human-annotated gold standards are indispensable for reliable evaluation .
Sentiment Analysis and Social Corpora
Processing sentiment in Yorùbá digital text requires handling informal language, code-switching between Yorùbá and Nigerian English (or Nigerian Pidgin), and unstandardized orthography.
To establish computational baselines, the community built AfriSenti, a large-scale sentiment analysis benchmark comprising over 110,000 annotated tweets across 14 African languages, including Yorùbá, Hausa, and Igbo . Led by Shamsuddeen Hassan Muhammad and colleagues, the dataset served as the foundation for the SemEval-2023 Task 12 shared task .
+-------------------------------------------------------------+
| AfriSenti Benchmark Architecture |
| |
| 110,000+ Annotated Tweets across 14 African Languages |
| Includes Yorùbá, Hausa, Igbo |
| Gold standard for multi-class sentiment classification |
+-------------------------------------------------------------+
AfriSenti highlighted the extent to which code-switching and un-diacritized social media expressions challenge standard monolingual classifiers . Evaluating models on this corpus proved that language-specific fine-tuning on informal corpora is essential for practical social listening and opinion mining systems .
Speech Processing: Foundations, Acoustic Modeling, and Corpora
Speech technologies for Yorùbá must model three distinct linguistic tone registers (high, mid, low) alongside standard phonemic vowel and consonant inventories. Pitch contours carry lexical and grammatical meaning; an Automatic Speech Recognition (ASR) system that transcribes segments without pitch fails to resolve lexical identity, while a Text-to-Speech (TTS) synthesizer that produces flat intonation sounds unnatural and unintelligible .
+-------------------------------------------------------------+
| Yorùbá Speech Architecture |
| |
| [F0 Fundamental Frequency] ---> [Pitch Contour Tracking] |
| [Acoustic Phonemes] ---> [Neural ASR / TTS Engine] |
| Outputs: Full Lexical Tone & Segmental Resolution |
+-------------------------------------------------------------+
Early Foundations: Tone Recognition and Syllable Segmentation
Early computational work on Yorùbá speech modeling was pioneered at Obafemi Awolowo University (OAU) by Ọdétúnjí Àjàdí Ọdẹ́jọbí . Ọdẹ́jọbí's research focused on pitch-contour fundamental frequencies ($F_0$) processed through artificial neural networks to achieve computational tone recognition and syllable segmentation .
Ọdẹ́jọbí demonstrated that isolating pitch contours and tracking fundamental frequency transitions allowed neural architectures to classify Yorùbá's three tone registers accurately from spoken speech samples, laying the acoustic groundwork for modern neural voice processing .
The ÌròyìnSpeech Corpus
For over a decade following Ọdẹ́jọbí's early experiments, progress in Yorùbá speech technology was constrained by the absence of open, diacritized speech corpora. In 2024, a collaborative team comprising Tolúlọpẹ́ Ógúnrẹ̀mí, Kọ́lá Túbọ̀sún, Anuoluwapo Aremu, Iroro Orife, and David Ifeoluwa Adelani released ÌròyìnSpeech .
+-------------------------------------------------------------+
| ÌròyìnSpeech Characteristics |
| |
| Total Duration: ~42 hours of diacritized speech |
| Speakers: 80 volunteer native speakers |
| Tasks Supported: ASR, Single-Speaker TTS, Multi-Speaker |
| Baseline ASR WER: 23.8% |
| TTS Data Req.: High-fidelity output at ~5 hours |
+-------------------------------------------------------------+
The ÌròyìnSpeech dataset consists of approximately 42 hours of high-quality, fully diacritized contemporary news speech recorded by 80 volunteer speakers . The authors established baseline benchmarks for both ASR and TTS:
- ASR Benchmark: Fine-tuning modern self-supervised speech models (such as wav2vec 2.0 and Whisper) on ÌròyìnSpeech established a baseline Word Error Rate (WER) of 23.8% on standard test splits .
- TTS Benchmark: The study proved that high-fidelity single-speaker TTS synthesizers could be trained with as little as 5 hours of clean, fully diacritized audio, producing natural pitch transitions across tone registers .
Dialectal Divergence and the Standard Written Constraint
A major structural blind spot in Yorùbá NLP is the near-exclusive focus on Standard Written Yorùbá (SWY), which is based historically on the Oyo and Lagos dialects (North-West Yorùbá) .
Regional varieties across the Yorùbá-speaking region, such as Ìjẹ̀bú, Èkìtì, Òndó, and Ọ̀wọ̀, exhibit marked phonological, lexical, and tonal shifts. For example, specific vowel shifts, elisions, and alternative lexical items are common in everyday speech across regional communities.
+-------------------------------------------------------------+
| Dialectal Performance Disparity |
| |
| Standard NW Yorùbá Training Data (Standard Models) |
| | |
| +---> High Accuracy on Standard News / Audio |
| | |
| +---> Severe Zero-Shot Performance Drop on |
| Regional Dialects (Ìjẹ̀bú, Èkìtì, Òndó) |
+-------------------------------------------------------------+
To measure this performance gap, Tolúlọpẹ́ Ógúnrẹ̀mí and colleagues developed YorùLect (published in the Voices Unheard initiative) . Their research evaluated mainstream pre-trained speech models, such as Meta's Massively Multilingual Speech (MMS) and OpenAI's Whisper, across regional Yorùbá dialects.
The results demonstrated that speech models trained exclusively on Standard North-West Yorùbá suffer catastrophic zero-shot performance drops when exposed to regional dialect audio . The authors argued that building inclusive speech technologies requires deliberate multi-dialect corpus collection rather than assuming standard models will generalize across dialect boundaries .
Industry Ecosystem and Open Infrastructure
Alongside academic research, an industrial and open-source infrastructure for Yorùbá language technology has emerged, characterized by collaborations between African startups, government bodies, and international research organizations.
+------------------------------------------------------------------+
| Yorùbá Language Tech Ecosystem |
| |
| Commercial & Startups: Spitch (STT/TTS via Cencori) |
| Open Models & Governance: Awarri & FMCIDE (Yoruba-ASR/N-ATLAS) |
| Academic / Grassroots: Masakhane, OAU Labs, ÌròyìnSpeech |
+------------------------------------------------------------------+
Commercial Implementations: Spitch
Spitch is an African voice-AI technology company providing production-grade speech infrastructure . Deployable via application programming interfaces such as the Cencori AI Gateway, Spitch offers:
- Speech-to-Text (STT) for Yorùbá, Hausa, Igbo, and Amharic.
- Text-to-Speech (TTS) synthesis capable of generating native tonal intonation.
- Automatic Diacritics Restoration integrated into voice-note processing and Interactive Voice Response (IVR) enterprise applications .
Proprietary training sets and exact out-of-domain error rates for closed commercial engines such as Spitch remain undocumented in peer-reviewed literature, leaving an empirical evaluation gap between commercial claims and public benchmarks.
Open-Source Government and Enterprise Partnerships: Awarri
Awarri, an AI research organization founded by Silas Adekunle, has advanced open-source foundational models for Nigerian languages in partnership with public institutions .
In collaboration with Nigeria's Federal Ministry of Communications, Innovation, and Digital Economy (FMCIDE) and the National Centre for Artificial Intelligence and Robotics (NCAIR), Awarri developed and released:
Yoruba-ASR: An open fine-tuned Whisper-Small model optimized for Nigerian speech contexts .N-ATLAS: An open-source, voice-first multilingual model covering Yorùbá, Hausa, Igbo, and Nigerian Pidgin, released to provide foundational linguistic infrastructure for civic and commercial developers .
Gaps, Methodological Contradictions, and Open Problems
Despite rapid advancement since 2020, Yorùbá NLP remains constrained by unresolved theoretical and methodological trade-offs:
- Strict Diacritic Enforcement vs. Robustness to Unmarked Text: Researchers disagree on whether text generation pipelines must strictly mandate automatic diacritic restoration as a mandatory pre-processing step or develop downstream encoders that are inherently robust to missing diacritics. Enforcing ADR introduces pipeline latency and compounds upstream restoration errors, whereas bypassing ADR preserves severe semantic ambiguities in the latent space .
- Subword Tokenizer Inefficiencies: Standard subword algorithms continue to segment diacritically marked Yorùbá vowels into fragmented characters and Unicode combiners. This increases sequence lengths and inflates memory usage during training. A dedicated, universally adopted Yorùbá subword tokenizer that natively represents composite vowel-diacritic units has not been standardized across major model registries .
- Absence of Dialect Orthographies: Outside of Standard Written Yorùbá, non-standard regional dialects lack standardized orthographic conventions. Consequently, computational work on regional dialects is largely restricted to speech audio, while text-based NLP remains exclusively monolingual in standard North-West Yorùbá .
- Scarcity of Peer-Reviewed Commercial Benchmarks: While commercial API platforms claim high operational accuracy on interactive voice applications, their training data distribution, handling of code-switching, and resilience to background noise have not been systematically verified in peer-reviewed comparative studies .
Summary of Key Benchmarks and Corpora
| Resource | Primary Domain | Core Task | Reference |
|---|---|---|---|
| MENYO-20k | Multi-domain text (News, Web, Literature) | Neural Machine Translation | Adelani et al. (2021) |
| MasakhaNER / 2.0 | News and Local Articles | Named Entity Recognition | Adelani et al. (2021) |
| AfriSenti | Social Media (Twitter) | Sentiment Classification | Muhammad et al. (2023) |
| ÌròyìnSpeech | Broadcast News Speech (~42 hrs) | ASR, TTS Synthesis | Ógúnrẹ̀mí et al. (2024) |
| YorùLect (Voices Unheard) | Regional Dialect Speech | Multi-dialect ASR Evaluation | Ógúnrẹ̀mí et al. (2024) |
| Yoruba-ASR / N-ATLAS | Multi-domain Voice / Language Models | Speech Recognition and Multimodal NLP | Awarri & FMCIDE (2024) |
Fuentes
- [1]Iroro Orife, Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text, Proceedings of Interspeech 2018 (International Speech Communication Association, 2018), pp. 2848–2852.
- [2]Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Tajudeen Kolawole, et al., Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages, Findings of the Association for Computational Linguistics: EMNLP 2020 (Association for Computational Linguistics, 2020), pp. 2144–2160.
- [3]David Ifeoluwa Adelani, Dana Ruiter, Jesujoba O. Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Awokoya, and Cristina España-Bonet, The Effect of Domain and Diacritics in Yorùbá–English Neural Machine Translation (MENYO-20k), Proceedings of the 18th Biennial Machine Translation Summit: Research Track (Association for Machine Translation in the Americas, 2021), pp. 61–70.
- [4]David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, et al., MasakhaNER: Named Entity Recognition for African Languages, Transactions of the Association for Computational Linguistics (MIT Press / ACL, 2021), Vol. 9, pp. 1116–1131.
- [5]Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum, David Ifeoluwa Adelani, Seid Muhie Yimam, et al., AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Association for Computational Linguistics, 2023), pp. 15788–15802.
- [6]Ọdétúnjí Àjàdí Ọdẹ́jọbí, Recognition of Tones in Yorùbá Speech: Experiments With Artificial Neural Networks, in Speech and Language Technologies, ed. Ivo Ipsic (IntechOpen, 2008), pp. 287–304.
- [7]Tolúlọpẹ́ Ógúnrẹ̀mí, Kọ́lá Túbọ̀sún, Anuoluwapo Aremu, Iroro Orife, and David Ifeoluwa Adelani, ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (ELRA and ICCL, 2024), pp. 4668–4679.
- [8]Tolúlọpẹ́ Ógúnrẹ̀mí, et al., Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024) (Association for Computational Linguistics, 2024), pp. 18231–18245.
- [9]Daniel Oreofe, et al., Spitch is Live on Cencori: Native African Voice Infrastructure, Cencori Developer Publications / Connecting Africa (2026).
- [10]Awarri Technologies and Federal Ministry of Communications, Innovation, and Digital Economy (FMCIDE), Yoruba-ASR-v1.0: Automatic Speech Recognition for Yoruba Language, Hugging Face Model Hub (2024).