Language Corpus: Section Guide
Guide to the five files in the language corpus section, covering text and speech datasets, natural language processing, lexicography, descriptive grammars, and language pedagogy and policy.
The language corpus section provides an exhaustive reference for the structural, computational, lexicographical, and educational infrastructure of the Yorùbá language. It documents the historical tools used to analyze Yorùbá speech and writing alongside contemporary digital resources, dataset benchmarks, grammatical frameworks, and institutional language policies. Five files examine how Yorùbá is recorded in textual and auditory archives, processed by computational systems, categorized in dictionaries, analyzed in formal linguistics, and taught in institutional school systems.
This section serves as the technical companion to 02-language, bridging the gap between descriptive sociolinguistics and the empirical materials necessary for computational modeling, academic analysis, and language planning [S1, S2]. Two files carry foundational weight across several sub-disciplines: Yorùbá Dictionaries and Lexicography (03), which uncovers how nineteenth-century missionary encounters introduced theological bias into lexical definitions, and Teaching and Learning Yorùbá (05), which charts the empirical success and political struggles of mother-tongue education from the Ifẹ̀ Six-Year Primary Project through the federal policy shifts of late 2025 [S3, S4].
The files
| # | File | One-line summary |
|---|---|---|
| 01 | Yorùbá Text and Speech Corpora | Parallel, monolingual, and speech datasets available for computational and linguistic research, from JW300 and MENYO-20k to NaijaVoices and Common Voice. |
| 02 | Yorùbá Natural Language Processing and Speech Technology | Computational methods, automatic diacritic restoration, tokenization challenges, speech recognition, and participatory machine translation architectures. |
| 03 | Yorùbá Dictionaries and Lexicography | Historical evolution of lexicography from Crowther's 1843 wordlists to Abraham's encyclopedic 1958 dictionary and modern electronic lexical databases. |
| 04 | Yorùbá Grammars and Linguistic Description | Descriptive frameworks from missionary sketches to structuralist and generative models, focusing on serial verbs, splitting verbs, vowel harmony, and tone. |
| 05 | Teaching and Learning Yorùbá | Language policy history in Nigeria, the Ifẹ̀ Six-Year Primary Project, classroom curricula, diaspora pedagogy, and the November 2025 mother-tongue policy reversal. |
Guided reading and thematic overview
01. Yorùbá Text and Speech Corpora
Yorùbá Text and Speech Corpora audits the open-access and proprietary text corpora, parallel datasets, and speech archives available for research [S1, S5]. It documents the dominance of religious domain parallel text, such as JW300 and multilingual Bible corpora, alongside multi-domain datasets such as MENYO-20k, which introduced standardized benchmarks for news, proverbs, books, and digital localization [S1,]. In the speech domain, the file details collections ranging from Mozilla Common Voice and BibleTTS to recent community-driven initiatives like ÌròyìnSpeech and NaijaVoices [S5,]. What remains undocumented and contested in corpus linguistics is the representation of non-standard varieties: available text and speech corpora overwhelmingly privilege Standard Written Yorùbá and Ọ̀yọ́ phonology, leaving Eastern Yorùbá, North-West peripheral dialects, and informal urban slang severely underrepresented in public computational repositories [S1, S5].
02. Yorùbá Natural Language Processing and Speech Technology
Yorùbá Natural Language Processing and Speech Technology examines the computational techniques and engineering challenges of applying machine learning models to a tonal, morphologically agglutinative language [S2, S6]. It provides a detailed technical breakdown of Automatic Diacritic Restoration (ADR), contrasting early syllable-based and n-gram methods developed by Nigerian computer scientists with contemporary sequence-to-sequence neural architectures [S6,, S7]. The file analyzes tokenization bottlenecks, showing how subword tokenizers designed for English fragment diacritized Yorùbá vowels into disconnected bytes, degrading downstream performance in Automatic Speech Recognition (ASR) and Neural Machine Translation (NMT) [S2,]. What remains unresolved is whether massively multilingual foundation models genuinely learn Yorùbá semantics or rely on superficial lexical matching that deteriorates whenever complex tonal puns, honorific constructions, or idiomatic proverbs are introduced [S2, S6].
03. Yorùbá Dictionaries and Lexicography
Yorùbá Dictionaries and Lexicography tracks the historical lineage of dictionary compilation across nearly two centuries [S3, S8]. Starting with Samuel Àjàyí Crowther's 1843 Vocabulary of the Yoruba Language and its 1852 revision, the file progresses through Thomas Jefferson Bowen's 1858 compilation, the 1913 Church Missionary Society (CMS) dictionary, and Roy Clive Abraham's landmark 1958 Dictionary of Modern Yoruba [S3, S8, S9]. It also surveys contemporary electronic lexicography, including Yíwọlá Awóyalé's multi-decade lexical database project . The file identifies how early missionary lexicographers distorted cultural and cosmological terms by shoehorning indigenous religious concepts into Victorian Christian categories, most notably translating Èṣù as "Satan" or "devil" and ẹbọ as "heathen sacrifice" [S3, S8]. What remains contested in contemporary lexicography is the proper lemma structure for phrasal compounds, ideophones, and splitting verbs, as well as the standard lexicographical representation of dialectal variants [S8, S10].
04. Yorùbá Grammars and Linguistic Description
Yorùbá Grammars and Linguistic Description details the development of descriptive and theoretical linguistic accounts of Yorùbá syntax, phonology, and morphology [S11, S12]. It traces the progression from nineteenth-century Latinate grammars to Ida Ward's 1952 phonetic studies, Ayọ̀ Bámgbóṣé's structuralist 1966 A Grammar of Yoruba, and Ọladele Awóbùlúyì's 1978 Essentials of Yoruba Grammar [S11, S12, S13]. Major theoretical debates are explored in detail, including the formal syntactic analysis of Serial Verb Constructions (SVCs) versus splitting verbs, the phonological rules of Advanced Tongue Root (ATR) vowel harmony across mid-vowels (/e, o/ versus /ẹ, ọ/), and the status of subject clitics [S12, S14]. What remains theoretically contested is whether splitting verbs constitute distinct morphological units or syntactic sequences of independent verbs operating under strict lexicalization constraints [S11, S12].
05. Teaching and Learning Yorùbá
Teaching and Learning Yorùbá examines language planning, institutional pedagogy, and sociolinguistic policy in Nigeria and the global diaspora [S4, S15]. It documents the history of mother-tongue education research, centering on Professor Aliu Babátúndé Fáfunwa's Ifẹ̀ Six-Year Primary Project (1970–1978), which empirically demonstrated that primary students taught all subjects in Yorùbá achieved greater cognitive development and higher academic performance in secondary school English than pupils taught in English immersion [S4, S15]. The file outlines the provisions of Nigeria's 1977 National Policy on Education and the 2022 National Language Policy, and analyzes the policy reversal announced in November 2025 that discontinued the mother-tongue primary mandate [S15, S16]. What remains contested is the pedagogical tension between psycholinguistic evidence favoring mother-tongue concept acquisition and the socio-economic prestige of English that drives parental resistance in urban centers [S4, S16].
Suggested reading orders
For computational linguists and NLP researchers: Start with Yorùbá Text and Speech Corpora to understand the distribution and limits of available data, continue to Yorùbá Natural Language Processing and Speech Technology for modeling architectures and tokenization challenges, and consult Yorùbá Grammars and Linguistic Description to master the grammatical constraints governing tone and verb serialization.
For theoretical linguists and lexicographers: Begin with Yorùbá Grammars and Linguistic Description for core structural arguments, move to Yorùbá Dictionaries and Lexicography to evaluate lexical classification and historical missionary bias, and reference Yorùbá Text and Speech Corpora for empirical verification across text genres.
For educators, historians, and policy analysts: Read Teaching and Learning Yorùbá first to grasp the political and pedagogical history of language planning in Nigeria, then turn to Yorùbá Dictionaries and Lexicography to understand the colonial codification of the language, and conclude with Yorùbá Natural Language Processing and Speech Technology to examine digital language preservation.
What this section does not cover
This section focuses strictly on linguistic analysis, text and speech corpora, lexicographical reference, computational tooling, and educational policy.
The formal oral literature of the Yorùbá tradition, such as the 256 Odù Ifá divination poetry, oríkì praise poetry, orature, and proverbs (òwe), belongs to 02-language, 05-ifa, and 06-orisa. The metaphysical principles of orí (destiny/inner head), àṣẹ (generative spiritual force), and ìwà (character) are treated in 04-cosmology. Sociolinguistic discussions regarding the wider Niger-Congo language family and historical dialect migration belong to 01-the-yoruba-language.
All claims regarding historical dates, grammar publications, corpus distributions, and official policy shifts are grounded directly in the scholarly literature documented below.