Language Corpus: Section Guide
Guide to the five files in the language corpus section, covering text and speech datasets, natural language processing, lexicography, descriptive grammars, and language pedagogy and policy.
Guide to the five files in the language corpus section, covering text and speech datasets, natural language processing, lexicography, descriptive grammars, and language pedagogy and policy.
The Yorùbá wey dem quote, the proverbs, oríkì, ẹsẹ Ifá, word list headwords, Odù names and citations dey exactly as the corpus record dem, for every language.
Dis language corpus section dey provide complete reference work for di structural, computational, lexicographical, and educational foundation of Yorùbá language. E document di old tools wey dem use take analyze Yorùbá speech and writing, alongside modern digital resources, dataset benchmarks, grammatical frameworks, and institutional language policies. Five files examine how dem dey record Yorùbá inside text and voice archives, how computer systems dey process am, how dictionary dey arrange am, how formal linguistics dey analyze am, and how dem dey teach am inside school systems.
Dis section dey serve as technical companion to 02-language, wey bridge di gap between descriptive sociolinguistics and di empirical materials wey dey necessary for computational modeling, academic analysis, and language planning [S1, S2]. Two files carry heavy foundation across several sub-disciplines: Yorùbá Dictionaries and Lexicography (03), wey expose how 19th-century missionary encounter bring religious bias enter lexical definitions, and Teaching and Learning Yorùbá (05), wey trace di empirical success and political struggles of mother-tongue education from di Ifẹ̀ Six-Year Primary Project reach di federal policy change dem do for late 2025 [S3, S4].
| # | File | One-line summary |
|---|---|---|
| 01 | Yorùbá Text and Speech Corpora | Parallel, monolingual, and speech datasets wey dey available for computational and linguistic research, from JW300 and MENYO-20k to NaijaVoices and Common Voice. |
| 02 | Yorùbá Natural Language Processing and Speech Technology | Computational methods, automatic diacritic restoration, tokenization challenges, speech recognition, and participatory machine translation architectures. |
| 03 | Yorùbá Dictionaries and Lexicography | How lexicography develop over time from Crowther 1843 wordlists to Abraham encyclopedic 1958 dictionary and modern electronic lexical databases. |
| 04 | Yorùbá Grammars and Linguistic Description | Descriptive frameworks from missionary sketches to structuralist and generative models, wey focus on serial verbs, splitting verbs, vowel harmony, and tone. |
| 05 | Teaching and Learning Yorùbá | History of language policy for Nigeria, di Ifẹ̀ Six-Year Primary Project, classroom curricula, diaspora pedagogy, and di November 2025 mother-tongue policy reversal. |
Yorùbá Text and Speech Corpora dey review di open-access and proprietary text corpora, parallel datasets, and voice archives wey dey available for research [S1, S5]. E document how religious domain parallel text dominate, such as JW300 and multilingual Bible corpora, alongside multi-domain datasets like MENYO-20k, wey bring standard benchmarks for news, proverbs, books, and digital localization [S1,]. For speech side, di file give details of collections wey start from Mozilla Common Voice and BibleTTS reach recent community initiatives like ÌròyìnSpeech and NaijaVoices [S5,]. Wetin still remain undocumented and contested for corpus linguistics na di representation of non-standard varieties: text and speech corpora wey dey available mostly favor Standard Written Yorùbá and Ọ̀yọ́ phonology, wey make Eastern Yorùbá, North-West peripheral dialects, and informal urban slang severely underrepresented inside public computational repositories [S1, S5].
Yorùbá Natural Language Processing and Speech Technology dey examine di computational techniques and engineering challenges of applying machine learning models to language wey get tone and agglutinative morphology [S2, S6]. E provide detailed technical breakdown of Automatic Diacritic Restoration (ADR), wey compare early syllable-based and n-gram methods wey Nigerian computer scientists develop with modern sequence-to-sequence neural architectures [S6,, S7]. Di file analyze tokenization bottlenecks, showing how subword tokenizers designed for English dey fragment diacritized Yorùbá vowels into disconnected bytes, wey dey reduce downstream performance inside Automatic Speech Recognition (ASR) and Neural Machine Translation (NMT) [S2,]. Wetin dem never resolve na weda massively multilingual foundation models genuinely dey learn Yorùbá semantics or dem just dey rely on surface lexical matching wey dey fail whenever complex tonal puns, honorific constructions, or idiomatic proverbs dey inside [S2, S6].
Yorùbá Dictionaries and Lexicography dey trace di history of dictionary compilation across almost two centuries [S3, S8]. Starting with Samuel Àjàyí Crowther 1843 Vocabulary of the Yoruba Language and di 1852 revision, di file go through Thomas Jefferson Bowen 1858 compilation, di 1913 Church Missionary Society (CMS) dictionary, reach Roy Clive Abraham landmark 1958 Dictionary of Modern Yoruba [S3, S8, S9]. E also review modern electronic lexicography, including Yíwọlá Awóyalé multi-decade lexical database project . Di file show how early missionary lexicographers alter cultural and cosmological terms wen dem force indigenous religious concepts enter Victorian Christian categories, especially wen dem translate Èṣù as "Satan" or "devil" and ẹbọ as "heathen sacrifice" [S3, S8]. Wetin people still dey debate for modern lexicography na di proper lemma structure for phrasal compounds, ideophones, and splitting verbs, alongside standard lexicographical representation of dialectal variants [S8, S10].
Yorùbá Grammars and Linguistic Description dey give detailed breakdown of how descriptive and theoretical linguistic accounts of Yorùbá syntax, phonology, and morphology take develop [S11, S12]. E trace di movement from nineteenth-century Latinate grammars go reach Ida Ward 1952 phonetic studies, Ayọ̀ Bámgbóṣé's structuralist 1966 A Grammar of Yoruba, and Ọladele Awóbùlúyì 1978 Essentials of Yoruba Grammar [S11, S12, S13]. E look deep into major theoretical debates, like formal syntactic analysis of Serial Verb Constructions (SVCs) versus splitting verbs, di phonological rules of Advanced Tongue Root (ATR) vowel harmony for mid-vowels (/e, o/ versus /ẹ, ọ/), and di status of subject clitics [S12, S14]. Wetin still dey under theoretical debate na weda splitting verbs na distinct morphological units or dem na syntactic sequences of independent verbs wey dey operate under strict lexicalization constraints [S11, S12].
Teaching and Learning Yorùbá dey examine language planning, institutional pedagogy, and sociolinguistic policy for Nigeria and across di global diaspora [S4, S15]. E document di history of mother-tongue education research, with focus on Professor Aliu Babátúndé Fáfunwa Ifẹ̀ Six-Year Primary Project (1970–1978), wey empirically show say primary students wey learn all subjects for Yorùbá achieve better cognitive development and higher academic performance for secondary school English pass pupils wey dem teach with English immersion [S4, S15]. Dis file outline wetin dey inside Nigeria 1977 National Policy on Education and di 2022 National Language Policy, and e analyze di policy reversal wey dem announce for November 2025 wey cancel di mother-tongue primary mandate [S15, S16]. Wetin still dey bring contention na di pedagogical tension between psycholinguistic evidence wey support mother-tongue concept acquisition and di socio-economic prestige of English wey dey make parents for urban centers dey resist am [S4, S16].
For computational linguists and NLP researchers: Start with Yorùbá Text and Speech Corpora to understand how di available data take spread and where di limits dey, continue go Yorùbá Natural Language Processing and Speech Technology for modeling architectures and tokenization challenges, and check Yorùbá Grammars and Linguistic Description to master di grammatical constraints wey dey control tone and verb serialization.
For theoretical linguists and lexicographers: Begin with Yorùbá Grammars and Linguistic Description for core structural arguments, move to Yorùbá Dictionaries and Lexicography to evaluate lexical classification and historical missionary bias, and reference Yorùbá Text and Speech Corpora for empirical verification across different text genres.
For educators, historians, and policy analysts: Read Teaching and Learning Yorùbá first to understand di political and pedagogical history of language planning for Nigeria, den turn to Yorùbá Dictionaries and Lexicography to understand di colonial codification of di language, and conclude with Yorùbá Natural Language Processing and Speech Technology to examine digital language preservation.
Dis section dey strictly focus on linguistic analysis, text and speech corpora, lexicographical reference, computational tooling, and educational policy.
Di formal oral literature of di Yorùbá tradition, like di 256 Odù Ifá divination poetry, oríkì praise poetry, orature, and proverbs (òwe), belong to 02-language, 05-ifa, and 06-orisa. Di metaphysical principles of orí (destiny/inner head), àṣẹ (generative spiritual force), and ìwà (character) dey treated for 04-cosmology. Sociolinguistic discussions about di wider Niger-Congo language family and historical dialect migration belong to 01-the-yoruba-language.
All claims about historical dates, grammar publications, corpus distributions, and official policy shifts dey grounded directly inside di scholarly literature wey dey documented below.
A comprehensive scholarly survey of documented Yorùbá lexical databases, parallel machine translation corpora, named entity benchmarks, and speech datasets.
A comprehensive scholarly survey of Yorùbá computational linguistics, covering automatic diacritic restoration, machine translation, speech recognition, speech synthesis, and community benchmarks.
The historical development of Yoruba dictionary-making from nineteenth-century missionary foundations through modern computational lexical databases.
A scholarly analysis of major reference grammars, foundational syntactic debates, and dialectological classifications in Yorùbá linguistics.
An analysis of Yorùbá language pedagogy, instructional materials, second-language tone acquisition research, and mother-tongue education policy in Nigeria.
An analysis of the Yorùbá three-level register tone system, demonstrating through minimal pairs and grammatical operations why tone is an essential phonemic component rather than optional decoration.