Ṣí ìlànà kan fún ìdí rẹ̀ àti orísun tí a fàyọ.
Compiled by Yíwọlá Awóyalé and distributed via the Linguistic Data Consortium, the database contains over 450,000 entries. It documents Standard Yorùbá alongside Atlantic diaspora varieties including Lucumí, Gullah, and Trinidadian Yorùbá, providing baseline lexical data for historical and computational linguistics.
Compiled by Yíwọlá Awóyalé and published through the Linguistic Data Consortium (LDC) in 2008, the *Global Yorùbá Lexical Database* is the largest structured electronic dictionary and lexical repository for Yorùbá and related varieties [S1].
MasakhaNER created the standard benchmark for African named entity recognition using human annotations from news publications. The benchmark demonstrated that models must account for orthographic tone and sub-dots to achieve reliable boundary detection and entity classification.
MasakhaNER addressed the persistent failure of multilingual zero-shot transfer from non-tonal Indo-European models by demonstrating that in-language, human-annotated tokens carrying correct tone and sub-dot marks produce superior boundary detection and classification accuracy [S2].
Previous translation resources were largely confined to narrow domains like religious prose. MENYO-20k established a multi-domain corpus of 20,100 parallel sentence pairs spanning news, literature, media transcripts, and proverbs.
Created by David Ifeoluwa Adelani, Dana Ruiter, Jesujoba O. Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, Cristina España-Bonet, and Dietrich Klakow, MENYO-20k was designed to resolve the extreme domain over-specialization that previously affected Yorùbá translation research [S3][S4].
Extracted from religious web publications, JW300 had provided an extensive parallel text resource across hundreds of low-resource languages. Formal copyright challenges by rights holders subsequently halted its public distribution.
The distribution of JW300 was subsequently halted and restricted due to formal copyright infringement challenges brought by the source copyright holders [S6].
Because everyday digital text is produced without dedicated keyboard layouts, authors typically omit tone marks and sub-dots. This omission causes severe lexical ambiguity where a single unmarked word can represent multiple unrelated meanings.
Because standard web scraping gathers Yorùbá text written on keyboards lacking native character support, web-mined text is overwhelmingly undiacritized, creating a severe bottleneck for downstream language processing [S1][S3].
Subword algorithms such as BPE and WordPiece frequently split accented characters into base Latin vowels and isolated Unicode diacritics. This fragmentation inflates sequence lengths and introduces errors during downstream neural model generation.
Standard tokenizers frequently split accented vowels into base characters and separate Unicode combining diacritics, fragmenting Yorùbá words into multiple sub-tokens.
The Masakhane network demonstrated that general massive multilingual models fail on African languages without community-centered curation. Involving native speakers in annotation and validation produces far more reliable models than uncurated web scraping.
The study established that standard massive multilingual models severely underperform on African languages unless local researchers and native speakers are directly involved in dataset curation and output validation [S2].
Diacritic restoration models trained on clean religious or news texts fail to generalize to informal digital registers. When applied to social media text containing slang and code-switching, these models experience steep performance declines.
Models trained on clean, diacritized religious texts (such as biblical translations) experience severe accuracy drops when evaluated on modern social media text or informal news reports containing code-switching and slang [S1][S3].
Crowther compiled the first major alphabetic lexicons of the language in 1843 and 1852. His work introduced the systematic use of sub-dots and accent marks for tone, creating the orthographic foundation for subsequent print culture.
The foundation of modern Yorùbá lexicography was laid by Samuel Àjàyí Crowther, a native Yorùbá speaker from Ọ̀ṣogun who was liberated from slavery in 1822, educated in Freetown and England, and later consecrated as the first African Anglican bishop [S1][S2].
Departing from utilitarian missionary compilations, Abraham's 1958 dictionary applied structural linguistic methodology to record every tone on every syllable. He also integrated comprehensive ethnographic essays and botanical identifications into entry definitions.
Abraham systematically marked every tone across every syllable in every entry, including mid tones and dynamic tonal glides/contours (such as high-to-low or low-to-high compound pitches) that previous lexicographers had overlooked or ignored [S5].
Delanọ published *Atúmọ̀ Èdè Yorùbá* to explain cultural concepts and kinship systems using Yorùbá metalanguage. He argued that defining terms through native communicative patterns prevented the semantic flattening inherent in direct English equivalents.
Delanọ recognized that explaining Yorùbá concepts through Yorùbá metalanguage allowed for precise cultural definitions that English equivalents inherently flattened [S6].
Applying M. A. K. Halliday's Scale and Category model, Bamgboṣe organized Yorùbá grammar into a rigorous hierarchy of sentence, clause, group, word, and morpheme. His work established standard analyses for clause classes, verbal particles, and tonal morphology.
Ayọ̀ Bamgboṣe's *A Grammar of Yoruba* represents the first full-scale structural description of Standard Yorùbá formulated within a contemporary general linguistic framework [S1].
Awobuluyi rejected the uncritical adoption of European parts of speech such as adjectives and adverbs in Yorùbá linguistics. Using distributional syntactic tests, he reorganized the lexicon into functional categories grounded directly in the language's own structural behavior.
Ọladele Awobuluyi's *Essentials of Yoruba Grammar* introduced a structural, distribution-based taxonomy designed explicitly to liberate Yorùbá grammatical description from Eurocentric and Latin-based categories [S2].
While earlier scholars treated sequences containing these markers as serial verb constructions, Awobuluyi showed that these items lack full verbal behavior. Because they cannot head independent predicates or undergo standard nominalization, he classified them as prepositions or auxiliaries.
In contrast, Awobuluyi applied strict distributional tests to demonstrate that items like *fi* and *bá* lack the full syntactic freedom of primary lexical verbs [S2].
Early university instruction focused heavily on structural grammar parsing, abstract phonological rules, and translation exercises. Over recent decades, curricula have realigned around functional task-based learning and four-skills development.
Over the past four decades, university-level instruction has shifted from historical grammar-translation models toward communicative, task-based frameworks aligned with international proficiency guidelines [S1][S2].
Schleicher aligned North American university instruction with ACTFL proficiency guidelines through foundational textbooks like *Jẹ́ K'Á Sọ Yorùbá*. Her materials replaced isolated verb memorization with authentic cultural dialogues and situational tasks.
This transition was led by Antonia Yétúndé Fọlárìn Schleicher, whose textbooks shifted classroom focus from rote memorization of verb paradigms to integrated four-skills development (speaking, listening, reading, and writing) situated within authentic cultural contexts [S1].