Yorùbá Dictionaries and Lexicography
The historical development of Yoruba dictionary-making from nineteenth-century missionary foundations through modern computational lexical databases.
The historical development of Yoruba dictionary-making from nineteenth-century missionary foundations through modern computational lexical databases.
Yorùbá lexicography is the systematic documentation, definition, and grammatical classification of the Yorùbá lexicon across printed dictionaries and digital databases. The tradition spans nearly two centuries, beginning with missionary-linguistic vocabularies designed for proselytization and standardisation in the 1840s and developing into modern computational corpora containing hundreds of thousands of entries. Across this trajectory, lexicographers have confronted structural challenges unique to the language, including the notation of phonological tone, the lexicographical handling of agglutinative morphology, dialectal divergence, and the recording of diaspora lexical varieties.
The formal documentation of the Yorùbá lexicon in alphabetic form began in the middle of the nineteenth century through the collaboration between indigenous Yorùbá speakers, liberated Africans in Sierra Leone, and European Christian missionaries .
The foundation of modern Yorùbá lexicography was laid by Samuel Àjàyí Crowther, a native Yorùbá speaker from Ọ̀ṣogun who was liberated from slavery in 1822, educated in Freetown and England, and later consecrated as the first African Anglican bishop .
Crowther’s first major lexical work was the Vocabulary of the Yoruba Language, published in 1843 by the Church Missionary Society in London . This text was compiled primarily among the Yorùbá-speaking diaspora (known as the Aku) in Freetown, Sierra Leone. The 1843 vocabulary established the initial Latin-based orthographic framework for Yorùbá. Crowther introduced lexical proverbs within entries to demonstrate how words functioned in natural discourse, recognising early on that Yorùbá words derived much of their contextual meaning from figurative usage . However, this initial volume lacked a comprehensive or fully systematic tonal notation system, leaving many homographs undifferentiated in print .
In 1852, Crowther published an extensively expanded version: A Vocabulary of the Yoruba Language, with introductory remarks by O. E. Vidal, issued by Seeleys in London . This 1852 edition was a watershed moment in West African linguistics. In it, Crowther introduced systematic diacritical marks:
Crowther’s 1852 volume stabilized nineteenth-century missionary print culture and created the orthographic foundation upon which subsequent educational, literary, and biblical translations were constructed .
Following Crowther's initial works, American Southern Baptist missionary T. J. Bowen published the Grammar and Dictionary of the Yoruba Language in 1858, sponsored and issued as Volume 10 of the Smithsonian Contributions to Knowledge series by the Smithsonian Institution in Washington, D.C.
Bowen’s work provided an independent, detailed lexicographical survey mapping Yorùbá vocabulary, grammatical rules, and English translations . Operating in central Yorùbá regions, Bowen gathered lexical material that supplemented Crowther’s texts, providing American and European scholarly audiences with an extensive bilingual glossary and structural grammatical sketch .
By the early twentieth century, the consolidation of colonial rule and the rapid expansion of mission schools in Southern Nigeria generated the need for a standardized, portable bilingual dictionary for students, teachers, and colonial administrative officers .
In 1913, the Church Missionary Society published A Dictionary of the Yoruba Language through the CMS Bookshop in Lagos . Later reprinted extensively by Oxford University Press in 1937 and 1950, this volume became the canonical reference dictionary for Yorùbá education for over half a century .
The CMS dictionary was organized into two distinct sections:
While the 1913 CMS dictionary achieved unmatched distribution and institutional standing, modern linguists and lexicographers note its methodological limitations . Its entries prioritized concise, utilitarian definitions suited for colonial schooling rather than deep structural or ethnolinguistic analysis. Furthermore, its tone marking was frequently inconsistent, leaving students and foreign researchers to infer tonal patterns and morphological boundaries that were not explicitly printed on the page .
The publication of Roy Clive Abraham’s Dictionary of Modern Yoruba in 1958 by the University of London Press transformed Yorùbá lexicography by applying structural linguistic methodology and extensive phonetic documentation to the language .
Abraham, an accomplished British colonial linguist who had previously worked on Hausa, Tiv, and Amharic, approached Yorùbá with a descriptive rigour that departed sharply from earlier missionary compilations .
Abraham’s dictionary remains one of the most cited lexical resources in African linguistics, establishing a benchmark for phonetic precision .
Despite the phonetic achievements of Abraham, indigenous scholars observed that European compilations often failed to capture the natural semantic structures, syntactic valency, and authentic idioms of native speakers . Chief Isaac Oluwole Delanọ, a prominent Yorùbá author, linguist, and cultural educator, addressed these limitations through two pioneering dictionaries .
Published by Oxford University Press in the same year as Abraham's work, Delanọ’s Atúmọ̀ Èdè Yorùbá: A Short Yoruba Grammar and Dictionary broke with the tradition of bilingual glossing by introducing monolingual Yorùbá definitions alongside concise English equivalents (a Yorùbá-Yorùbá-English format) .
Delanọ recognized that explaining Yorùbá concepts through Yorùbá metalanguage allowed for precise cultural definitions that English equivalents inherently flattened . Atúmọ̀ Èdè Yorùbá provided:
In 1969, the Institute of African Studies at the University of Ife published Delanọ’s two-volume work, A Dictionary of Yoruba Monosyllabic Verbs .
Recognizing that the monosyllabic verb root is the structural engine of the Yorùbá language, Delanọ isolated verbs and organized entries strictly by root consonants from b through y . For each verb root, Delanọ documented:
Delanọ’s focus on verbal syntax addressed a major gap in the historical record, providing syntactic models that informed later structural grammars of the language .
The transition from print lexicography to digital database architecture culminated in 2008 with the publication of the Global Yorùbá Lexical Database v. 1.0, compiled by Yíwọlá Awóyalé and published by the Linguistic Data Consortium (LDC) at the University of Pennsylvania (Catalog No. LDC2008L03) .
Awóyalé’s lexical database represents the largest computational corpus of the Yorùbá language assembled to date, containing over 368,000 distinct entries in its core database and expanding past 450,000 annotated lexical records across related computational sets .
The Global Yorùbá Lexical Database was constructed not merely as an electronic word list, but as a fully parsed linguistic resource delivered in computational formats, including SIL Toolbox and structured XML . Its major structural dimensions include:
Awóyalé’s work created a bridge between historical descriptive linguistics, diaspora studies, and natural language processing (NLP), serving as the foundational resource for digital Yorùbá lexicology .
In the twenty-first century, the expansion of the internet, mobile technologies, and open-access research led to the emergence of community-driven, web-native lexical projects .
Founded and directed by linguist and writer Kọ́lá Túbọ̀sún, the Yorùbá Name Project launched YorubaName.com in 2015, followed by the expanded lexical platform YorubaWord.com .
This project addresses the limitations of static print dictionaries through a multimedia, open-access lexical database . Its features include:
Collaborative digital repositories, such as the Yorùbá Wiktionary and the Glosbe Yorùbá lexical database, provide crowdsourced, bidirectional English–Yorùbá translations . These platforms host open-source collections of translated sentence pairs, morphological notes, and modern technological vocabulary, though their crowdsourced nature results in variable diacritic consistency .
The history of Yorùbá lexicography is marked by enduring debates concerning linguistic authority, orthographic conventions, dialectal inclusion, and the limits of non-native documentation.
A major debate in African lexicography concerns the comparative standing of non-native linguistic scholars versus native-speaker compilers .
In a critical retrospective published in the Yorùbá Studies Review, Toyin Falola and Michael Oladejo Afolayan (2021) evaluated the contributions of Isaac Oluwole Delanọ relative to European predecessors such as R. C. Abraham . Falola and Afolayan observed that despite Abraham’s phonetic precision, his status as a non-native outsider produced recurring structural shortcomings:
Conversely, Delanọ’s native speaker intuition allowed him to capture the fluid idiomatic nuances of Yorùbá proverbs, colloquial banter, and syntactic valency without distorting the cultural register of the language .
A historical question persists regarding the division of authorship in early CMS publications . Scholars debate the extent to which Samuel Àjàyí Crowther’s early manuscripts were independently derived from his linguistic analysis versus shaped by the editorial interventions and Latinate preconceptions of European CMS supervisors, such as John Raban and C. A. Gollmer . While Crowther was the native speaker and primary compiler, the mission committee exercised editorial oversight over grammatical categories and scriptural glosses .
The method and density of tone marking remain subjects of debate across educational and computational domains .
The historical record displays a clear regional bias. From the 1840s through the mid-twentieth century, missionary and colonial lexicographers focused almost exclusively on the central-western dialect varieties, specifically Ọ̀yọ́ and Ègbá . This focus formed the basis of what became codified as "Standard Yorùbá" (Yorùbá Àjùmọ̀lò) .
Consequently, the historical record is largely silent regarding systematic, comprehensive lexical documentation of eastern and southeastern dialect clusters, including:
These regional varieties possess distinct phonological structures, verbal particles, and lexical items that were largely excluded from early canonical dictionaries .
In digital lexicography, scholars debate how diaspora registers should be classified within Yorùbá databases . While Yíwọlá Awóyalé integrated Lucumí, Trinidadian Yorùbá, and Gullah within the Global Yorùbá Lexical Database as historical continuations of the Yorùbá lexicon, some linguists argue that these registers have undergone significant syncretism, grammatical restructuring, and lexical creolization through contact with Spanish, French Creole, and English, requiring careful distinction from historical West African Yorùbá .
As Yorùbá lexicography operates in digital and computational spaces, specific technical and linguistic hurdles continue to affect the language .
As demonstrated by Natural Language Processing (NLP) researchers, including Irorun Orife et al. (2020) in Improving Yorùbá Diacritic Restoration, the majority of digitized Yorùbá text across the internet completely omits tone marks and subdots .
Because Yorùbá relies on pitch and vowel quality for lexical differentiation, the absence of diacritics creates severe homographic ambiguity . For example, the un-diacriticized string owo can represent at least four entirely unrelated words:
Natural language processing tools, search engines, and automated translation platforms frequently fail to process Yorùbá texts accurately because base digital vocabularies lack consistent diacritical restoration .
The scholarly record reflects that while several academic institutions and independent research groups have created semantic ontology prototypes and bilingual machine-translation wordlists, a complete, standardized, and fully unified open-access Yorùbá WordNet (equivalent to the Princeton WordNet for English) has not yet been finalized . The development of such a semantic network remains one of the major uncompleted tasks in modern Yorùbá computational lexicography.
| Year | Author / Compiler | Title | Publisher / Institution | Key Structural Features | Tone Notation System |
|---|---|---|---|---|---|
| 1843 | Samuel Àjàyí Crowther | Vocabulary of the Yoruba Language | Church Missionary Society (London) | First native-compiled vocabulary; introduced proverbs into entries . | Rudimentary / non-systematic . |
| 1852 | Samuel Àjàyí Crowther | A Vocabulary of the Yoruba Language | Seeleys / CMS (London) | Orthographic milestone; established subdots (ẹ, ọ, ṣ) and grammatical elements . | High (´) and Low (`) marked; Mid unmarked . |
| 1858 | T. J. Bowen | Grammar and Dictionary of the Yoruba Language | Smithsonian Institution (Washington, D.C.) | Comprehensive early survey published for academic research in the Americas . | Systematic basic marking . |
| 1913 | Church Missionary Society (CMS) | A Dictionary of the Yoruba Language | CMS Bookshop (Lagos) / Oxford University Press | Standard bilingual pedagogical reference for schools and colonial civil service . | High and Low marked; frequent omissions . |
| 1958 | Roy Clive Abraham | Dictionary of Modern Yoruba | University of London Press | Exhaustive phonetic, morphosyntactic, and encyclopedic/cultural entries . | Comprehensive: all syllables, contour tones, and mid pitches fully marked . |
| 1958 | Isaac Oluwole Delanọ | Atúmọ̀ Èdè Yorùbá | Oxford University Press (London) | Pioneered monolingual Yorùbá definitions (Yorùbá-Yorùbá-English format) . | High and Low marked; standard pedagogical . |
| 1969 | Isaac Oluwole Delanọ | A Dictionary of Yoruba Monosyllabic Verbs | Institute of African Studies, University of Ife | Focused strictly on verb root behavior, syntactic valency, and complementation . | Fully marked on verb roots and examples . |
| 2008 | Yíwọlá Awóyalé | Global Yorùbá Lexical Database v. 1.0 | Linguistic Data Consortium, Univ. of Pennsylvania | Over 368,000 computational entries; morphemic parsing; Atlantic diaspora coverage . | Fully encoded three-level computational tonology . |
| 2015– | Kọ́lá Túbọ̀sún et al. | YorubaName.com / YorubaWord.com | Yorùbá Name Project | Open-access multimedia database; etymological parsing; audio recordings / TTS . | Fully marked tone diacritics paired with audio . |
The CMS mission and Crowther, how and why Yoruba people converted to two world religions, the British annexation, indirect rule and what it did to kingship, and the making of "Yoruba" as one identity.
A scholarly analysis of major reference grammars, foundational syntactic debates, and dialectological classifications in Yorùbá linguistics.
How Yorùbá came to be written in Arabic and then Roman script, who decided the spelling rules, and why the subdots and tone marks are not optional.
A comprehensive scholarly survey of documented Yorùbá lexical databases, parallel machine translation corpora, named entity benchmarks, and speech datasets.
A comprehensive scholarly survey of Yorùbá computational linguistics, covering automatic diacritic restoration, machine translation, speech recognition, speech synthesis, and community benchmarks.
An analysis of Yorùbá language pedagogy, instructional materials, second-language tone acquisition research, and mother-tongue education policy in Nigeria.