Yorùbá Dictionaries and Lexicography
The historical development of Yoruba dictionary-making from nineteenth-century missionary foundations through modern computational lexical databases.
Yorùbá lexicography is the systematic documentation, definition, and grammatical classification of the Yorùbá lexicon across printed dictionaries and digital databases. The tradition spans nearly two centuries, beginning with missionary-linguistic vocabularies designed for proselytization and standardisation in the 1840s and developing into modern computational corpora containing hundreds of thousands of entries. Across this trajectory, lexicographers have confronted structural challenges unique to the language, including the notation of phonological tone, the lexicographical handling of agglutinative morphology, dialectal divergence, and the recording of diaspora lexical varieties.
Nineteenth-Century Foundations: Crowther and Bowen
The formal documentation of the Yorùbá lexicon in alphabetic form began in the middle of the nineteenth century through the collaboration between indigenous Yorùbá speakers, liberated Africans in Sierra Leone, and European Christian missionaries .
Samuel Àjàyí Crowther (1843, 1852)
The foundation of modern Yorùbá lexicography was laid by Samuel Àjàyí Crowther, a native Yorùbá speaker from Ọ̀ṣogun who was liberated from slavery in 1822, educated in Freetown and England, and later consecrated as the first African Anglican bishop .
Crowther’s first major lexical work was the Vocabulary of the Yoruba Language, published in 1843 by the Church Missionary Society in London . This text was compiled primarily among the Yorùbá-speaking diaspora (known as the Aku) in Freetown, Sierra Leone. The 1843 vocabulary established the initial Latin-based orthographic framework for Yorùbá. Crowther introduced lexical proverbs within entries to demonstrate how words functioned in natural discourse, recognising early on that Yorùbá words derived much of their contextual meaning from figurative usage . However, this initial volume lacked a comprehensive or fully systematic tonal notation system, leaving many homographs undifferentiated in print .
In 1852, Crowther published an extensively expanded version: A Vocabulary of the Yoruba Language, with introductory remarks by O. E. Vidal, issued by Seeleys in London . This 1852 edition was a watershed moment in West African linguistics. In it, Crowther introduced systematic diacritical marks:
- Subdots beneath open vowels: ẹ (representing the open-mid front unrounded vowel [ɛ]) and ọ (representing the open-mid back rounded vowel [ɔ]) .
- A subdot beneath the consonant ṣ (representing the voiceless postalveolar fricative [ʃ]) to distinguish it from the alveolar sibilant s .
- Tone diacritics: acute accents (´) for high pitch, grave accents (`) for low pitch, and the absence of marks or explicit indicators for mid pitch .
Crowther’s 1852 volume stabilized nineteenth-century missionary print culture and created the orthographic foundation upon which subsequent educational, literary, and biblical translations were constructed .
T. J. Bowen (1858)
Following Crowther's initial works, American Southern Baptist missionary T. J. Bowen published the Grammar and Dictionary of the Yoruba Language in 1858, sponsored and issued as Volume 10 of the Smithsonian Contributions to Knowledge series by the Smithsonian Institution in Washington, D.C.
Bowen’s work provided an independent, detailed lexicographical survey mapping Yorùbá vocabulary, grammatical rules, and English translations . Operating in central Yorùbá regions, Bowen gathered lexical material that supplemented Crowther’s texts, providing American and European scholarly audiences with an extensive bilingual glossary and structural grammatical sketch .
The Pedagogical and Missionary Standard: CMS (1913)
By the early twentieth century, the consolidation of colonial rule and the rapid expansion of mission schools in Southern Nigeria generated the need for a standardized, portable bilingual dictionary for students, teachers, and colonial administrative officers .
In 1913, the Church Missionary Society published A Dictionary of the Yoruba Language through the CMS Bookshop in Lagos . Later reprinted extensively by Oxford University Press in 1937 and 1950, this volume became the canonical reference dictionary for Yorùbá education for over half a century .
The CMS dictionary was organized into two distinct sections:
- Yorùbá–English: Designed to provide concise definitions, lexical categories, and idiomatic phrases for Yorùbá speakers learning English and for non-native missionaries reading Yorùbá texts .
- English–Yorùbá: Designed as a reference for translation from the colonial administrative language into Yorùbá .
While the 1913 CMS dictionary achieved unmatched distribution and institutional standing, modern linguists and lexicographers note its methodological limitations . Its entries prioritized concise, utilitarian definitions suited for colonial schooling rather than deep structural or ethnolinguistic analysis. Furthermore, its tone marking was frequently inconsistent, leaving students and foreign researchers to infer tonal patterns and morphological boundaries that were not explicitly printed on the page .
Phonological Rigour and Encyclopedic Lexicography: R. C. Abraham (1958)
The publication of Roy Clive Abraham’s Dictionary of Modern Yoruba in 1958 by the University of London Press transformed Yorùbá lexicography by applying structural linguistic methodology and extensive phonetic documentation to the language .
Abraham, an accomplished British colonial linguist who had previously worked on Hausa, Tiv, and Amharic, approached Yorùbá with a descriptive rigour that departed sharply from earlier missionary compilations .
Key Innovations of Abraham (1958)
- Systematic Tone Marking: Abraham systematically marked every tone across every syllable in every entry, including mid tones and dynamic tonal glides/contours (such as high-to-low or low-to-high compound pitches) that previous lexicographers had overlooked or ignored .
- Morphosyntactic Classification: Rather than forcing Yorùbá words into rigid European grammatical categories, Abraham mapped complex morphosyntactic behaviours, noting split-verbs, verb-nominal collocations, and grammatical particle distributions .
- Encyclopedic and Cultural Scope: Abraham transformed his dictionary into an ethnographic reference work. Entries for botanical specimens, cultural ceremonies, religious offices, anatomical terms, and historical events were accompanied by extensive cultural essays, line drawings, and botanical cross-references .
Abraham’s dictionary remains one of the most cited lexical resources in African linguistics, establishing a benchmark for phonetic precision .
Native-Speaker Syntax and Monolingual Lexicography: Isaac Oluwole Delanọ
Despite the phonetic achievements of Abraham, indigenous scholars observed that European compilations often failed to capture the natural semantic structures, syntactic valency, and authentic idioms of native speakers . Chief Isaac Oluwole Delanọ, a prominent Yorùbá author, linguist, and cultural educator, addressed these limitations through two pioneering dictionaries .
Atúmọ̀ Èdè Yorùbá (1958)
Published by Oxford University Press in the same year as Abraham's work, Delanọ’s Atúmọ̀ Èdè Yorùbá: A Short Yoruba Grammar and Dictionary broke with the tradition of bilingual glossing by introducing monolingual Yorùbá definitions alongside concise English equivalents (a Yorùbá-Yorùbá-English format) .
Delanọ recognized that explaining Yorùbá concepts through Yorùbá metalanguage allowed for precise cultural definitions that English equivalents inherently flattened . Atúmọ̀ Èdè Yorùbá provided:
- Native-speaker descriptions of culturally specific terms, rituals, kinship relations, and social institutions .
- Grammatical explanations framed through Yorùbá communicative patterns rather than Latinate pedagogical models .
- Clear orthographic guides suited for secondary schools and higher education .
A Dictionary of Yoruba Monosyllabic Verbs (1969)
In 1969, the Institute of African Studies at the University of Ife published Delanọ’s two-volume work, A Dictionary of Yoruba Monosyllabic Verbs .
Recognizing that the monosyllabic verb root is the structural engine of the Yorùbá language, Delanọ isolated verbs and organized entries strictly by root consonants from b through y . For each verb root, Delanọ documented:
- Verb-object collocations and syntactic complementation rules .
- Transitive and intransitive alternations .
- Semantic shifts that occur when monosyllabic roots combine with different nominal arguments .
Delanọ’s focus on verbal syntax addressed a major gap in the historical record, providing syntactic models that informed later structural grammars of the language .
Computational Databases and Atlantic Diaspora Lexicography: Yíwọlá Awóyalé (2008)
The transition from print lexicography to digital database architecture culminated in 2008 with the publication of the Global Yorùbá Lexical Database v. 1.0, compiled by Yíwọlá Awóyalé and published by the Linguistic Data Consortium (LDC) at the University of Pennsylvania (Catalog No. LDC2008L03) .
Awóyalé’s lexical database represents the largest computational corpus of the Yorùbá language assembled to date, containing over 368,000 distinct entries in its core database and expanding past 450,000 annotated lexical records across related computational sets .
Structural Architecture of the Database
The Global Yorùbá Lexical Database was constructed not merely as an electronic word list, but as a fully parsed linguistic resource delivered in computational formats, including SIL Toolbox and structured XML . Its major structural dimensions include:
- Morphemic Decomposition: Every entry is parsed into its constituent morphological elements, explicitly marking root morphemes, prefixes, infixes, reduplication patterns, and elided vowels .
- Full Tonological Notation: The database uses rigorous three-level tone encoding, resolving ambiguities inherent in standard orthography .
- Atlantic Diaspora Lexical Mapping: Unlike all previous dictionaries, Awóyalé incorporated vocabulary from Yorùbá-derived diaspora varieties across the Atlantic basin, explicitly cataloguing:
- Lucumí (the liturgical and cultural register preserved in Cuba) .
- Trinidadian Yorùbá (lexical survivals and ritual language documented in Trinidad and Tobago) .
- Gullah lexical retentions in the Sea Islands of the southeastern United States .
Awóyalé’s work created a bridge between historical descriptive linguistics, diaspora studies, and natural language processing (NLP), serving as the foundational resource for digital Yorùbá lexicology .
Contemporary Digital and Multimedia Lexical Projects
In the twenty-first century, the expansion of the internet, mobile technologies, and open-access research led to the emergence of community-driven, web-native lexical projects .
YorubaName.com and YorubaWord.com (2015–Present)
Founded and directed by linguist and writer Kọ́lá Túbọ̀sún, the Yorùbá Name Project launched YorubaName.com in 2015, followed by the expanded lexical platform YorubaWord.com .
This project addresses the limitations of static print dictionaries through a multimedia, open-access lexical database . Its features include:
- Etymological and Morphological Deconstructions: Breaking down personal names and compound nouns into their component phrases, verbs, and nominal qualifiers .
- Phonetic Tone Marks and Audio Pronunciation: Pairing every entry with both standard tone diacritics and human-recorded audio files as well as text-to-speech (TTS) samples, ensuring that tonal melodies are preserved and accessible to diaspora learners and non-fluent speakers .
- Dialect Tracking: Documenting dialectal variants across Yorùbá-speaking sub-regions .
Collaborative Glossaries: Yorùbá Wiktionary and Glosbe
Collaborative digital repositories, such as the Yorùbá Wiktionary and the Glosbe Yorùbá lexical database, provide crowdsourced, bidirectional English–Yorùbá translations . These platforms host open-source collections of translated sentence pairs, morphological notes, and modern technological vocabulary, though their crowdsourced nature results in variable diacritic consistency .
Methodological Debates and Gaps in the Record
The history of Yorùbá lexicography is marked by enduring debates concerning linguistic authority, orthographic conventions, dialectal inclusion, and the limits of non-native documentation.
1. Native Authority versus Non-Native Compilation
A major debate in African lexicography concerns the comparative standing of non-native linguistic scholars versus native-speaker compilers .
In a critical retrospective published in the Yorùbá Studies Review, Toyin Falola and Michael Oladejo Afolayan (2021) evaluated the contributions of Isaac Oluwole Delanọ relative to European predecessors such as R. C. Abraham . Falola and Afolayan observed that despite Abraham’s phonetic precision, his status as a non-native outsider produced recurring structural shortcomings:
- Unidiomatic Examples: Abraham frequently constructed illustrative sentences that, while grammatically possible under structural rules, were culturally unnatural or unidiomatic to native speakers .
- Over-Complicated Cross-Referencing: Abraham’s organizational system relied on complex, idiosyncratic linguistic symbols and excessive sub-entry routing, making the dictionary difficult for general speakers and students to navigate .
Conversely, Delanọ’s native speaker intuition allowed him to capture the fluid idiomatic nuances of Yorùbá proverbs, colloquial banter, and syntactic valency without distorting the cultural register of the language .
2. Missionary Influence versus Indigenous Innovation in Crowther’s Work
A historical question persists regarding the division of authorship in early CMS publications . Scholars debate the extent to which Samuel Àjàyí Crowther’s early manuscripts were independently derived from his linguistic analysis versus shaped by the editorial interventions and Latinate preconceptions of European CMS supervisors, such as John Raban and C. A. Gollmer . While Crowther was the native speaker and primary compiler, the mission committee exercised editorial oversight over grammatical categories and scriptural glosses .
3. Tone Marking Conventions and the 1974 Orthography Reform
The method and density of tone marking remain subjects of debate across educational and computational domains .
- Full Phonetic Marking: Structural linguists (Abraham 1958, Awóyalé 2008) advocate for marking every tone, including mid tones and tonal contours, to eliminate ambiguity for computational algorithms and foreign learners .
- Pedagogical Parsimony: Standard educational practice, stemming from the CMS tradition and endorsed in popular print, marks only high (´) and low (`) tones, leaving mid tones unmarked to minimize visual clutter on the printed page .
- The 1974 Orthography Committee Divergence: The formalization of standardized rules by the Western State Yoruba Orthography Committee in 1974 introduced changes regarding vowel elision, word boundary separation, and the spelling of contracted verb-noun compounds. As a result, older canonical dictionaries (CMS 1913, Abraham 1958) diverge orthographically from post-1974 publications, requiring modern readers to navigate conflicting spelling standards .
4. Regional Bias and Historical Silences: The Ọ̀yọ́/Ègbá Standard
The historical record displays a clear regional bias. From the 1840s through the mid-twentieth century, missionary and colonial lexicographers focused almost exclusively on the central-western dialect varieties, specifically Ọ̀yọ́ and Ègbá . This focus formed the basis of what became codified as "Standard Yorùbá" (Yorùbá Àjùmọ̀lò) .
Consequently, the historical record is largely silent regarding systematic, comprehensive lexical documentation of eastern and southeastern dialect clusters, including:
- Ondó
- Ọ̀wọ̀
- Ìkálẹ̀
- Ìjẹ̀bú (outside of scattered comparative notes)
These regional varieties possess distinct phonological structures, verbal particles, and lexical items that were largely excluded from early canonical dictionaries .
5. Diaspora Varieties: Historical Retention versus Creolization
In digital lexicography, scholars debate how diaspora registers should be classified within Yorùbá databases . While Yíwọlá Awóyalé integrated Lucumí, Trinidadian Yorùbá, and Gullah within the Global Yorùbá Lexical Database as historical continuations of the Yorùbá lexicon, some linguists argue that these registers have undergone significant syncretism, grammatical restructuring, and lexical creolization through contact with Spanish, French Creole, and English, requiring careful distinction from historical West African Yorùbá .
Computational Lexical Resources: Challenges and Open Problems
As Yorùbá lexicography operates in digital and computational spaces, specific technical and linguistic hurdles continue to affect the language .
Diacritic Inconsistency and Lexical Ambiguity
As demonstrated by Natural Language Processing (NLP) researchers, including Irorun Orife et al. (2020) in Improving Yorùbá Diacritic Restoration, the majority of digitized Yorùbá text across the internet completely omits tone marks and subdots .
Because Yorùbá relies on pitch and vowel quality for lexical differentiation, the absence of diacritics creates severe homographic ambiguity . For example, the un-diacriticized string owo can represent at least four entirely unrelated words:
- owó (money)
- ọwọ́ (hand / broom / group)
- òwò (trade / commerce)
- ọ̀wọ̀ (respect / honor)
Natural language processing tools, search engines, and automated translation platforms frequently fail to process Yorùbá texts accurately because base digital vocabularies lack consistent diacritical restoration .
The Absence of a Unified Yorùbá WordNet
The scholarly record reflects that while several academic institutions and independent research groups have created semantic ontology prototypes and bilingual machine-translation wordlists, a complete, standardized, and fully unified open-access Yorùbá WordNet (equivalent to the Princeton WordNet for English) has not yet been finalized . The development of such a semantic network remains one of the major uncompleted tasks in modern Yorùbá computational lexicography.
Chronological Overview of Major Lexicographical Works
| Year | Author / Compiler | Title | Publisher / Institution | Key Structural Features | Tone Notation System |
|---|---|---|---|---|---|
| 1843 | Samuel Àjàyí Crowther | Vocabulary of the Yoruba Language | Church Missionary Society (London) | First native-compiled vocabulary; introduced proverbs into entries . | Rudimentary / non-systematic . |
| 1852 | Samuel Àjàyí Crowther | A Vocabulary of the Yoruba Language | Seeleys / CMS (London) | Orthographic milestone; established subdots (ẹ, ọ, ṣ) and grammatical elements . | High (´) and Low (`) marked; Mid unmarked . |
| 1858 | T. J. Bowen | Grammar and Dictionary of the Yoruba Language | Smithsonian Institution (Washington, D.C.) | Comprehensive early survey published for academic research in the Americas . | Systematic basic marking . |
| 1913 | Church Missionary Society (CMS) | A Dictionary of the Yoruba Language | CMS Bookshop (Lagos) / Oxford University Press | Standard bilingual pedagogical reference for schools and colonial civil service . | High and Low marked; frequent omissions . |
| 1958 | Roy Clive Abraham | Dictionary of Modern Yoruba | University of London Press | Exhaustive phonetic, morphosyntactic, and encyclopedic/cultural entries . | Comprehensive: all syllables, contour tones, and mid pitches fully marked . |
| 1958 | Isaac Oluwole Delanọ | Atúmọ̀ Èdè Yorùbá | Oxford University Press (London) | Pioneered monolingual Yorùbá definitions (Yorùbá-Yorùbá-English format) . | High and Low marked; standard pedagogical . |
| 1969 | Isaac Oluwole Delanọ | A Dictionary of Yoruba Monosyllabic Verbs | Institute of African Studies, University of Ife | Focused strictly on verb root behavior, syntactic valency, and complementation . | Fully marked on verb roots and examples . |
| 2008 | Yíwọlá Awóyalé | Global Yorùbá Lexical Database v. 1.0 | Linguistic Data Consortium, Univ. of Pennsylvania | Over 368,000 computational entries; morphemic parsing; Atlantic diaspora coverage . | Fully encoded three-level computational tonology . |
| 2015– | Kọ́lá Túbọ̀sún et al. | YorubaName.com / YorubaWord.com | Yorùbá Name Project | Open-access multimedia database; etymological parsing; audio recordings / TTS . | Fully marked tone diacritics paired with audio . |