Yorùbá Dictionaries and Lexicography
The historical development of Yoruba dictionary-making from nineteenth-century missionary foundations through modern computational lexical databases.
The historical development of Yoruba dictionary-making from nineteenth-century missionary foundations through modern computational lexical databases.
The Yorùbá wey dem quote, the proverbs, oríkì, ẹsẹ Ifá, word list headwords, Odù names and citations dey exactly as the corpus record dem, for every language.
Yorùbá lexicography na di systematic documentation, definition, and grammatical classification of Yorùbá vocabulary across printed dictionaries and digital databases. Dis tradition don reach almost two centuries, e start from missionary-linguistic vocabularies wey dem design for preaching and standardisation for di 1840s, reach modern computational corpora wey get hundreds of thousands of entries. Throughout dis journey, lexicographers don face structural challenges wey specific to di language, including how to mark phonological tone, how lexicography go handle agglutinative morphology, dialect differences, and how to record diaspora lexical varieties.
Di formal documentation of Yorùbá vocabulary for alphabet form start for di middle of di nineteenth century through collaboration between indigenous Yorùbá speakers, liberated Africans for Sierra Leone, and European Christian missionaries .
Na Samuel Àjàyí Crowther lay di foundation of modern Yorùbá lexicography. Him be native Yorùbá speaker from Ọ̀ṣogun wey dem liberate from slavery for 1822, e go school for Freetown and England, and later become di first African Anglican bishop .
Crowther first major lexical work na di Vocabulary of the Yoruba Language, wey Church Missionary Society publish for London for 1843 . Dem compile dis text mainly among di Yorùbá-speaking diaspora (wey dem dey call di Aku) for Freetown, Sierra Leone. Di 1843 vocabulary establish di first Latin-based orthographic framework for Yorùbá. Crowther put proverbs inside entries to show how words dey work for everyday talk, as e understand early say plenty contextual meaning of Yorùbá words dey come from figurative usage . However, dis first volume no get complete or fully systematic tone notation system, so many words wey dey spelled di same way (homographs) no get difference for print .
For 1852, Crowther publish one version wey e expand well well: A Vocabulary of the Yoruba Language, with introductory remarks from O. E. Vidal, wey Seeleys for London publish . Dis 1852 edition na turning point for West African linguistics. Inside am, Crowther introduce systematic diacritical marks:
Crowther 1852 volume stabilise nineteenth-century missionary print culture and build di orthographic foundation wey dem take do later educational, literary, and Bible translations .
After Crowther first works, American Southern Baptist missionary T. J. Bowen publish di Grammar and Dictionary of the Yoruba Language for 1858, wey Smithsonian Institution for Washington, D.C. sponsor and release as Volume 10 of di Smithsonian Contributions to Knowledge series .
Bowen work provide independent and detailed lexicographical survey wey map out Yorùbá vocabulary, grammar rules, and English translations . As e dey work for central Yorùbá regions, Bowen gather lexical material wey add to Crowther texts, and e give American and European academic audience wide bilingual glossary and structural grammar sketch .
By early twentieth century, as colonial rule strong and mission schools spread quick-quick for Southern Nigeria, need come dey for standardized, portable bilingual dictionary for students, teachers, and colonial administrative officers .
For 1913, Church Missionary Society publish A Dictionary of the Yoruba Language through CMS Bookshop for Lagos . Later on, Oxford University Press reprint am plenty times for 1937 and 1950, and dis volume come turn di main reference dictionary for Yorùbá education for over fifty years .
Dem arrange di CMS dictionary into two clear sections:
Even though di 1913 CMS dictionary spread reach everywhere and get strong institutional standing, modern linguists and lexicographers don point out di limitations for di method wey dem use . Di entries focus more on short, practical definitions wey fit colonial schooling instead of deep structural or ethnolinguistic analysis. Apart from dat, di way dem mark tone no dey consistent many times, wey make students and foreign researchers dey guess tone patterns and morphological boundaries wey dem no clearly print for di page .
Di publication of Roy Clive Abraham Dictionary of Modern Yoruba for 1958 by University of London Press change Yorùbá lexicography well well as e apply structural linguistic methodology and extensive phonetic documentation to di language .
Abraham, wey be accomplished British colonial linguist wey don work on Hausa, Tiv, and Amharic before, study Yorùbá with descriptive rigour wey different totally from earlier missionary compilations .
Abraham dictionary still remain one of di lexical resources wey people cite pass for African linguistics, as e set standard for phonetic precision .
Even with di phonetic achievements of Abraham, local scholars notice say European compilations often fail to capture di natural semantic structures, syntactic valency, and original idioms of native speakers . Chief Isaac Oluwole Delanọ, wey be prominent Yorùbá author, linguist, and cultural educator, tackle dese limitations through two pioneering dictionaries .
As Oxford University Press publish am for di same year as Abraham work, Delanọ’s Atúmọ̀ Èdè Yorùbá: A Short Yoruba Grammar and Dictionary break away from di tradition of bilingual glossing as e introduce monolingual Yorùbá definitions join concise English equivalents (a Yorùbá-Yorùbá-English format) .
Delanọ understand say to explain Yorùbá concepts through Yorùbá metalanguage make am possible to get accurate cultural definitions wey English equivalents naturally dey flatten . Atúmọ̀ Èdè Yorùbá provide:
For 1969, di Institute of African Studies for University of Ife publish Delanọ’s two-volume work, A Dictionary of Yoruba Monosyllabic Verbs .
As e understand say di monosyllabic verb root na di structural engine of di Yorùbá language, Delanọ separate verbs and arrange entries strictly by root consonants from b go reach y . For each verb root, Delanọ document:
Di focus wey Delanọ’s put on top verbal syntax solve one major gap for historical record, as e provide syntactic models wey guide later structural grammars of di language .
Di movement from paper lexicography go reach digital database architecture reach peak for 2008 when dem publish di Global Yorùbá Lexical Database v. 1.0, wey Yíwọlá Awóyalé compile and Linguistic Data Consortium (LDC) for University of Pennsylvania publish (Catalog No. LDC2008L03) .
Awóyalé lexical database na di biggest computational corpus of di Yorùbá language wey dem don gather till today, as e get pass 368,000 distinct entries for im main database and pass 450,000 annotated lexical records across related computational sets .
Dem build di Global Yorùbá Lexical Database no be just as ordinary electronic word list, but as fully parsed linguistic resource wey dey available in computational formats, including SIL Toolbox and structured XML . Di main structural dimensions include:
Awóyalé work form bridge between historical descriptive linguistics, diaspora studies, and natural language processing (NLP), as e serve as foundational resource for digital Yorùbá lexicology .
For di twenty-first century, as internet, mobile technology, and open-access research dey spread, e lead to community-driven, web-native lexical projects .
Linguist and writer Kọ́lá Túbọ̀sún found and direct di project, as di Yorùbá Name Project launch YorubaName.com for 2015, before dem come launch di expanded lexical platform YorubaWord.com .
Dis project dey solve di problems of normal printed dictionary wey no fit change, by using multimedia, open-access word database . Di things wey dey inside include:
Digital storehouse wey people dey contribute togeda, like Yorùbá Wiktionary and Glosbe Yorùbá word database, dey give crowdsourced translation between English and Yorùbá wey dey go both sides . Dese platforms dey host open-source collection of translated sentence pairs, notes about how words form, and modern technology vocabulary, even though say because na crowdsource, di way dem dey use tone mark no dey consistent .
Di history of Yorùbá dictionary work get long-standing debate about who get authority over language, spelling rules, how dem include dialects, and di limits of record wey people wey no be native speaker make.
One big debate for African dictionary work dey touch di position of linguistics scholar wey no be native speaker compared to people wey be native speaker wey compile words .
Inside one critical review wey dem publish for Yorùbá Studies Review, Toyin Falola and Michael Oladejo Afolayan (2021) assess wetin Isaac Oluwole Delanọ contribute compared to European scholars wey come before am like R. C. Abraham . Falola and Afolayan observe say even though Abraham tone and sound accurate well, because im be outsider wey no be native speaker, im work get some structural problem wey dey repeat:
On di oda hand, Delanọ’s native speaker understanding make am fit capture di natural meaning of Yorùbá proverbs, everyday gist, and sentence patterns without changing di cultural tone of di language .
One historical question still dey about who write wetin for early CMS publications . Scholars dey debate how much of Samuel Àjàyí Crowther early manuscript come from im own language analysis versus how much European CMS supervisors like John Raban and C. A. Gollmer change am through their editing and Latin-based ideas . Even though na Crowther be di native speaker and main compiler, di mission committee still control di grammar categories and bible explanations wey enter am .
Di method and how much tone mark to use still be matter of debate for education and computer language processing .
Di historical record show clear bias towards some regions. From di 1840s reach mid-twentieth century, missionary and colonial dictionary writers put almost all their attention on central-western dialects, especially Ọ̀yọ́ and Ègbá . Na dis focus form di foundation of wetin come become "Standard Yorùbá" (Yorùbá Àjùmọ̀lò) .
As result of dat, di historical record almost no talk about proper, complete dictionary documentation of eastern and southeastern dialect groups, including:
Dese regional dialects get their own sound structure, verb particles, and words wey early major dictionaries mostly leave out .
For digital dictionary work, scholars dey debate how to classify diaspora language forms inside Yorùbá databases . While Yíwọlá Awóyalé put Lucumí, Trinidadian Yorùbá, and Gullah inside di Global Yorùbá Lexical Database as historical continuation of Yorùbá vocabulary, some linguists argue say dese language forms don undergo heavy mixing, grammar changes, and word creolization through contact with Spanish, French Creole, and English, wey mean say dem need careful separation from historical West African Yorùbá .
As Yorùbá lexicography dey operate for digital and computational space, specific technical and linguistic challenge dem dey continue to affect di language .
As Natural Language Processing (NLP) researchers don show, including Irorun Orife et al. (2020) inside Improving Yorùbá Diacritic Restoration, majority of digitized Yorùbá text across internet completely leave out tone mark dem and subdot dem .
Because Yorùbá depend on pitch and vowel quality to distinguish words, di lack of diacritics dey cause serious homographic ambiguity . For example, di word owo wey dem write without diacritics fit stand for at least four different words wey no relate at all:
Natural language processing tools, search engines, and automated translation platforms dey often fail to process Yorùbá texts correctly because base digital vocabularies no get consistent diacritical restoration .
Di scholarly record show say even though several academic institutions and independent research groups don create semantic ontology prototypes and bilingual machine-translation wordlists, complete, standardized, and fully unified open-access Yorùbá WordNet (wey match Princeton WordNet for English) never ready . To develop dat kind semantic network still remain one of di major uncompleted work for modern Yorùbá computational lexicography.
| Year | Author / Compiler | Title | Publisher / Institution | Key Structural Features | Tone Notation System |
|---|---|---|---|---|---|
| 1843 | Samuel Àjàyí Crowther | Vocabulary of the Yoruba Language | Church Missionary Society (London) | First vocabulary wey native person compile; e introduce proverbs inside entries . | Rudimentary / non-systematic . |
| 1852 | Samuel Àjàyí Crowther | A Vocabulary of the Yoruba Language | Seeleys / CMS (London) | Orthographic milestone; e establish subdots (ẹ, ọ, ṣ) and grammatical elements . | High (´) and Low (`) dey marked; Mid unmarked . |
| 1858 | T. J. Bowen | Grammar and Dictionary of the Yoruba Language | Smithsonian Institution (Washington, D.C.) | Early comprehensive survey wey dem publish for academic research inside di Americas . | Systematic basic marking . |
| 1913 | Church Missionary Society (CMS) | A Dictionary of the Yoruba Language | CMS Bookshop (Lagos) / Oxford University Press | Standard bilingual pedagogical reference for schools and colonial civil service . | High and Low marked; frequent omissions . |
| 1958 | Roy Clive Abraham | Dictionary of Modern Yoruba | University of London Press | Exhaustive phonetic, morphosyntactic, and encyclopedic/cultural entries . | Comprehensive: all syllables, contour tones, and mid pitches fully marked . |
| 1958 | Isaac Oluwole Delanọ | Atúmọ̀ Èdè Yorùbá | Oxford University Press (London) | E pioneer monolingual Yorùbá definitions (Yorùbá-Yorùbá-English format) . | High and Low marked; standard pedagogical . |
| 1969 | Isaac Oluwole Delanọ | A Dictionary of Yoruba Monosyllabic Verbs | Institute of African Studies, University of Ife | E focus strictly on verb root behavior, syntactic valency, and complementation . | Fully marked on verb roots and examples . |
| 2008 | Yíwọlá Awóyalé | Global Yorùbá Lexical Database v. 1.0 | Linguistic Data Consortium, Univ. of Pennsylvania | Pass 368,000 computational entries; morphemic parsing; Atlantic diaspora coverage . | Fully encoded three-level computational tonology . |
| 2015– | Kọ́lá Túbọ̀sún et al. | YorubaName.com / YorubaWord.com | Yorùbá Name Project | Open-access multimedia database; etymological parsing; audio recordings / TTS . | Fully marked tone diacritics paired with audio . |
The CMS mission and Crowther, how and why Yoruba people converted to two world religions, the British annexation, indirect rule and what it did to kingship, and the making of "Yoruba" as one identity.
A scholarly analysis of major reference grammars, foundational syntactic debates, and dialectological classifications in Yorùbá linguistics.
How Yorùbá came to be written in Arabic and then Roman script, who decided the spelling rules, and why the subdots and tone marks are not optional.
A comprehensive scholarly survey of documented Yorùbá lexical databases, parallel machine translation corpora, named entity benchmarks, and speech datasets.
A comprehensive scholarly survey of Yorùbá computational linguistics, covering automatic diacritic restoration, machine translation, speech recognition, speech synthesis, and community benchmarks.
An analysis of Yorùbá language pedagogy, instructional materials, second-language tone acquisition research, and mother-tongue education policy in Nigeria.