Three layers, three positions. The heritage material is free and stays free; the software is not.
The corpus text is free to individuals, permanently.
No paywall on the heritage material. No account required to read. Institutions are asked to pay for institutional access; individuals are not, ever.
Everything the project charges for sits outside that line: software, licensing, convenience features, teaching, and commissioned work. The date this commitment is formally published, and will not be moved from, goes here once the entity that makes it exists.
Prose, structured data and software are licensed separately, because they are three different kinds of work with three different reuse cases.
| Layer | Licence | Effect |
|---|---|---|
| Corpus prose | CC BY-SA 4.0 | Free to read, reuse and adapt with attribution; derivatives stay open. |
| Structured data | CC BY 4.0 | Free for research and NLP use, including African language technology work. |
| Software platform | Proprietary | Licensed separately, and it is the project's main commercial line. |
Intended, not yet granted. A licence is granted by publishing it with the work, and the corpus has not yet been published under these terms. The choices above are the project's settled intent and match the practice of comparable corpora; they become binding when the files ship carrying them.
These sit alongside the licences as a statement of intent rather than as licence terms.
Ifá verse content is not licensed for generative AI training
The foreseeable downstream harm outweighs any revenue, and the decision is not a close one.
Attribution must survive
Reuse that strips the citations reproduces exactly the problem this project exists to fix, since the citations are the substance and not decoration.
The pattern is consistent: share-alike on prose, a looser licence or an outright waiver on structured data. This project follows it rather than inventing a position.
Wikimedia
Wikidata's structured claims and statements are released under CC0. Wikipedia's article prose is CC BY-SA.
Europeana
All metadata across Europeana is CC0. Prose in the blogs and exhibitions is CC BY-SA, and the heritage media objects carry their own individual rights statements.
Universal Dependencies
New treebank repositories default to CC BY-SA 4.0 for structured syntactic annotation, with providers free to use CC BY 4.0 instead.
CLARIN
CLARIN identifies CC BY 4.0 and CC BY-SA 4.0 as the primary open licences for text corpora, covering both copyright and database rights, and names CC0 as a waiver tool rather than a licence.
Academic use depends on a citation that resolves to a specific passage in a specific version, and stays resolvable.
Ìpilẹ̀ṣẹ̀, "<document title>", https://ipilese.com/read/<section>/<slug>, accessed <date>.
Not in place. There is no DOI, no versioned corpus release, and no snapshot deposited in a repository. A citation to this work is currently a URL plus an access date, which is weaker than what a scholarly resource should offer and is named as weaker rather than dressed up.
The practice this is aiming at is the one Perseus and DraCor established: a dereferenceable identifier that resolves to a passage within a named version, and a versioned corpus snapshot deposited with a DOI so a computational analysis can cite the exact release it ran against.
One REST API is the source of truth for every surface, and it is the same one anyone else would consume.
Documents, sections and their metadata are served over a versioned JSON API. Responses carry the version they were served under, so a client can assert the version it actually reached rather than the one it thinks it requested. The implemented endpoints start at ipilese.com/api/v1/documents