About
A research community for Ethiopian languages
EthioNLP began at COLING 2018 in Santa Fe, among researchers from Addis Ababa University, the University of Hamburg and the University of Trento who proposed a shared community for Ethiopian language technology rather than separate individual efforts.
- 80+ languages in Ethiopia
- 6 with an open model or dataset from us
- 15 people listed
- 2018 founded
Why it exists
NLP work on Ethiopian languages has been carried out in many places without formal communication between the researchers involved. Groups often worked on the same problems, rebuilt similar corpora, and published in a literature that had not been collected in one place.
The community is organised around collaboration and training as well as output. Members co-author, review each other's work, supervise theses and run workshops. Datasets and models follow from that, and training a researcher to run their own experiments has value beyond a single release.
The same applies to anyone who contributes language data. A speaker, annotator or collector is a collaborator on the work, named in it where they wish to be, and paid fairly where a project has funding. The expectations are set out in the community standards.
Ethiopia is multinational and multilingual, with more than 80 languages. The catalogue lists 100 entries, following Glottolog, which also includes extinct, pidgin and signed varieties. These languages are under-represented in language technology. Of the four Ethiopian languages with enough indexed NLP papers to measure, 71.4% of the work is on Amharic alone. Many of the languages in the catalogue have no corpus, no baseline and no paper to cite.
How it works
Everything here is public and regenerated from public sources. Models and datasets come from the Hugging Face API, code from GitHub, publications from OpenAlex, Semantic Scholar, DBLP and the ACL Anthology, and the language catalogue from Glottolog and Wikidata. Profiles are not maintained by hand; a member supplies a few public identifiers once, and the site updates from them.
Two distinctions are enforced. Work built by this community is kept apart from Ethiopian-language work by others, and work built for an Ethiopian language is kept apart from massively multilingual work that lists one among many.
Members are listed only after they confirm it. A Hugging Face organisation has more people in it than this site lists; only the 15 who have agreed to appear are shown.
Part of AfricaNLP
EthioNLP takes part in the wider African NLP community, including Masakhane, the AfricaNLP workshop series and the Deep Learning Indaba. The progress dashboard tracks Ethiopian-language research against African-language research as a whole.
Taking part
There is no vetting step and no fee. Students are welcome; many of the open topics are the size of a first thesis rather than a longer research programme. See how to join, the open thesis topics, or how a lab registers.
What we work on
Six tracks, described in full on the research page.