Progress · 2010–2025
Research output for Ethiopian languages
Published NLP research on four Ethiopian languages, counted and set beside four widely-studied African languages measured by the same query against the same source, alongside a timeline of milestones. The totals are close, while the distribution across languages is uneven.
The comparison
Volume and growth against four African languages
Across the whole period the totals are close, 1208 papers against 1291, so the gap is not simply a continent-wide effect. The comparison group is growing faster: it grew 6.1× over the period against Ethiopia's 4.6×, and in 2025 the Ethiopian four accounted for 41.0% of the pair.
NLP papers per year
Indexed works classified under Natural Language Processing
Inside Ethiopia
Concentration on Amharic
71.4% of the indexed NLP research on these four languages is on Amharic. Tigrinya, Afaan Oromo and Somali share the remainder between them, and fourteen of the 100 catalogue entries have at least one indexed paper.
NLP papers per year, by Ethiopian language
Same source and method as above
| Language | 2010 | 2025 | Total since 2010 | Group |
|---|---|---|---|---|
| Amharic | 19 | 139 | 863 | Ethiopia |
| Tigrinya | 1 | 13 | 77 | Ethiopia |
| Afaan Oromo | 2 | 31 | 165 | Ethiopia |
| Somali | 3 | 14 | 103 | Ethiopia |
| Swahili | 16 | 93 | 515 | comparison |
| Yoruba | 3 | 71 | 309 | comparison |
| Hausa | 8 | 76 | 279 | comparison |
| Igbo | 3 | 43 | 188 | comparison |
Milestones
EthioNLP milestones, interleaved with the AfricaNLP milestones the community grew out of and contributes to. Each entry links to a source.
-
2025
First EthioNLP workshop, at ICES22
Twenty-one accepted contributions on Ethiopian-language NLP, co-located with the 22nd International Conference of Ethiopian Studies at Hawassa University.
-
2025
LT4All 2025 and deRSE 2025
"Advancing Language Technology for Ethiopia's Diverse Linguistic Landscape through the EthioNLP Collaborative Effort" presented in February 2025.
-
2025
EthioNLP on Voice of America
Community members discussed Ethiopian-language NLP for a general audience on VoA, in February 2025.
-
2024
Walia-LLM and Amharic instruction data
Task-specific and generative datasets combined into an Amharic LLaMA, with the dataset pipeline, models and evaluation outputs released openly.
-
2024
EthioLLM at LREC-COLING 2024
Multilingual language models for five Ethiopian languages, Amharic, Ge'ez, Afaan Oromo, Somali and Tigrinya, plus English, released with a benchmark suite and task evaluation.
-
2023
AfriSenti and SemEval-2023 Task 12
Sentiment analysis data for fourteen African languages, Amharic, Tigrinya and Oromo among them, run as a shared task with wide participation.
-
2023
Survey, NLP in Ethiopian languages
A community survey of the state, challenges and opportunities of NLP for Amharic, Afaan Oromo, Tigrinya and Wolaytta, with a public repository of existing resources.
-
2022
NLLB-200 and FLORES-200
Machine translation and evaluation data extended to 200 languages, including Amharic, Tigrinya, Afaan Oromo and Somali.
-
2021
MasakhaNER
Named-entity recognition data and models for ten African languages, built by more than fifty researchers, an example of participatory dataset work that Ethiopian projects have followed since.
-
2020
Lacuna Fund launches
Dedicated funding for labelled datasets in under-served languages, which funded part of the African-language data released since.
-
2020
First AfricaNLP workshop
The workshop series that gave African-language NLP a recurring venue at a major conference, beginning at ICLR 2020, held online after the Addis Ababa edition moved.
-
2019
Masakhane begins
An open, participatory research effort for African-language NLP, run by African researchers. EthioNLP members have contributed since its early years.
-
2018
EthioNLP founded at COLING 2018
At COLING in Santa Fe, researchers from Addis Ababa University, the University of Hamburg and the University of Trento agreed to build a formal community for Ethiopian-language NLP.
-
2017
First Deep Learning Indaba
A pan-African machine-learning gathering that became a regular meeting point for the field, and where many EthioNLP members first met.
Add a milestone by appending to
_data/milestones.yml.
Entries need a year, a track, a one-sentence description and a source URL.
How these numbers are produced
OpenAlex works matching a title/abstract search for the language name, intersected with the Natural language processing concept (C204321447). Counts are of indexed works, not of a curated bibliography. OpenAlex's coverage of the last year or two of any given period is uneven, so read the right-hand end of every line as a lower bound. The comparison between the lines is unaffected: both are counted the same way.
The intersection is what makes the count meaningful: an unfiltered search for Amharic also returns clinical studies of Amharic-speaking patients and sociolinguistic work, which would roughly double the apparent figure without adding any NLP papers.
Three limitations:
- A paper is counted when the language name appears in its title or abstract. Work that covers Ethiopian languages inside a hundred-language multilingual benchmark, without naming them, is missed.
- OpenAlex indexes the ACL Anthology well but not perfectly, and local Ethiopian venues and theses are largely absent, so the true figure is higher than the one plotted, in both directions of the comparison.
- The running year is excluded: it is incomplete, and OpenAlex indexes in bursts, so including it produces a spike that is a collection artefact rather than a finding.
These figures come from one script,
scripts/sync_progress.py,
and re-running it reproduces this page.
Source: OpenAlex. Generated 10 September 2026.