Progress · 2010–2025

Research output for Ethiopian languages

Published NLP research on four Ethiopian languages, counted and set beside four widely-studied African languages measured by the same query against the same source, alongside a timeline of milestones. The totals are close, while the distribution across languages is uneven.

1208 NLP papers on Ethiopian languages since 2010
4.6× growth, first three years to last three
71.4% of that is Amharic alone
6 Ethiopian languages with an EthioNLP model or dataset of 100 in the catalogue

The comparison

Volume and growth against four African languages

Across the whole period the totals are close, 1208 papers against 1291, so the gap is not simply a continent-wide effect. The comparison group is growing faster: it grew 6.1× over the period against Ethiopia's 4.6×, and in 2025 the Ethiopian four accounted for 41.0% of the pair.

NLP papers per year

Indexed works classified under Natural Language Processing

Ethiopian languages (Amharic, Tigrinya, Afaan Oromo, Somali) Comparison: Swahili, Yoruba, Hausa, Igbo
0 50 100 150 200 250 300201020122014201620182020202220242025

Inside Ethiopia

Concentration on Amharic

71.4% of the indexed NLP research on these four languages is on Amharic. Tigrinya, Afaan Oromo and Somali share the remainder between them, and fourteen of the 100 catalogue entries have at least one indexed paper.

NLP papers per year, by Ethiopian language

Same source and method as above

Amharic Tigrinya Afaan Oromo Somali
0 50 100 150201020122014201620182020202220242025
Language 2010 2025 Total since 2010 Group
Amharic 19 139 863 Ethiopia
Tigrinya 1 13 77 Ethiopia
Afaan Oromo 2 31 165 Ethiopia
Somali 3 14 103 Ethiopia
Swahili 16 93 515 comparison
Yoruba 3 71 309 comparison
Hausa 8 76 279 comparison
Igbo 3 43 188 comparison

Milestones

EthioNLP milestones, interleaved with the AfricaNLP milestones the community grew out of and contributes to. Each entry links to a source.

Ethiopia Africa-wide
  1. 2025

    First EthioNLP workshop, at ICES22

    Twenty-one accepted contributions on Ethiopian-language NLP, co-located with the 22nd International Conference of Ethiopian Studies at Hawassa University.

  2. 2025

    LT4All 2025 and deRSE 2025

    "Advancing Language Technology for Ethiopia's Diverse Linguistic Landscape through the EthioNLP Collaborative Effort" presented in February 2025.

  3. 2025

    EthioNLP on Voice of America

    Community members discussed Ethiopian-language NLP for a general audience on VoA, in February 2025.

  4. 2024

    Walia-LLM and Amharic instruction data

    Task-specific and generative datasets combined into an Amharic LLaMA, with the dataset pipeline, models and evaluation outputs released openly.

  5. 2024

    EthioLLM at LREC-COLING 2024

    Multilingual language models for five Ethiopian languages, Amharic, Ge'ez, Afaan Oromo, Somali and Tigrinya, plus English, released with a benchmark suite and task evaluation.

  6. 2023

    AfriSenti and SemEval-2023 Task 12

    Sentiment analysis data for fourteen African languages, Amharic, Tigrinya and Oromo among them, run as a shared task with wide participation.

  7. 2023

    Survey, NLP in Ethiopian languages

    A community survey of the state, challenges and opportunities of NLP for Amharic, Afaan Oromo, Tigrinya and Wolaytta, with a public repository of existing resources.

  8. 2022

    NLLB-200 and FLORES-200

    Machine translation and evaluation data extended to 200 languages, including Amharic, Tigrinya, Afaan Oromo and Somali.

  9. 2021

    MasakhaNER

    Named-entity recognition data and models for ten African languages, built by more than fifty researchers, an example of participatory dataset work that Ethiopian projects have followed since.

  10. 2020

    Lacuna Fund launches

    Dedicated funding for labelled datasets in under-served languages, which funded part of the African-language data released since.

  11. 2020

    First AfricaNLP workshop

    The workshop series that gave African-language NLP a recurring venue at a major conference, beginning at ICLR 2020, held online after the Addis Ababa edition moved.

  12. 2019

    Masakhane begins

    An open, participatory research effort for African-language NLP, run by African researchers. EthioNLP members have contributed since its early years.

  13. 2018

    EthioNLP founded at COLING 2018

    At COLING in Santa Fe, researchers from Addis Ababa University, the University of Hamburg and the University of Trento agreed to build a formal community for Ethiopian-language NLP.

  14. 2017

    First Deep Learning Indaba

    A pan-African machine-learning gathering that became a regular meeting point for the field, and where many EthioNLP members first met.

Add a milestone by appending to _data/milestones.yml. Entries need a year, a track, a one-sentence description and a source URL.

How these numbers are produced

OpenAlex works matching a title/abstract search for the language name, intersected with the Natural language processing concept (C204321447). Counts are of indexed works, not of a curated bibliography. OpenAlex's coverage of the last year or two of any given period is uneven, so read the right-hand end of every line as a lower bound. The comparison between the lines is unaffected: both are counted the same way.

The intersection is what makes the count meaningful: an unfiltered search for Amharic also returns clinical studies of Amharic-speaking patients and sociolinguistic work, which would roughly double the apparent figure without adding any NLP papers.

Three limitations:

  • A paper is counted when the language name appears in its title or abstract. Work that covers Ethiopian languages inside a hundred-language multilingual benchmark, without naming them, is missed.
  • OpenAlex indexes the ACL Anthology well but not perfectly, and local Ethiopian venues and theses are largely absent, so the true figure is higher than the one plotted, in both directions of the comparison.
  • The running year is excluded: it is incomplete, and OpenAlex indexes in bursts, so including it produces a spike that is a collection artefact rather than a finding.

These figures come from one script, scripts/sync_progress.py, and re-running it reproduces this page.

Source: OpenAlex. Generated 10 September 2026.