Research

Research tracks

The tracks follow a sequence: text, then annotation, then a model, then applications built on it, then evaluation, alongside the training of people for each stage.

Datasets & corpora

Annotated corpora, parallel text, speech and benchmark suites. Released so far for Amharic, Afaan Oromo, Tigrinya, Somali and Ge'ez; in progress for Khimtagne, Awngi and other languages with few existing resources.

  • EthioSenti, EthioHate, EthioPOS
  • Amharic instruction-tuning data
  • Gender-bias MT benchmarks

Open resources in this track

Amharic Hate Speech - Dataset and classification Models

Jupyter Notebook
83 forksupdated Jul 2024

113 papers in datasets & corpora →

All 26 resources in datasets & corpora →

Language models

Pretrained encoders, multilingual generative models and tokenisers built for Ethiopic scripts and morphologically rich Semitic and Cushitic languages.

  • EthioLLM (s/b/l, 70K & 250K vocab)
  • Walia-LLM / Amharic LLaMA
  • Walia tokenizers, mT5 adaptations

Open resources in this track

By an EthioNLP member

text generationAmharic
552 downloads 0updated May 2025apache-2.0

By an EthioNLP member

text generationAmharic
343 downloads 1updated May 2025apache-2.0

Different semantic models for Amharic

Jupyter Notebook
2211 forksupdated Jan 2024

91 papers in language models →

All 40 resources in language models →

Applications

Applied systems: translation, search, content moderation, agricultural and legal question answering, and accessibility for Ethiopian Sign Language.

  • Machine translation systems
  • Hate-speech and misinformation tooling
  • Retrieval-augmented QA in Amharic

Open resources in this track

Natural Language Processing in Ethiopian Languages: Current State, Challenges, and Opportunities

africaamharicethiopiaethiopian-languages
178 forksupdated Jun 2025MIThomepage

124 papers in applications →

All 7 resources in applications →

Evaluation & benchmarks

Shared tasks, leaderboards and error analysis, so that claims about Ethiopian-language systems can be checked independently.

  • SemEval and AfricaNLP shared tasks
  • Bias and safety evaluation suites
  • Task-level error analyses

Open resources in this track

AmharicStoryQA a long-sequence story question answering benchmark grounded in culturally diverse narratives from Amharic-speaking regions

0updated Jan 2026CC0-1.0

88 papers in evaluation & benchmarks →

Capacity building

Mentoring MSc and PhD students, running workshops and tutorials with Ethiopian universities, and improving access to NLP research training within the country.

  • EthioNLP workshop at ICES22
  • Student mentoring and reading groups
  • University collaborations

Open resources in this track

collection of some interesting personal projects. Afaratlas , Amharic-NLP,General-NLP,Generative-Learning-algorithms,RL

Jupyter Notebookamharic-nlpcourseragdagenerative-learning-algorithms
1updated May 2020MIT

5 papers in capacity building →

Language & linguistics

Computational description of Ethiopian languages, morphology, dialect variation, historical reconstruction, and the documentation work the areas above depend on.

  • Morphological analysis for Ethio-Semitic
  • Dialect discovery in under-resourced varieties
  • Proto-language reconstruction

Open resources in this track

Dedicated repository to list all tech startups, companies, co-working spaces etc. in Ethiopia.

41 forkupdated Jun 2022MIT

24 papers in language & linguistics →

Why EthioNLP

Origins and goals

There are efforts throughout the world to conduct NLP for Ethiopian languages, but without formal communication among researchers. At COLING 2018 in Santa Fe, researchers from Addis Ababa University, the University of Hamburg and the University of Trento took the initiative to create a formal society for Ethiopian-language NLP research.

Our goals

  • Identify, prioritise and focus NLP research topics for Ethiopian languages.
  • Organise workshops, seminars and conferences for Ethiopic NLP research.
  • Join efforts among Ethiopian NLP researchers around the world.
  • Support and mentor students in NLP and data science research.
  • Coordinate resources for Ethiopian language research.
  • Collaborate with Ethiopian universities and assist education quality.