Community

Teams & collaborators

A group registers its Hugging Face and GitHub organisations, not individual models. From then on, work it publishes that covers an Ethiopian language is listed on this site automatically.

Registering is what makes a group's work count as EthioNLP's rather than as related work.

Collaborators

Communities we work with

Peer communities and research groups. Their work is their own and is not counted in the totals on this site.

  • Masakhane

    pan africansince 2019

    Africa-wide

    An open, participatory research effort for African-language NLP, run by African researchers. One of the largest of the continent's language communities.

    Together: EthioNLP members are contributors and co-authors on Masakhane datasets and benchmarks, including MasakhaNER and the AfriSenti and AfriHate lines of work.

  • Bahir Dar, Ethiopia

    Research on information and communication technologies for development, including Amharic language resources and annotation work, at one of Ethiopia's largest universities.

    Together: Joint work on Amharic corpus construction and annotation, and a route for students in Ethiopia into the community's projects.

  • HausaNLP

    nationalsince 2022

    Nigeria

    A research community for Hausa and other Nigerian languages, and the group behind the Hausa language catalogue on which this site's own catalogue is modelled.

    Together: Shared work on African-language sentiment and hate-speech benchmarks, and a common approach to cataloguing what exists for a country's languages.

Groups working with us that are not listed can add an entry to _data/collaborators.yml, or open an issue.

Teams

Registered teams

4 labs and projects have registered their organisations. Work they publish that covers an Ethiopian language is listed here on the next nightly sync. Models are not submitted by hand.

  • EthioNLP

    community

    The community itself. Datasets, benchmarks and the EthioLLM model family for Ethiopian languages.

    19 models and datasets listed here · registered 2018-08

  • The lab behind AmRoBERTa, one of the first RoBERTa models pretrained from scratch for Amharic, and a body of Amharic hate-speech and annotation work.

    5 models and datasets listed here · registered 2018-08

  • AfriHate

    project

    A pan-African hate-speech benchmark, with Amharic, Tigrinya, Afaan Oromo and Somali among its languages.

    1 model and dataset listed here · registered 2024-01

  • EthioSpeech

    project

    Speech corpora and recognition models for Ethiopian languages.

    2 models and datasets listed here · registered 2024-01