About AfriSpeech
We're a community of researchers and engineers passionate about bringing African languages into the AI era.
Our Mission
There are over 2,000 African languages, yet very few have been included in modern NLP research and tooling. AfriSpeech exists to change that. We build open-source models, datasets, and tools that make it easy to work with African languages in NLP.
We focus on data curation, model training, and creating easy-to-use tools. Our flagship project, Africa Corpus, is the largest open-source African language dataset on HuggingFace with data from 693+ languages.
What We've Built
Language Identification
A CPU-friendly language ID model covering 1,386 African languages. It runs on a phone or laptop and requires just 5 seconds of audio.
Corpus Builder
Tools to access monolingual and parallel data for 693 African languages from multiple sources.
G2P Conversion
Rule-based grapheme-to-phoneme conversion for 400+ African languages, converting text to native orthography or IPA.
8 Open Datasets
Curated speech and text datasets freely available on HuggingFace, covering hundreds of African languages.
Contact
Have questions or want to collaborate? Reach out to us: