Every African language deserves representation in AI
Open-source NLP models, datasets, and tools built for African languages. Try our demos right here, or dive into our research.
What we do
Building NLP for Africa
From speech recognition to phoneme conversion, we create tools that make it easy to work with African languages in NLP.
Language Identification
A CPU-friendly language ID model covering 1,386 African languages. It runs on a phone or laptop and requires just 5 seconds of audio.
Corpus Builder
Tools to access monolingual and parallel data for 693 African languages from multiple sources.
G2P Conversion
Rule-based grapheme-to-phoneme conversion for 400+ African languages, converting text to native orthography or IPA.
8 Open Datasets
Curated speech and text datasets freely available on HuggingFace, covering hundreds of African languages.
Live demo
Try It Now
Identify which African language is being spoken in an audio clip. Uses a fast version of Omnilingual ASR to transcribe speech, then a lightweight classifier names the language. Runs entirely on CPU.
African Speech ID
1,386 African languages
African Speech ID
Pick a sample audio below to see what the model makes of it.
Five seconds minimum, ten or more works best.
Open data
Open Datasets
Curated speech and text datasets covering hundreds of African languages, freely available on HuggingFace.
Open Bible Speech African
553K rows · 1.09K likes · Bible audio across African languages
View Dataset →Open source
Our Projects
Tools built by the community, for the community. All code is open-source on GitHub.
africa-corpus-builder
14 starsGet access to monolingual and parallel data for 693 African languages
View Repo →afrispeech-selector
14 starsEasy access to speech data across 142 African languages for TTS and ASR
View Repo →Get Involved
We welcome contributions from developers, linguists, and researchers working on African language technology.