The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP

Sheriff Issaka1, Keyi Wang2, Yinka Ajibola3, Oluwatumininu Samuel-Ipaye3, Zhaoyi Zhang3, Nicte Aguillon Jimenez3, Evans Kofi Agyei4, Abraham Lin5, Rohan Ramachandran3, Sadick Abdul Mumin7, Faith Nchifor3, Mohammed Shuraim6, Erick Rosas Gonzalez1, Lieqi Liu1, Sylvester Kpei8, Jemimah Osei8, Carlene Ajeneza3, Persis Boateng9, Prisca Adwoa Dufie Yeboah10, Saadia Gabriel1

1University of California, Los Angeles  ·  2Georgia Institute of Technology  ·  3University of Wisconsin–Madison  ·  4University of Cape Coast  ·  5Carleton University  ·  6Stetson University  ·  7Northwestern University in Qatar  ·  8Cornell University  ·  9Soka University of America  ·  10Columbia University

ACL 2026 (Long Paper) Oral Presentation Best Paper Nomination

Abstract

Despite representing nearly one-third of the world's languages, African languages remain critically underserved by modern NLP technologies, with 88% classified as severely underrepresented or completely ignored in computational linguistics. We present the African Languages Lab (All Lab), a comprehensive research initiative that addresses this technological gap through systematic data collection, model development, and empirical analysis. Our contributions include: (1) a quality-controlled data collection pipeline, yielding the largest validated African multi-modal speech and text dataset spanning 40 languages with 19 billion text tokens and 12,628 hours of aligned speech data; (2) extensive experimental validation demonstrating that even modest-scale models, when fine-tuned on targeted language data, achieve substantial improvements over untrained baselines, averaging +23.69 ChrF++, +0.33 COMET, and +15.34 BLEU points across 31 evaluated languages; and (3) a comparative analysis against Google Translate in which a 1B-parameter model matched or surpassed the commercial system in several languages including Yoruba and Twi, revealing that data scarcity, rather than model scale, constitutes the primary bottleneck for low-resource NLP, and suggesting that systematic dataset development yields disproportionate returns for low-resource languages.

BibTeX

@inproceedings{issaka-etal-2026-african,
    title = "The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African {NLP}",
    author = "Issaka, Sheriff and Wang, Keyi and Ajibola, Yinka and Samuel-Ipaye, Oluwatumininu and Zhang, Zhaoyi and Jimenez, Nicte Aguillon and Agyei, Evans Kofi and Lin, Abraham and Ramachandran, Rohan and Mumin, Sadick Abdul and Nchifor, Faith and Shuraim, Mohammed and Liu, Lieqi and Gonzalez, Erick Rosas and Kpei, Sylvester and Osei, Jemimah and Ajeneza, Carlene and Boateng, Persis and Yeboah, Prisca Adwoa Dufie and Gabriel, Saadia",
    booktitle = "Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2026",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.acl-long.1965/",
}