New: High-level announcements are live
    Open
    Yapr journallearn igbo by speaking why most

    Learn Igbo by Speaking: Why Most Apps Get Igbo Wrong

    Igbo has consonants that need two closures in your mouth happening at once, not one after another. Type them out and they look like any other two-letter cluster. Say them wrong and you've said a different word.

    Yapr editorial7 min read
    Pixel art language studio with a glowing book, microphone, headphones, globe, and conversation bubbles
    Ideas become conversation
    Start reading
    In this article

    Igbo gets introduced to English speakers mostly through its tone system — plenty of guides warn you that pitch changes meaning, and that's true. But there's a consonant problem that shows up before tone even enters the picture, and it's one a transcript-matching app has no way to catch: Igbo has two consonants that aren't built the way any English sound is built, and the letters used to spell them look completely ordinary.

    The approach

    One Sound, Two Places in Your Mouth, at the Same Time

    English consonants are made in one place: your lips close for "b," the back of your tongue hits your soft palate for "g," and so on. Igbo has two consonants, written gb and kp, where both of those things happen at once: lips closing and the back of the tongue touching the velum, simultaneously, as a single sound rather than two sounds said back to back. Linguists call these labial-velar consonants, and they're rare enough globally that most English speakers have never had to produce one.

    The instinct when you see "gb" or "kp" written down is to read it the way English reads consonant clusters: say the g, then say the b, one after another, the way you'd read "gb" in an invented English word. That's not what the spelling is asking for. Àgbà, "jaw," and àkpà, "bag," are a real minimal pair in Igbo. The two words differ only in whether the double closure is voiced (gb) or voiceless (kp). Say either one as two separate consonants strung together instead of one simultaneous closure, and you've produced a sound that isn't quite either intended word. It's close enough that a transcript-based system may still guess the right spelling from context, but not the sound an Igbo speaker actually made or expects to hear back. The word "Igbo" itself contains this same consonant: phonetically [iɡ͡boː], the gb pronounced as one simultaneous closure, not "ig" followed by "bo."

    An app that shows you "agba" and has you read it aloud is testing whether you can find those letters on a page and produce some approximation of a "g" and a "b" near each other. It isn't testing whether you closed your lips and the back of your tongue at the same instant, the one thing that actually makes the sound correct. Text-to-speech-trained recognition, built overwhelmingly on languages where every consonant has one place of articulation, doesn't have a clean acoustic model for a doubly-articulated stop at all.


    The approach

    Pitch That Shifts as the Sentence Goes On

    Igbo tone doesn't just mark individual syllables high or low and stop there. When a low-tone syllable sits between two high-tone syllables, the second high tone doesn't return to the pitch of the first — it steps down to a new, slightly lower ceiling, and every following high tone in the phrase measures itself against that lower ceiling instead of the original one. Linguists call this downstep, and a 2026 Speech Prosody paper (Nwosu, "Automatic and Non-Automatic Downstep in Igbo Are Not Realized the Same") found that Igbo actually has two distinct kinds of it (one triggered automatically by an audible low tone in between, and one that shows up with no low tone pronounced at all), and that the two aren't acoustically identical, which is exactly the kind of fine-grained pitch behavior a speech system needs dedicated training to handle rather than a generic tone-detection model.

    The practical result is that a High tone late in an Igbo sentence isn't a fixed pitch you can learn once and reuse. It's relative to whatever downstep has already happened earlier in the same phrase, and it keeps ratcheting downward each time the trigger condition occurs. A learner (or a recognition system) expecting every "High tone" to sound the same regardless of position will misjudge the later ones as too low, when the speaker producing them is doing exactly what a fluent Igbo phrase requires.


    Pixel art conversation partners practicing together in a warm evening cafe

    Practice that feels human

    Real fluency grows in the pauses, reactions, and small moments that make a conversation feel alive.

    The approach

    What Igbo Learners Actually Need

    Feedback on whether a doubly-articulated stop was actually produced as one closure, not credit for a transcript that happens to spell "gb" or "kp" correctly. Àgbà and àkpà are different words, not the same word said two ways. That distinction is worth catching in the moment.

    Listening practice built around real connected speech, where downstep's cumulative, phrase-relative pitch-lowering actually shows up, rather than isolated single-word tone drills that never let a High tone sit after a downstep trigger.

    Real conversation, not read-aloud drills. Both the gb/kp contrast and downstep are about what you actually produce and hear in the flow of speech, not about whether you can decode Igbo spelling into some approximation of the right sounds.

    Room for the fact that Igbo tone behavior compounds across a sentence, instead of treating every High-marked syllable as an identical, isolated pitch target.


    The approach

    How Yapr Handles Igbo

    Audio-level pronunciation feedback. Yapr's speech-to-speech pipeline evaluates the sound you actually produced, so a two-beat "g-then-b" standing in for Igbo's single simultaneous gb closure is something it can catch, instead of accepting whatever word your transcript looks closest to.

    Real conversation, not read-aloud drills. The gb/kp contrast and downstep's phrase-level pitch shifts both need to be produced and recognized in live, connected speech — Yapr's conversation practice puts you in that context instead of isolated flashcard reps where the hard part gets skipped by just seeing the word on screen.

    Whisper mode. Practice the double closure and the pitch shifts of downstep quietly, without an audience, while still getting real feedback on what you said.

    $12.99/month, 47 languages. One subscription covers Igbo and anything else you're learning, against the cost of finding and scheduling an Igbo-speaking tutor.


    Yapr processes what you actually said in Igbo, including the simultaneous gb/kp closure and downstep's shifting pitch ceiling, instead of just checking whether your transcript matches the text on the screen. 47 languages at yapr.ca.


    Quick answers

    Frequently asked questions

    Turn reading into speaking

    Yapr processes what you actually said in Igbo, including the simultaneous gb/kp closure and downstep's shifting pitch ceiling, instead of just checking whether your transcript matches the text on the screen.

    47 languages at yapr.ca.

    Practice privately, at your pace, with zero judgment.

    Pixel art learner practicing aloud with headphones and a phone in a cozy room at night