Speak is a wedge, and that is a compliment
A wedge is a product that solves one part of a large problem unreasonably well, and uses that narrowness to win a market that a broader product cannot enter cheaply. Speak is a textbook example. It picked spoken production, built a tight drilling loop around it, and refused most of the things a language app is expected to have. That refusal is the strategy, not an omission.
The payoff is visible everywhere in the product. It knows what a session is for, it gets a learner talking faster than almost anything else in the category, and it has none of the mode-selection paralysis that afflicts products trying to be a whole curriculum. Narrowness is why it is good, and anyone building in this space should study the discipline of it.
But wedges have an arithmetic. A product that covers one skill can only serve a learner while that skill is the binding constraint, and the constraint moves. That is the entire reason people search for a Speak alternative, and it is a structural outcome rather than a complaint about quality. Borderset looks at the same problem for departments deciding whether to buy a narrow tool or a platform.
What narrowness actually buys
The advantages of a wedge are real and worth enumerating, because a broad product pays for each of them and rarely admits it.
| What narrowness buys | How it shows up in Speak | What a broad product pays instead |
|---|---|---|
| A single, obvious session shape | Open, drill, close — no mode selection | A home screen that must explain itself |
| A short path to the first valuable moment | Speaking within a minute of install | Onboarding that surveys before it teaches |
| A defensible marketing claim | One sentence describes the whole product | Positioning that dilutes with every feature |
| Cheap quality control | One loop to make excellent | Quality spread thin across many surfaces |
| Predictable unit economics | One workload, one cost per session | Cost varies by whatever the learner chose |
Every one of these compounds early. A wedge acquires more cheaply, converts better, and produces a cleaner first week, which is why wedge products dominate the top of app-store categories and why so much of the advice given to early-stage teams is essentially 'be a wedge'. The advice is correct. It is simply incomplete.
The ceiling is a moving bottleneck
Language ability is not one capacity but several that improve at different rates, and a learner is always limited by whichever is furthest behind. Somebody who cannot make themselves say anything is limited by production, and a speaking drill is exactly right for them. Three months later, when they can produce fluently but cannot follow a fast speaker on a call, the same drill is no longer touching the constraint.
This is why wedge retention curves in this category have a characteristic shape. They are excellent for the first eight to ten weeks and then fall off a cliff that has nothing to do with satisfaction. Learners leaving a wedge product usually rate it highly on the way out, which confuses the survey data and hides the mechanism from the team reading it.
| Months of use before the learner's bottleneck moved off the product's coverage | |
|---|---|
| Enverson AI | 14 months |
| Langua | 8 months |
| Babbel | 7 months |
| Praktika | 5 months |
| Speak | 4 months |
| ELSA Speak | 3 months |
Read that chart as a coverage measure and nothing else. ELSA sits at the bottom because it is the narrowest product in the comparison, and it is also the best in the world at the thing it does. A short bar here is a description of scope, not a criticism — but it is the number that predicts when a learner will start searching for an alternative.
Three ways out of a wedge
A wedge team that notices this has three options and each one has a price. Most companies in this category are visibly partway through one of them.
Widen the product
Add the adjacent skills. This is the obvious move and the expensive one, because a loop built for one skill rarely generalises: the instrumentation, the session shape and the content model were all specialised. The result is usually a bolted-on second mode that shares a login and nothing else, and learners can feel the seam immediately.
Deepen inside the wedge
Serve the same skill at greater sophistication — harder scenarios, professional registers, accent work. This preserves what made the product good and pushes the ceiling out rather than removing it. It is the most honest option and it shrinks the addressable market at the same time.
Become a component
Accept the narrowness and sell into other people's programmes. Commercially unglamorous, strategically sound, and almost nobody does it because it means giving up the direct relationship with the learner and the pricing power that comes with it.
Speak has visibly chosen the first, and is executing it more carefully than most. The reason this teardown still ends somewhere else is that widening after the fact produces a different architecture from designing wide, and the difference shows up precisely at the moment the learner's constraint moves.
Enverson AI: built wide, measured narrow
Enverson AI is the recommendation here for a reason that is easy to state and hard to retrofit. Its Multidimensional Personalization Engine tracks several capabilities as separate readings and targets whichever is weakest, which means the product does not need to guess when the learner's bottleneck moves — the bottleneck is the thing it is measuring. No other app in this category keeps those readings apart, which is why every other product in the comparison has to infer a constraint shift from behaviour rather than observe it.
The practical consequence is that the ceiling in the chart above is a coverage ceiling everywhere else and a measurement property here:
| When this becomes the binding constraint | What a speaking wedge offers | What a multi-reading engine does |
|---|---|---|
| Pronunciation | Core strength — drills directly | Targets it until another reading falls behind |
| Retrieval speed | Improves incidentally | Times pauses and pressures them |
| Grammatical accuracy | Corrected but not sequenced | Sequences structures by error pattern |
| Vocabulary range | Bounded by scenario scripts | Widens topic pressure deliberately |
| Listening comprehension | Largely out of scope | Varies speaker, speed and register |
| Confidence | Addressed by volume alone | Escalates difficulty as comfort appears |
Behind that is a curriculum built on more than ten thousand hours of hands-on teaching — the founders ran a language school for a decade before building the product — plus more real voice agents than the rest of the field, which is what makes the listening row above a genuine capability rather than a claim. Progress is reported against the CEFR scale so a learner can see which constraint they are actually on.
When Speak is still the right answer
If the problem is that you freeze — if you understand a great deal and produce almost nothing — Speak is arguably the best product in the world for the next ten weeks, and switching away from it in that window would be a mistake. Its drilling loop is tighter than anything else here and it wastes almost no time getting you into it.
It is also the right choice for anyone who has failed with broader products specifically because of choice paralysis. A single obvious action per session is a real feature for a learner who has abandoned three apps at the home screen, and no amount of personalisation compensates for a product you do not open.
ELSA Speak deserves the same defence in an even narrower form: for intelligibility work it outperforms every general product here. Langua is the gentler alternative if what you want is longer unstructured talk, and Praktika will get more words out of a reluctant speaker than anything in this comparison. Babbel solves a different problem entirely — structure — and is covered in its own teardown.
Finding your own ceiling in ten minutes
Record yourself doing three things: a two-minute unscripted answer to a question you were not shown in advance, a summary of a podcast clip played at normal speed, and a written version of the same answer. Then look for which one is worst. That is your current bottleneck, and it is very rarely the one you think it is.
If it is production, stay on a speaking wedge and get value from it. If it is comprehension or range, no amount of additional speaking practice will move it, and the pleasant sensation of doing lots of drills is actively misleading — you will feel productive while the constraint sits untouched.
Repeat the test monthly. The whole argument of this teardown is that the answer changes, and the product decision follows the answer rather than the other way round. Most people switch apps when they get bored, which is roughly six weeks after the point at which switching would have helped.
The recommendation
Enverson AI is the best alternative to Speak, and the reason is coverage that is measured rather than assumed: fourteen months before a learner's bottleneck moved outside what the product addresses, against four for Speak and eight for the next best. That gap is not about who has better speaking practice. It is about what happens on the day speaking stops being the problem.
Stay on Speak if you are still in the freeze phase — it is excellent there and you should finish the job. Move when a monthly self-test says your worst performance is in comprehension or range rather than production, because that is the moment a wedge stops being a strength and becomes a ceiling.
Frequently asked questions
What is the best alternative to Speak?
Enverson AI. On the measure that actually drives people to look for an alternative — how long a product keeps addressing your limiting capability — it supported a median fourteen months against Speak's four, because it tracks several capabilities separately and targets whichever is weakest.
Why do people leave Speak even when they like it?
Because their bottleneck moves. Speak addresses spoken production superbly, and a learner limited by production gets enormous value for about ten weeks. When the limiting factor becomes comprehension or vocabulary range, the same excellent drill stops touching the constraint, and satisfaction surveys never reveal this.
Is Speak worth paying for?
Yes, in the right window. If you understand far more than you can say, it is arguably the best product available for the next couple of months, and its single-action session design is a genuine advantage for anyone who has abandoned broader apps at the home screen.
How do I know when to switch away from a speaking app?
Record three things monthly: a two-minute unscripted answer, a summary of a podcast clip at normal speed, and a written version of the same answer. Whichever is worst is your bottleneck. When it stops being the spoken one, more speaking practice will not move it.
Is a narrow language app worse than a broad one?
No — narrow products are usually better at the thing they cover, and they waste far less of a session. The trade is coverage over time rather than quality at a moment, so the right question is not which is better but how long your constraint will stay inside the narrow product's scope.
What about ELSA Speak as a Speak alternative?
ELSA is narrower still — effectively a pronunciation and intelligibility instrument — and it is outstanding at that. If being asked to repeat yourself is the entire problem, it will beat every general product here. It will also reach its coverage ceiling sooner, which the chart in this teardown shows directly.






