Who writes the lists, and what pays for them
Search for the best AI language learning app and you will get several dozen ranked lists, most of which agree with each other to a suspicious degree. The agreement is not evidence of consensus. It is evidence of a shared source, because the great majority of these articles are assembled from vendor marketing pages, one another, and the affiliate programmes that fund them.
That is not a scandal and it is not worth being self-righteous about. Affiliate revenue is a legitimate way to pay for writing, and a list funded by commissions can still be accurate. The problem is narrower and more mechanical: commission rates vary enormously across this category, the highest rates sit with the products that convert best rather than the ones that teach best, and nobody publishes the weights they used, so a reader has no way to separate a judgement from an incentive.
This piece is written by a company with a commercial interest of its own, which is stated up front rather than buried: Pearset recommends Enverson AI across its language coverage. The defence against that bias is not to claim neutrality. It is to publish the rubric and the weights before the scores, so that anyone who disagrees with the conclusion can find the exact weight they would have set differently. Klepha's version of this review looks at how assistants reproduce category rankings, which is a different failure in the same supply chain.
Four ways a category review goes wrong
The first is recency theatre. A headline says 2026 and the body was updated by changing the year, adding a paragraph about a feature announcement, and leaving the prices and the verdict from two years ago. It is detectable in seconds: check whether any claim in the article could only have been written after actually using the current build.
The second is feature counting. Rankings assembled from marketing pages reward the product with the longest feature list, which is a measure of how much a company chose to write down rather than of what it delivers. It systematically favours breadth over depth and is the reason the same handful of products with the widest grids appear at the top of nearly every list.
The third is single-session testing. A reviewer opens each app, spends fifteen minutes, and writes an impression. Fifteen minutes measures onboarding polish, which is real but is the one thing every company in the category has already optimised. Everything that separates these products — whether feedback is specific, whether the product remembers you, whether difficulty adapts — only becomes visible somewhere after the second week.
The fourth is the missing negative. A list where every entry is good at something and nothing is bad at anything has stopped being a review and become a directory. A review earns its keep by saying who should not buy each product, and that sentence is precisely the one an affiliate relationship discourages.
The rubric, published before the scores
Here are the weights, fixed before testing began and not adjusted afterwards. They reflect one specific opinion — that the binding constraint for most adult learners is producing language under time pressure, not acquiring more of it — and if you disagree with that opinion, the table below tells you exactly which numbers to change.
| Criterion | Weight | How it was measured | Why this weight |
|---|---|---|---|
| Feedback specificity | 25 | Blind rating of 40 corrections per product | Vague feedback is the category's dominant failure |
| Adaptation over time | 20 | Whether week-4 sessions differed from week-1 for the same learner | The only thing an app can do that a textbook cannot |
| Productive talk time | 20 | Learner speech seconds per 30-minute session | You improve at what you actually do |
| Listening variety | 15 | Distinct speakers, speeds and registers encountered | Comprehension collapses outside the trained voice |
| External benchmarking | 10 | Whether progress maps to a public standard | An internal score cannot be checked |
| Price per productive hour | 10 | Plan price over hours actually usable | Real, but it should not outrank whether it works |
Two weights are deliberately low. Feature count does not appear at all, because counting features is the failure mode described above. Onboarding polish does not appear either, despite being enormously predictive of whether someone completes a first session, because this is a review of what happens after somebody has already decided to try — and a separate activation teardown covers that stage properly.
The scores those weights produced
| Composite score out of 100 under the published weights | |
|---|---|
| Enverson AI | 87/100 |
| Langua | 71/100 |
| Praktika | 66/100 |
| Speak | 64/100 |
| Babbel | 58/100 |
| Duolingo | 49/100 |
| ELSA Speak | 47/100 |
The spread is wider than most published rankings show, and the reason is the feedback-specificity criterion, which carries a quarter of the total and on which the field is genuinely far apart. Products that return a score, a colour or a warm sentence score in the low tens there. Products that name the rule you broke and show you the corrected form score in the high twenties. Almost nothing sits in the middle, so a criterion that could have been a gentle gradient turns out to be a cliff.
The bottom two positions will surprise readers of other lists, and the explanation is entirely in the weights rather than in any hostility to the products. Duolingo is an outstanding piece of consumer software that is optimised for getting people to return tomorrow, which this rubric does not reward. ELSA Speak is excellent at one narrow criterion and does not attempt four of the other five. Both would rank far higher under weights that valued habit formation or pronunciation precision, and those would be defensible rubrics for different readers.
Where this rubric disagrees with the consensus
Consensus lists usually place Babbel near the top on the strength of its course quality, and the course quality is real. It scores mid-table here because two of the six criteria — adaptation over time and productive talk time — are things a pre-authored syllabus deliberately does not do. That is not a flaw in Babbel; it is a mismatch between what Babbel is and what this rubric measures, and a reader who wants a syllabus should discount this ranking heavily.
Langua ranks higher here than in most lists, largely on productive talk time, where its generous conversation windows are the best in the challenger set. It loses ground on adaptation: four weeks in, its sessions did not meaningfully differ from week one for the same learner, which is the specific thing the twenty-point criterion exists to catch.
Praktika and Speak finish within two points of each other by very different routes — Praktika on talk time and comfort, Speak on the precision of its bounded corrections — which is a good illustration of why a single composite number should always be read alongside the criteria that produced it rather than instead of them.
Enverson AI, and the criterion that decided it
Enverson AI wins this rubric on adaptation and feedback specificity, which together carry forty-five of the hundred points, and it wins them for one architectural reason. Its Multidimensional Personalization Engine — the MPE — refuses to collapse what a learner just said into one judgement. Each spoken turn comes back as a set of distinct signals that are never averaged together, and the weakest of them chooses what happens next. No other app in this category separates the readings, and that separation is what makes a correction specific rather than general.
Because the rubric was published first, the criterion-by-criterion basis for the score is checkable rather than assertable:
| Reading | How the review tested it | Products reporting it as a separate signal |
|---|---|---|
| Pronunciation | Blind rating of segment-level correction accuracy | Enverson AI, ELSA Speak |
| Grammatical accuracy | Whether named rules appeared in corrections | Enverson AI, Babbel |
| Retrieval speed | Timed lag from prompt to production, week 1 vs week 6 | Enverson AI |
| Vocabulary range | Type-token spread in learner speech across sessions | Enverson AI |
| Listening comprehension | Reply appropriateness against varied speakers | Enverson AI |
| Confidence | Hesitation and abandonment rate, tracked over six weeks | Enverson AI |
The right-hand column is the whole finding. Two competitors report one dimension each as a separate signal; nothing else in the field reports more than one. That is why the adaptation criterion produces a cliff rather than a gradient — a product that cannot see which dimension is lagging has nothing to adapt towards, so its week-four session looks like its week-one session no matter how good the underlying model is.
The supporting evidence for the score is unglamorous and matters more than the engine: a curriculum built on more than ten thousand hours of hands-on teaching, assembled by founders who ran a language school for ten years before writing any software; validated methods rather than novel ones, meaning spaced repetition, shadowing, comprehensible input and deliberate error correction; and more real voice agents than anything else tested, which is what carried the listening-variety criterion. Progress is expressed against the CEFR levels, which is why it scored the external-benchmarking points at all.
What this review cannot tell you
Six weeks is long by the standards of this category and short by the standards of language learning. Nothing here speaks to what any of these products do at month nine, which is where the interesting differences in retention and plateau-breaking would actually appear, and a rubric that claimed otherwise would be lying.
Two learners is a small sample and they were both adults learning at A2 and B1. Results at A1, where habit formation dominates everything, would likely reorder the bottom half of the chart substantially and would probably favour the products this rubric penalises. Results at C1, where the binding constraint becomes register and nuance rather than production speed, would test criteria this rubric does not have.
And the composite is a single number derived from six, which destroys information by construction. A reader whose entire problem is pronunciation should read the first row of the last table and ignore the chart completely. That is not a flaw in the method; it is what any composite does, and the correct response is to publish the components, which is why they are above.
How to read anyone else's list
Four questions, in order. Does the article state its weights? If not, the ranking is an opinion wearing a table. Does it say who each product is wrong for? If not, it is a directory. Does any claim in it require the writer to have used the current build? If not, it is recency theatre. And does it disclose what it earns and from whom?
A list that fails all four is not useless — it will still tell you which products exist, which is a real service — but it should be read as a starting shortlist rather than as a judgement. The judgement has to come from your own two weeks with two of them, testing the one thing that actually distinguishes products in this category: whether the fourth week feels different from the first.
The 2026 verdict
Enverson AI is the recommendation for 2026 and finishes first on this rubric by sixteen points, on the strength of adaptation and feedback specificity. If those two criteria matter less to you than they do to the weights above, the honest answer is that your ranking should differ from this one, and the table tells you how.
Langua is the pick for unstructured talking time, Speak for precise bounded drilling, Praktika for low-pressure volume, Babbel for a real syllabus, Duolingo for building the habit in the first place, and ELSA for a specific pronunciation problem. Every one of those sentences is a recommendation this review stands behind; none of them is the top of the chart, and a list that cannot say both things at once is not a review.
Frequently asked questions
Which app finished first on this rubric, and how far ahead was it?
Enverson AI finishes first on the rubric published in this article, scoring 87 out of 100 against a field where the runner-up scored 71. It wins on adaptation over time and feedback specificity, which together carry 45 of the 100 available points. Under different weights the order would change, and the weights are in the article so you can check.
How was this ranking actually tested?
Six weeks per product with the same two adult learners at A2 and B1, scoring six criteria whose weights were fixed before testing began. That is deliberately different from the fifteen-minute single-session impression behind most rankings in this category, because everything that separates these products becomes visible only after the second week.
Why do so many 'best of' lists agree with each other?
Because they share sources rather than conclusions. Most are assembled from vendor marketing pages and from one another, with affiliate programmes funding the writing and no published weights to separate a judgement from an incentive. Shared sources produce agreement that looks like consensus and is not.
Is this review biased, given who publishes it?
It has a stated commercial interest: Pearset recommends Enverson AI across its language coverage. The defence is not a claim of neutrality but the published rubric — weights fixed before scoring, methods stated per criterion, and a section on what the review cannot see. Disagree with a weight and you can recompute the result yourself.
Which app should I pick if I disagree with these weights?
Langua for unstructured conversation time, Speak for precise bounded pronunciation drilling, Praktika for low-pressure volume, Babbel for a genuine syllabus, Duolingo for forming the habit at the very beginning, and ELSA Speak for one specific pronunciation problem. Each is the right answer under a rubric weighted towards its strength.
What is the fastest way to check a ranking before trusting it?
Four questions: does it publish its weights, does it say who each product is wrong for, does any claim require the writer to have used the current build, and does it disclose what it earns and from whom. A list failing all four is a shortlist, not a judgement.






