Paid AI dictation tools have quietly become a real software category, and in 2026 the professional shortlist comes down to three names: Wispr Flow, superwhisper and Aqua Voice. The comparison content circulating about them is mostly unreliable — prices misread, architecture claims copied from one vendor’s competitor page to the next, accuracy percentages printed with no source at all. This guide rebuilds the comparison from vendor primary sources and marks which numbers are evidence and which are marketing.
The stakes are larger than a ten-dollar subscription suggests. Dictation sits between your microphone and everything you write — email, code, prompts, client documents — so the architecture question (does the audio leave the machine?) is a procurement question, not a preference. The category also has real money behind it now: Wispr, the company behind Wispr Flow, had raised $81M at a $700M post-money valuation by November 2025, per TechCrunch.
What follows: what each tool actually is, the cloud-versus-on-device split that decides most purchases, every listed price labeled by the surface it appears on, an honest read of the accuracy evidence (there is far less of it than the search results suggest), a full comparison matrix with an evidence-grade column of our own, and a verdict framework for writers, developers and privacy-bound teams.
- 01Architecture is the real decision axis, not accuracy.superwhisper is the only one of the three that transcribes on-device. Wispr Flow processes audio in the cloud at every tier, including Enterprise. Aqua Voice lists no on-device mode. If audio leaving the machine is a problem, the shortlist is one item long.
- 02None of the three publishes a head-to-head accuracy benchmark.Wispr Flow names no underlying model and no error rate on its product pages. superwhisper publishes no accuracy figure either. Aqua Voice publishes a proprietary benchmark plus one external leaderboard placement it reports itself. Every tidy percentage table in the search results is unsourced.
- 03superwhisper is the cheapest listed subscription; Wispr Flow has the largest free tier.superwhisper Pro is listed at $8.49/month, or $84.99/year billed annually. Wispr Flow's free tier allows 2,000 words a week on desktop — twice Aqua Voice's entire 1,000-word lifetime free allotment, every single week.
- 04The circulating superwhisper lifetime price increase is not real.The vendor's own pricing surface lists Lifetime at $249.99 and Pro at $8.49 a month. The inflated figure repeated across several review sites traces to a misreading of the monthly rate — a live example of aggregate content propagating an error primary sources never contained.
- 05Platform reach splits the field cleanly.Wispr Flow is the only one of the three on Android. superwhisper covers macOS, Windows 10/11 and iOS. Aqua Voice is Mac and iOS only. For a mixed-device team, platform coverage can end the evaluation before price or accuracy is discussed.
01 — The FieldThree tools making three different bets.
These three are not variations on one product. Each has picked a different constraint to optimise against, and each has accepted a different cost for it. Wispr Flow bought platform reach and team governance by staying entirely in the cloud. superwhisper bought privacy and configurability by shipping local models and accepting a heavier setup. Aqua Voice bought vocabulary accuracy for developers by going narrow — Mac and iOS only — with a proprietary model behind it.
Read the grid below as three answers to the same question, not as a ranking. The ranking comes in section 06, and it depends entirely on who is asking.
Wispr Flow
Cloud dictation with an AI rewriting layer, built out toward teams. Enterprise adds SOC 2 Type II, ISO 27001, a HIPAA-ready BAA, SSO/SAML and enforced Privacy Mode with zero data retention. A separate Notetaker product handles meeting transcription and speaker ID, and is Mac-only for now.
superwhisper
Runs local AI models on-device, with Apple Silicon recommended for the best offline performance, and optionally calls cloud models instead. Modes are user-configurable — from pure transcription to a screen-aware Super Mode — with custom system prompts, and Pro users can bring their own API keys.
Aqua Voice
Positioned explicitly at developers, with integration branding for Cursor, Windsurf, VS Code and Warp, and homepage copy about dictating code and prompting coding agents. Its proprietary Avalon model is the only one in this comparison with any outside leaderboard placement.
02 — ArchitectureCloud or on-device — the question that eliminates options fastest.
Before price, before accuracy: does the audio leave the machine? Answer that and the shortlist usually collapses on its own.
superwhisper is the only one that can answer no. It runs local AI models on-device — the vendor recommends Apple Silicon Macs for the best offline performance — and separately supports a list of selectable cloud language models when you want them, with bring-your-own-key support on Pro. That dual path is the product’s defining feature and the reason it survives procurement conversations the other two do not.
Wispr Flow cannot. It processes audio in the cloud at every tier, Enterprise included. What Enterprise buys is a governance wrapper around cloud processing — enforced Privacy Mode with zero data retention, SOC 2 Type II, ISO 27001, a HIPAA-ready BAA, SSO/SAML — not local inference. Wispr Flow does state that users can opt out of model training at any time, which is a meaningful control, but it is a policy control rather than an architectural one.
Aqua Voice lists no on-device mode. Its differentiator is that transcription is informed by screen and app context, which is a cloud-shaped design. Treat it as cloud unless the vendor publishes otherwise.
There is a fourth answer worth naming, because it is the one regulated teams keep landing on: if no vendor tier satisfies the audio-residency requirement, the honest next step is a self-hosted stack rather than a different subscription. We compare that field against these paid leaders in open-source voice dictation versus Wispr Flow and friends, and the batch-transcription side of the same problem — meeting recordings, archives, pipelines rather than live input — is covered in our guide to self-hosted Whisper transcription pipelines.
03 — Listed PricingWhat each one costs, labeled by the surface it is published on.
Every figure below is the price listed on the vendor’s own pricing page at the time of writing. Vendor list prices are not enterprise agreements, App Store prices, or regional prices — they are the published starting point and nothing more.
Wispr Flow
Per Wispr Flow’s pricing page: free at $0, with 2,000 words a week on desktop, 1,000 words a week on iPhone, and unlimited on Android. Pro at $15/user/month billed monthly, or $12/user/month billed annually — a 20% annual discount — which unlocks unlimited dictations, extended meeting retention, access to advanced AI thinking models, and centralised billing. Enterprise is custom-priced.
superwhisper
Per superwhisper.com: a permanent free tier that does not expire, covering voice-to-text in any app, meeting recording, 100-plus languages and unlimited use of the small models. Pro at $8.49/month, or $84.99/year billed annually, plus a one-time Lifetime licence at $249.99. A 40% student discount and a 30-day no-questions-asked refund apply to paid plans. Enterprise is custom-priced and SOC 2 Type II certified.
Aqua Voice
Per the Aqua Voice pricing FAQ: Starter is free with 1,000 lifetime words — a one-time allotment, not a weekly or monthly one. Pro at $10/month, or $8/month billed annually ($96/year) unlocks unlimited dictation and the Avalon model. Max at $30/month, or $24/month billed annually ($288/year) adds Realtime mode and voice commands. Team is $12/user/month billed annually; Enterprise is custom with SSO/SAML, SCIM and Zero Data Retention. Students get 70% off Pro and Max.
Annualised list cost per seat · vendor-published rates
Source: vendor pricing pages; annualised figures are our own arithmetic from the listed monthly ratesAnnualised, the spread is wider than the monthly headline suggests. Wispr Flow Pro on annual billing is $144 per seat per year; superwhisper Pro on annual billing is $84.99. That is a $59.01 per-seat annual gap — $1,475.25 across twenty-five seats, before anyone negotiates anything. Meanwhile Aqua Voice’s Team tier lands on exactly the same $144 per seat per year as Wispr Flow Pro, which makes platform coverage rather than price the tiebreaker between those two.
superwhisper’s lifetime licence is the one genuinely different commercial shape in the category. At $249.99 against $84.99 a year, it pays for itself just under the three-year mark — three years of annual billing costs $254.97. Whether that is a good trade is a judgement about how confident you are that a Mac-first utility is still maintained and still competitive in 2029, not a calculation. For an individual who has already made dictation a daily habit, it is the strongest value in this comparison. For a team, the lack of a per-seat middle tier between Pro and custom Enterprise is the constraint that matters more.
04 — Evidence GradeThe accuracy numbers that are real, and the ones that are not.
Here is the finding that should reframe every accuracy comparison you read about this category: none of these three vendors publishes a head-to-head accuracy benchmark against a named competitor, and two of the three publish no accuracy figure at all.
Wispr Flow names no underlying model and no error rate on its product pages. The one comparative number in circulation is Wispr-stated and reached the public through funding coverage: roughly a 10% error rate against 27% for OpenAI’s Whisper and 47% for Apple’s native transcription. No methodology accompanied it. Treat it as a vendor claim, because that is exactly what it is.
superwhisper publishes no accuracy figure either. Its marketing competes on architecture, modes and price rather than on a leaderboard number — which is arguably the more honest posture, though it leaves buyers with nothing to compare.
Aqua Voice is the only one with anything external. Its own materials claim 97.3% accuracy on “AISpeak” — Aqua’s proprietary in-house benchmark, not an industry standard, and not comparable to anything. Separately, and far more usefully, Aqua reports that its Avalon model debuted on the independent Hugging Face Open ASR Leaderboard in October 2025 at 6.24% average word error rate — sixth overall and first among proprietary systems, ahead of Whisper-large-v3, ElevenLabs Scribe v1 and Rev AI Fusion in that ranking. That placement is reported in Aqua’s own recap of the result, not re-pulled from a live leaderboard table here, so read it as an external placement the vendor reports rather than a figure we re-confirmed.
Apple Dictation vs Aqua Voice
9to5Mac's Ben Lovejoy read the same passage into the same Mac microphone twice. Apple's built-in dictation made 17 errors; Aqua Voice made one. A single passage, run once, by one journalist — not a statistically robust benchmark, but a real outside test.
Avalon average WER
Aqua reports that Avalon entered the independent Open ASR Leaderboard at 6.24% average word error rate — sixth overall, first among proprietary systems. Reported via Aqua's own recap rather than a live re-pull, so weigh it accordingly.
Between these three tools
Neither Wispr Flow nor superwhisper publishes an accuracy figure, and no vendor in this comparison benchmarks against a named competitor. Every tidy percentage table ranking all three is, as far as we can establish, unsourced.
The 9to5Mac comparison is worth understanding precisely, because it is the most concrete accuracy data point anywhere in this category and it is also easy to over-read. Ben Lovejoy read Steve Jobs’ Stanford commencement address into the same Mac microphone twice, once through Apple’s built-in dictation and once through Aqua Voice. Apple’s pass produced 17 errors; Aqua’s produced one. It is one passage, one run, one reviewer — so do not read it as a benchmark. What it does show cleanly is formatting behaviour, which is arguably the more useful signal for daily writing anyway.
“You don't have to specify punctuation, paragraph breaks, or other formatting. Aqua Voice does it all intelligently and automatically.”— Ben Lovejoy, 9to5Mac, April 17, 2026
In an April 2026 follow-up, Lovejoy reported having dictated 352,000 words over 105 days on Mac with Aqua Voice — roughly 3,350 words a day, sustained — and reviewed the iPhone release positively, while noting that Apple’s on-device security model forces a brief app-switch on first use, friction he flagged as potentially disqualifying for some disabled users. Long-run usage documented by a named journalist is a materially better signal than an unsourced percentage, and it is worth more than any of the vendor accuracy claims in this section.
Our reading of the vacuum: the absence of head-to-head benchmarks is not an oversight, it is a rational choice. Publishing a comparative word-error-rate invites a rebuttal benchmark, a methodology argument, and a permanent obligation to re-run the test on every model update. Two of the three vendors have concluded the reputational downside outweighs the marketing upside. Buyers pay for that decision by having nothing to compare — which is precisely why your own two-week trial (section 08) is worth more than any table, including this one.
05 — Decision MatrixThe full comparison, with an evidence column.
Most comparison tables for this category print an accuracy row as if all three numbers came from the same test. None of them do. The table below keeps commercial terms, architecture and evidence in separate blocks, and adds the row nobody else publishes: a grade for how much each vendor’s accuracy claim is actually worth.
Prices are vendor-listed at the time of writing. Annualised figures are our arithmetic from the listed monthly rates. The grade in the final row is our assessment, not a vendor position.
| Attribute | Wispr Flow | superwhisper | Aqua Voice |
|---|---|---|---|
| Commercial terms · vendor-listed | |||
| Entry price, billed monthly | $15/user/mo | $8.49/mo | $10/mo (Pro) |
| Best listed recurring rate | $12/user/mo billed annually — $144/yr | $84.99/yr — about $7.08/mo | $8/mo billed annually — $96/yr |
| One-time licence | Not listed | Lifetime, $249.99 | Not listed |
| Free tier shape | 2,000 words/week desktop · 1,000/week iPhone · unlimited Android | Permanent free tier; unlimited use of the small models | 1,000 lifetime words, one-time |
| Team and enterprise | Enterprise custom · SOC 2 Type II, ISO 27001, HIPAA-ready BAA, SSO/SAML, enforced Privacy Mode | Enterprise custom · SOC 2 Type II, model access controls, enterprise-hosted model options | Team $12/user/mo billed annually · Enterprise custom with SSO/SAML, SCIM, Zero Data Retention |
| Architecture and platform reach | |||
| Processing | Cloud only, at every tier | On-device capable (Apple Silicon recommended); cloud models optional | Cloud; no on-device mode listed |
| Platforms | Mac, Windows, iOS, Android | macOS, Windows 10/11, iOS | Mac, iOS |
| Underlying model disclosed | No model named | Selectable cloud models, plus bring-your-own API key on Pro | Proprietary — Avalon |
| Evidence grade for accuracy claims · our assessment | |||
| Vendor-stated figure | ~10% error rate vs 27% Whisper, 47% Apple — Wispr-stated via TechCrunch, no methodology published | None published | 97.3% on AISpeak — Aqua’s own proprietary benchmark |
| Outside evidence | None found | None found | Open ASR Leaderboard placement Aqua reports (6.24% avg WER, #6 overall, #1 proprietary) · 9to5Mac side-by-side vs Apple Dictation |
| Our grade | Claimed, unverifiable | Not claimed | Partly external |
Read down the evidence block and the category’s marketing posture becomes legible. The vendor with the strongest funding and the widest distribution has the weakest published evidence. The cheapest tool makes no accuracy claim at all. The narrowest tool — Mac and iOS, developer-focused — is the only one that has submitted anything to an outside scoreboard. That inverse relationship between market position and evidentiary rigour is common in young software categories, and it is the single best reason to distrust every ranking article you find on this topic, this one included, in favour of your own trial.
06 — Who Picks WhatMatch the tool to the workflow, not to a leaderboard.
Because there is no credible accuracy ranking to lean on, the verdict has to come from architecture, platform coverage and what you actually dictate. Four buyer profiles cover most real situations.
Wispr Flow
The only one of the three on Android, the only one with a meaningful weekly free allowance to trial with, and the most developed enterprise wrapper — SOC 2 Type II, ISO 27001, HIPAA-ready BAA, SSO/SAML, enforced Privacy Mode. Accept that audio is processed in the cloud at every tier; if that is fine, this is the least friction across a heterogeneous team.
superwhisper
The only genuine on-device option, the cheapest listed subscription, the only lifetime licence, and the most configurable mode system — custom system prompts, screen-aware Super Mode, bring-your-own API key on Pro. Reviewers flag it as the steepest setup of the popular apps, and the Windows and iOS experiences trail the Mac app.
Aqua Voice
The only tool here with an outside leaderboard placement, and the only one whose positioning is explicitly about technical vocabulary — dictating code, refactoring functions, prompting coding agents, with integration branding for Cursor, Windsurf, VS Code and Warp. The constraint is reach: Mac and iOS only, with no Windows or Android client.
None of the three, as sold
If audio cannot leave your infrastructure and a vendor attestation is not sufficient, superwhisper’s local mode is the only paid path — and beyond that the answer is a self-hosted stack. We score that field against these paid leaders in open-source voice dictation versus Wispr Flow and friends.
One cross-cutting note for engineering teams: if the point of dictation is to drive an AI coding agent rather than to write prose, the tool matters less than the loop you build around it. We walk through that end to end in dictating straight into Claude Code. And if your interest is voice in a customer-facing product rather than voice as a personal input method, that is a different stack entirely — see voice AI for customer-facing products and, for the output side, the open-source text-to-speech side of the voice stack.
07 — Market SignalWhy dictation suddenly has venture money behind it.
A category that looked like a utility three years ago is now venture-funded, and that changes what buyers should expect from roadmaps, pricing and support.
Wispr raised $30M in June 2025 led by Menlo Ventures, then a further $25M in November 2025 led by Notable Capital — bringing total funding to $81M at a $700M post-money valuation, per TechCrunch. The same coverage reported that Wispr Flow was in use at more than 270 Fortune 500 companies and named Nvidia and Amazon among its enterprise customers. Those adoption figures are company-supplied and reported rather than independently audited, so weigh them as direction rather than measurement.
“We were still not planning to raise anytime soon because we had a really long runway and the team's really lean.”— Tanay Kothari, CEO of Wispr, speaking to TechCrunch, November 2025
Our read on where this goes: subscription prices in this category are unstable in one direction, and it is not upward. Speech recognition itself is commoditising — the open-weight Whisper family and its descendants keep improving and cost nothing per hour on your own hardware — so the durable margin is not in turning audio into text. It sits in the layer above: formatting intelligence, screen and app context, mode configuration, and the governance wrapper a security team will sign off on. That is exactly where all three of these vendors already compete.
Two consequences follow for buyers. First, expect the next round of differentiation to be about what the tool does with the text after it hears it, not about word error rate — which means the accuracy vacuum is likely to persist rather than resolve. Second, expect on-device processing to migrate from an enthusiast preference to a procurement checkbox as more regulated teams adopt dictation, which would make superwhisper’s current architectural position more valuable than its price suggests. Neither of those is a prediction with a date attached; both are reasons to re-evaluate this shortlist annually rather than treating a 2026 decision as settled.
08 — RolloutA two-week trial that produces a real answer.
Because there is no trustworthy public benchmark, the only reliable evaluation is a structured trial on your own vocabulary. The plan below is ours, not a vendor methodology, and it fits inside the free tiers for the first week — Wispr Flow’s 2,000 desktop words a week and superwhisper’s permanent free tier both carry a real trial. Aqua Voice’s 1,000 lifetime words will not, so budget a single paid month there if it is on your shortlist.
Constraint screening
Answer the disqualifying questions first: does audio have to stay on the machine, do you need Windows or Android, and does your security review need SOC 2 Type II or a BAA. In our experience this removes at least one candidate outright and often two, before anyone has evaluated output quality.
Vocabulary and formatting
Dictate an identical set of five passages through each survivor: one of ordinary prose, one dense with product and client names, one with code identifiers, one long-form to test paragraph breaks, and one in the second language your team actually uses. Count errors by hand. This is the entire evaluation.
Score on friction, not percentages
The right metric is not word error rate, it is how long you spend fixing output — including punctuation and paragraphing, which is where the biggest observable gap between tools shows up. Then re-check the pricing page, because list prices in this category move without announcement.
One rollout warning, and it belongs to Wispr Flow specifically rather than to the field: its user reviews report reliability complaints clustering after purchase rather than during the trial — dictation that works less consistently once it is load-bearing — and Windows freezes on some machines. superwhisper reviewers separately note that its Windows and iOS experience lags the Mac app. Review-aggregator sweeps also surface a notable divergence for Wispr Flow between a high mobile app-store score and a markedly lower Trustpilot score, though neither figure was pulled from a primary store API here, so we would not build a decision on either. The practical mitigation is to keep the trial running for a full billing cycle before rolling seats out to a team, and to prefer monthly billing for the first period even at the higher rate.
If dictation is one piece of a wider push to get AI tools into daily practice rather than a standalone purchase, our AI and digital transformation engagements start with exactly this kind of structured tool evaluation — scoped to your workflows, your compliance constraints and your existing stack.
09 — ConclusionBuy the architecture, not the accuracy claim.
There is no accuracy leaderboard here — so decide on architecture, reach and price.
The honest summary of this category in 2026 is that the most confident claims are the least supported. Wispr Flow has the broadest platform reach, the most developed enterprise wrapper and the biggest free tier, and it processes audio in the cloud at every tier without publishing a model name or an error rate. superwhisper is the cheapest listed subscription, the only genuine on-device option, and the only one offering a lifetime licence — and it makes no accuracy claim at all. Aqua Voice is the narrowest in reach and the only one with any outside evidence behind it.
That leaves architecture, platform coverage and price as the decidable axes, which is fine — they are the axes that actually determine whether a dictation tool survives contact with a real workflow. A tool that will not run on half your team’s devices is not a candidate at any accuracy. A tool that sends audio to a cloud your security review will not approve is not a candidate at any price. Work the eliminations first and the shortlist usually resolves itself.
The broader lesson is worth carrying beyond dictation. When a category’s comparison content converges on tidy percentages that no vendor actually published, the numbers are being manufactured somewhere in the aggregation chain — the same mechanism that turned a monthly price into a phantom lifetime increase across this category’s review sites. Re-fetch the primary source. Grade the evidence, not the claim. Then run your own two-week trial, because for a tool you will use every working hour, the only benchmark that matters is your own vocabulary.