Tuned for better dictation.
The results are here.
Better writing takes careful work. Here’s how our local polish models performed, what we graded, and the decisions behind the results.
The details make
the difference.
Explore the strengths and tradeoffs of each model. This is our archived English text-polish comparison, not a test of complete dictation apps.
Complete English results by category
Pass rate includes pass + minor. Serious errors are in parentheses. Bold marks the highest pass rate, including ties.
| Category | Cases | EG-1Pass (serious) | S1-miniPass (serious) | Fluid-1Pass (serious) |
|---|---|---|---|---|
| Overall | 1,462 | 90.3% (66) | 86.5% (64) | 82.9% (81) |
| Clean speech | 309 | 96.1% (8) | 93.9% (14) | 93.5% (14) |
| Self-correction | 219 | 77.6% (22) | 59.8% (24) | 51.1% (21) |
| Fillers only | 200 | 91.0% (9) | 95.0% (3) | 94.0% (8) |
| Voice at risk | 134 | 91.8% (6) | 92.5% (6) | 86.6% (9) |
| Unfinished thought | 120 | 91.7% (3) | 91.7% (3) | 93.3% (1) |
| Spoken list | 114 | 85.1% (7) | 82.5% (9) | 54.4% (4) |
| Topic shift | 102 | 94.1% (5) | 93.1% (2) | 94.1% (2) |
| Inline enumeration | 87 | 92.0% (0) | 72.4% (1) | 97.7% (1) |
| Connected prose | 74 | 94.6% (3) | 98.6% (0) | 95.9% (2) |
| Numbers and dates | 73 | 97.3% (2) | 100.0% (0) | 93.2% (5) |
| Quoted instruction | 30 | 80.0% (1) | 70.0% (2) | 43.3% (14) |
Serious errors changed meaning or dropped content. Lower is better. All models were scored on the same cases.
One request at a time.
Local text-polish timing, including the runtime. Lower is faster.
EG-1 v2
312 msMedian606 ms at p95
S1-mini
87 msMedian176 ms at p95
Fluid-1
267 msMedian1,746 ms at p95
200 cases per model. Text-polish time only, not recording-stop-to-paste latency. Runtime differences limit the comparison.
EG-1 v2 · S1-mini in lists mode · Fluid-1 with reasoning
Built around the way
people actually speak.
Corrections. Unfinished thoughts. Spoken lists. The test checks more than whether a sentence looks tidy.
Define good writing.
Write the policy for meaning, voice and structure before creating answer keys.
Bring speech into it.
Speak realistic scenarios with TTS, then transcribe them with a real recognizer.
Check the answer keys.
Two authors work independently. Resolve disagreements and exclude undecidable cases.
Grade consistently.
Use the same judge and rubric for every model, with an adjudication pass.
Explore the benchmark methodology +
The exam was sealed August 15, 2026. It contains 1,462 English cases across 11 categories; 11 undecidable cases were excluded. Initial answer-author agreement was 77.94%.
The August 26 run used an Apple M5 Max with 64 GB memory and macOS 26. GPT-5.6-luna graded the model outputs with the same rubric. Pass rates include pass and minor; serious errors are reported separately.
Envious Labs builds EG-1 and selected this test. Model grading has uncertainty. Fluid-1’s production instructions and custom decoder were not replicated, so this is not its official app performance.
Read the benchmark details ↗More care for
what you meant.
We tune for the hard parts of spoken writing: keeping your voice, resolving corrections, and creating structure only where it belongs. We test those changes against the same written standards.
The thought you meant.
With AI polishLet’s meet Thursday, actually Friday.
Let’s meet Friday.


Try it with your own words.
Free, private dictation for Mac. No account. No subscription.
Download for MacmacOS 14+ · Apple Silicon