Journal

A smaller speech model earns its place on the iPhone

Aidana transcribed the 47 matched clips 3.35 times faster, with observed English WER of 2.15% and Spanish WER of 5.37%.

Vlad / experimentos.
The boundary

This is short public read speech on one iPhone 15 Pro. English’s paired uncertainty interval crosses zero, and first-use latency favored ONNX.

Measurements

Recorded result

Fewer observed transcription errors in both languages

iPhone 15 Pro · 24 English + 24 Spanish clips · 325 + 242 reference words

Fewer observed transcription errors in both languagesEnglish · ONNX · 2 CPU threads: 3.69 word error rate (%); English · Aidana · MLX GPU: 2.15 word error rate (%); Spanish · ONNX · 2 CPU threads: 9.50 word error rate (%); Spanish · Aidana · MLX GPU: 5.37 word error rate (%). Short public read speech, one completed pass per stack. English’s paired interval crosses zero; this does not establish all-accent or noisy-microphone accuracy.English · ONNX · 2 CPU threads3.69English · ONNX · 2 CPU threads: 3.69 word error rate (%)English · Aidana · MLX GPU2.15English · Aidana · MLX GPU: 2.15 word error rate (%)Spanish · ONNX · 2 CPU threads9.50Spanish · ONNX · 2 CPU threads: 9.50 word error rate (%)Spanish · Aidana · MLX GPU5.37Spanish · Aidana · MLX GPU: 5.37 word error rate (%)0word error rate (%)
  1. English · ONNX · 2 CPU threads3.69
  2. English · Aidana · MLX GPU2.15
  3. Spanish · ONNX · 2 CPU threads9.50
  4. Spanish · Aidana · MLX GPU5.37

word error rate (%)

Short public read speech, one completed pass per stack. English’s paired interval crosses zero; this does not establish all-accent or noisy-microphone accuracy.

View data & source
Fewer observed transcription errors in both languages · word error rate (%)
ConfigurationValue
English · ONNX · 2 CPU threads3.69
English · Aidana · MLX GPU2.15
Spanish · ONNX · 2 CPU threads9.50
Spanish · Aidana · MLX GPU5.37

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The Spanish improvement is clearer on this sample

Aidana − ONNX · 5,000 paired clip bootstrap samples · 95% intervals

The Spanish improvement is clearer on this sampleEnglish: -1.54 WER difference (percentage points), 95 percent interval -3.49 to 0.29; Spanish: -4.13 WER difference (percentage points), 95 percent interval -8.55 to -0.63. Lower favors Aidana. English crosses zero; the Spanish interval is below zero. This small sample does not establish performance across all speech.English-1.5495% interval: -3.49 to 0.29Spanish-4.1395% interval: -8.55 to -0.63Zero difference
  1. English-1.54

    95% interval: -3.49 to 0.29 · zero marked

  2. Spanish-4.13

    95% interval: -8.55 to -0.63 · zero marked

Lower favors Aidana. English crosses zero; the Spanish interval is below zero. This small sample does not establish performance across all speech.

View data & source
The Spanish improvement is clearer on this sample · WER difference (percentage points)
ConfigurationValue95% low95% high
English-1.54-3.490.29
Spanish-4.13-8.55-0.63

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The same 47 clips took less time to transcribe

47 matched clips · 279.724 seconds of audio · iPhone 15 Pro

The same 47 clips took less time to transcribeONNX · 2 CPU threads: 41.575 decode seconds; Aidana · MLX GPU: 12.400 decode seconds. One background-suspended clip is excluded from both engines. The first-use cost remains: Aidana’s first clip took 2.865 s versus 0.573 s for ONNX.ONNX · 2 CPU threads41.575ONNX · 2 CPU threads: 41.575 decode secondsAidana · MLX GPU12.400Aidana · MLX GPU: 12.400 decode seconds0decode seconds
  1. ONNX · 2 CPU threads41.575
  2. Aidana · MLX GPU12.400

decode seconds

One background-suspended clip is excluded from both engines. The first-use cost remains: Aidana’s first clip took 2.865 s versus 0.573 s for ONNX.

View data & source
The same 47 clips took less time to transcribe · decode seconds
ConfigurationValue
ONNX · 2 CPU threads41.575
Aidana · MLX GPU12.400

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

The model that feels fastest after loading may still make you wait for its first words. This paired iPhone experiment compared English and Spanish transcription, kept a suspended timing sample out of both engines, and weighed accuracy against first-use latency before selecting the app’s new default.

From the original notebook

Selected: Aidana mixed 8-bit weights, float16 computation, MLX GPU. On the actual iPhone 15 Pro it produced the lowest observed English and Spanish WER among the completed paired phone passes. The included SpokenKeep app source uses that exact checkpoint and pinned native runtime as its default.

Phone configuration English errors / words English WER Spanish errors / words Spanish WER Model creation
ONNX baseline, two CPU threads 12 / 325 3.69% 23 / 242 9.50% 1.536 s
Aidana mixed 8-bit / float16 / MLX GPU 7 / 325 2.15% 13 / 242 5.37% 1.421 s

The 47 valid matched timing clips showed 3.35× faster Aidana transcription. Creation timings are one observed launch each, not repeated medians or app startup. The native short-clip benchmark does not prove all accents, noisy microphone accuracy, or the fastest possible first-use implementation.

Evidence

Redux was tested using its supplied Photon runtimes on the Mac. The supplied Photon distribution does not provide the tested iOS runtime; it is not the phone’s selected backend. Aidana’s native runtime preserves packed integer weights instead of casting the entire checkpoint to floating point.

Build the included native app

cd speech-to-text/parakeet-iphone-2026-10-03/app
python3 scripts/fetch_build_assets.py
open ParakeetNotes.xcodeproj

Choose your own development team in Xcode and run scheme ParakeetNotes on a physical iPhone. The source retains iOS 18.0 and Swift 6. The package lock pins the tested dependency versions. The asset helper downloads the public model, Silero VAD and Sherpa-ONNX 1.13.2 distribution and verifies their SHA-256 values. Weights, framework binaries, signing material, credentials, private recordings and Handy history are excluded from this upload.

The app starts preparing the model during initialization and shows a loading bar. It records audio independently, saves it before deferred transcription, and retains audio on STT failure. The recording view places timer and waveform at the top, gives the transcript more height, follows new text, and briefly highlights its latest changed words. App source also includes native StoreKit and draft launch materials from the ongoing product work; the upload is not an App Store release. Subscription products and release/legal identity still require the owner’s App Store Connect setup.

iOS Simulator cannot execute MLX’s required Metal features. It supports recording/UI and resource verification; speech inference reports that limitation recoverably. Use the physical phone tests for MLX accuracy and startup timing.

Reproduce the model comparison

Follow the pinned environments and commands in the full comparison. Runners live in app/scripts/benchmarks. Keep each device’s runs serial, retain excluded timings, and compare regenerated corpus hashes against the captured manifest. Phone filenames marked cold mean no explicit warm-up, not cleared filesystem or GPU caches.

The default app’s assets are fully downloadable through the helper. Repeating the legacy ONNX comparison additionally requires its four baseline files at app/ParakeetNotes/ParakeetResources/Models/parakeet-sherpa/, matching the captured hashes. Those baseline files are not bundled in the upload or the new app; the earlier experiment used the source workstation’s existing assets.

The source snapshot was prepared from groxaxo/ParakeetNotes during this task, preserving its bundle identity. Experimentos stores a self-contained buildable snapshot and evidence; its root is not substituted for the original app repo.

Current app validation

The signed iPhone Release build succeeded and passed strict code-sign checks. The built weight file’s full SHA-256 matches the benchmark checkpoint. It was installed successfully on the iPhone 15 Pro; a new production launch and warm-up matrix remain blocked by the phone lock.

The simulator regression suite completed 54 tests: 47 passed, 7 skipped, zero failures. Four skipped checks require a physical MLX GPU; three native StoreKit checks skip the documented iOS 26.5 test-session configuration defect. Native manual QA also saved a 30-second recording while STT was unavailable. Validation records and fresh recording-view screenshots are included. UI screenshots use synthetic test text; the observed model WER comes from the separate completed real-phone benchmark.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (3)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS