Back to the experimentSupporting notebook

Parakeet on iPhone: English/Spanish accuracy and Aidana default

Source: speech-to-text/parakeet-iphone-2026-10-03/README.md · revision 6550ead3945b

Selected: Aidana mixed 8-bit weights, float16 computation, MLX GPU. On the actual iPhone 15 Pro it produced the lowest observed English and Spanish WER among the completed paired phone passes. The included SpokenKeep app source uses that exact checkpoint and pinned native runtime as its default.

Phone configuration English errors / words English WER Spanish errors / words Spanish WER Model creation
ONNX baseline, two CPU threads 12 / 325 3.69% 23 / 242 9.50% 1.536 s
Aidana mixed 8-bit / float16 / MLX GPU 7 / 325 2.15% 13 / 242 5.37% 1.421 s

The 47 valid matched timing clips showed 3.35× faster Aidana transcription. Creation timings are one observed launch each, not repeated medians or app startup. The native short-clip benchmark does not prove all accents, noisy microphone accuracy, or the fastest possible first-use implementation.

Evidence

Redux was tested using its supplied Photon runtimes on the Mac. The supplied Photon distribution does not provide the tested iOS runtime; it is not the phone’s selected backend. Aidana’s native runtime preserves packed integer weights instead of casting the entire checkpoint to floating point.

Build the included native app

cd speech-to-text/parakeet-iphone-2026-10-03/app
python3 scripts/fetch_build_assets.py
open ParakeetNotes.xcodeproj

Choose your own development team in Xcode and run scheme ParakeetNotes on a physical iPhone. The source retains iOS 18.0 and Swift 6. The package lock pins the tested dependency versions. The asset helper downloads the public model, Silero VAD and Sherpa-ONNX 1.13.2 distribution and verifies their SHA-256 values. Weights, framework binaries, signing material, credentials, private recordings and Handy history are excluded from this upload.

The app starts preparing the model during initialization and shows a loading bar. It records audio independently, saves it before deferred transcription, and retains audio on STT failure. The recording view places timer and waveform at the top, gives the transcript more height, follows new text, and briefly highlights its latest changed words. App source also includes native StoreKit and draft launch materials from the ongoing product work; the upload is not an App Store release. Subscription products and release/legal identity still require the owner’s App Store Connect setup.

iOS Simulator cannot execute MLX’s required Metal features. It supports recording/UI and resource verification; speech inference reports that limitation recoverably. Use the physical phone tests for MLX accuracy and startup timing.

Reproduce the model comparison

Follow the pinned environments and commands in the full comparison. Runners live in app/scripts/benchmarks. Keep each device’s runs serial, retain excluded timings, and compare regenerated corpus hashes against the captured manifest. Phone filenames marked cold mean no explicit warm-up, not cleared filesystem or GPU caches.

The default app’s assets are fully downloadable through the helper. Repeating the legacy ONNX comparison additionally requires its four baseline files at app/ParakeetNotes/ParakeetResources/Models/parakeet-sherpa/, matching the captured hashes. Those baseline files are not bundled in the upload or the new app; the earlier experiment used the source workstation’s existing assets.

The source snapshot was prepared from groxaxo/ParakeetNotes during this task, preserving its bundle identity. Experimentos stores a self-contained buildable snapshot and evidence; its root is not substituted for the original app repo.

Current app validation

The signed iPhone Release build succeeded and passed strict code-sign checks. The built weight file’s full SHA-256 matches the benchmark checkpoint. It was installed successfully on the iPhone 15 Pro; a new production launch and warm-up matrix remain blocked by the phone lock.

The simulator regression suite completed 54 tests: 47 passed, 7 skipped, zero failures. Four skipped checks require a physical MLX GPU; three native StoreKit checks skip the documented iOS 26.5 test-session configuration defect. Native manual QA also saved a 30-second recording while STT was unavailable. Validation records and fresh recording-view screenshots are included. UI screenshots use synthetic test text; the observed model WER comes from the separate completed real-phone benchmark.