Recommendation: Aidana 8-bit with float16 computation in the native Swift runtime. It produced the lowest observed English/Spanish word error rates and faster transcription on the actual iPhone 15 Pro. On 3 October 2026, the user selected this exact configuration as the production default. The app now loads the pinned Aidana checkpoint through the same pinned MLX Audio Swift runtime, using float16 computation and preserving its packed mixed tensors. Preparation starts during app initialization and includes visible loading stages and a one-second synthetic-silence GPU warm-up. No microphone access is used to preload. See the default-model implementation and validation.
These results compare model and runtime combinations. They do not establish the fastest possible implementation across every iPhone backend. Repeated phone reload, precision and explicit warm-up trials remain incomplete after the phone locked and its remote process connection became unavailable.
Completed iPhone tests
Target: YeahRight, iPhone 15 Pro / iPhone16,1, iOS 27.0. Both completed runs used the same bundled 16 kHz mono WAVs. The QA app never opened the microphone. Each accuracy pass contains 24 English and 24 Spanish clips.
| Native engine | English errors / words | English WER | Spanish errors / words | Spanish WER | Model creation |
|---|---|---|---|---|---|
| Existing int8 ONNX, 2 CPU threads | 12 / 325 | 3.69% | 23 / 242 | 9.50% | 1.536 s |
| Aidana mixed 8-bit, float16 computation, MLX GPU | 7 / 325 | 2.15% | 13 / 242 | 5.37% | 1.421 s |
Creation times are one measured launch each, not repeated medians or whole app startup times. Creation ends after the loader materializes weights; MLX evaluation and GPU synchronization are explicit. The harness times model creation separately from SpokenKeep’s VAD preparation and UI startup.
Aidana’s first 3.072-second clip took 2.865 s to transcribe, including WAV reading and initial GPU execution/compilation. ONNX’s first clip took 0.573 s. Load plus first decode is approximately 4.286 s for Aidana and 2.109 s for ONNX. Faster steady transcription does not by itself prove faster first-use readiness. No iPhone explicit warm-up trial has completed yet.
The ONNX pass was background-suspended during one English clip, giving a 141.39-second wall-clock measurement. It resumed and completed all 48 transcriptions, so its accuracy remains usable. Whole-pass throughput is invalid. Excluding that same clip from both engines gives a paired speed comparison:
| Engine | Matched clips | Matched audio | Decode time | Audio seconds per processing second |
|---|---|---|---|---|
| ONNX, 2 threads | 47 | 279.724 s | 41.575 s | 6.73× |
| Aidana, float16 | 47 | 279.724 s | 12.400 s | 22.56× |
Aidana is 3.35× faster on this subset, including its first-use cost. The exact excluded clip and raw measurements are retained. Its slower first transcription has not been removed from these numbers.
A four-thread ONNX run loaded in 1.066 s and logged 47 finished clips before the CoreDevice console connection was invalidated. It did not produce a completed JSON. Resource caches were already warm; this partial result is not a valid repeated loading comparison or a complete accuracy pass. The earlier production two-thread setting provided the baseline; the user subsequently selected Aidana based on the completed paired accuracy and throughput results.
Aidana reported nominal thermal state at start and finish. Its sampled RSS at readiness and maximum recorded sample was 737 MB, versus 1,394 MB and 1,468 MB for ONNX. RSS is not a complete GPU footprint or a continuously sampled allocation peak. A revised harness adds physical-footprint and MLX allocation counters and an active-view idle timer; it built successfully, but installation was blocked by the phone connection state. RSS alone does not establish lower total iPhone memory use.
A 5,000-sample paired clip bootstrap gives Aidana-minus-ONNX WER differences with 95% intervals of −3.49 to +0.29 percentage points for English and −8.55 to −0.63 points for Spanish. English’s interval crosses zero. Spanish’s observed improvement is stronger on this sample. Neither establishes accuracy across all accents or noisy microphone conditions.
Both requested checkpoints on the Mac
Host: Apple M5, 24 GB, macOS 26.6.1. Full passes use the same 48 clips. Load medians below use three fresh processes; filesystem and Metal shader caches remain present. Python imports are measured separately from model loading.
| Model and Python runtime | Median model ready | English WER | Spanish WER | English speed | Spanish speed |
|---|---|---|---|---|---|
| Existing ONNX, 2 CPU threads | 0.707 s | 3.38% | 9.50% | 29.73× | 29.57× |
| Aidana, MLX GPU | 0.278 s | 2.46% | 7.85% | 51.49× | 45.52× |
| Aidana, MLX CPU | 0.195 s | 2.46% | 7.85% | 1.11× | 1.09× |
| Redux, Photon CPU, 4 threads | 2.571 s | 6.46% | 10.33% | 48.26× | 49.04× |
| Redux, Photon Metal | 2.567 s | 6.77% | 9.50% | 28.52× | 21.08× |
Additional two-trial load medians: ONNX 1 thread 0.673 s; ONNX 4 threads 0.694 s; Redux CPU 2 threads 2.243 s. These shorter runs use the first three clips for reload/first-call measurements and do not replace full accuracy passes.
Redux’s weight file is 177,774,490 bytes; Aidana’s is 778,033,428 bytes. The current ONNX weights and token file total 670,478,772 bytes. Smaller disk size did not give Redux the fastest readiness in the measured Mac runtime.
Redux uses ternary weights and Photon. Photon’s local requirements list Apple Silicon Macs and NVIDIA desktop/server platforms. The downloaded native binaries also declare macOS rather than iOS. The supplied Photon distribution was tested on the Mac and was not installed as an iPhone runtime; an iOS port or supported mobile SDK would be required.
Aidana uses mixed quantization. The Python loader quantizes only layers with scales in the checkpoint and preserves packed integer tensors. The native benchmark uses the unmodified pinned mlx-audio-swift loader, which successfully ran Aidana on the real iPhone. The production app now pins that same runtime revision and checkpoint; the benchmark itself remains isolated.
Native Swift precision and warm-up on the Mac
Each precision has one full 48-clip pass and two fresh-process, three-clip reloads.
| Aidana computation | Median model ready | Median first clip | English WER | Spanish WER |
|---|---|---|---|---|
| bfloat16 | 0.145 s | 0.053 s | 2.15% | 5.37% |
| float16 | 0.161 s | 0.051 s | 2.15% | 5.37% |
| float32 | 0.156 s | 0.056 s | 2.15% | 5.37% |
All precisions produced the same aggregate error counts. The first full passes had first-clip costs of 3.741 s, 2.045 s and 1.932 s respectively; subsequent processes benefited from shader caches. Order differed between reload trials. Initial precision costs cannot be treated as a comparison with all GPU caches cleared.
Three float16 startup warm-up trials used one second of silence. Medians were 0.157 s loading, 0.044 s warm-up and 0.038 s first speech decoding. With existing caches, warming added 0.044 s to preparation and reduced the median first speech call by only about 0.013 s. This does not show an improvement in total startup latency or settle iPhone warm-up behavior.
Sampling and scoring
Seed 20261002 selected uniform random rows within predefined accent/gender strata before recognition. Seed + 1 shuffled the paired execution order.
- English: 24 clips spanning all 11 configurations of ylacombe/english_dialects, recorded by volunteers speaking varieties of English from the British Isles.
- Spanish: four clips per male/female configuration of Argentinian, Chilean and Colombian Spanish, totaling 24 clips.
Total audio: 287.574 seconds, 325 English and 242 Spanish reference words. Inputs use band-limited resampling to 16 kHz mono PCM16, without amplitude normalization. Every paired engine receives the same canonical WAV bytes.
WER and CER use Unicode NFKC, case folding and punctuation removal; accents and numeric/unit formatting are preserved. Written digits and number words can count as different despite conveying the same quantity. These scores therefore include formatting disagreements. This is a small sample of short read speech.
The dataset cards declare CC BY-SA 4.0. Attribution/source links, dataset revisions, configurations, row indices, anonymized speaker IDs and source/WAV hashes are retained in the manifest. Reused public reference text retains that license. Sources are OpenSLR/Google speech collections redistributed by Ylacombe. The repository evidence contains no private recording or signed download URL.
Evidence and reproducibility
- Summary: trial counts, paired throughput, confidence intervals and exclusions.
- Raw results: 35 completed result files with per-clip predictions and decode times.
- Model hashes and sizes.
- Python package versions.
- Timing exclusion.
- Paired throughput clip exclusion.
| Repository | Pinned revision |
|---|---|
| moondream/parakeet-redux | 2bf128600aac4b16946f7ed8372e56117fe5e23b |
| kyr0/aidana-parakeet-tdt-0.6b-8bit | af2f86c1a83a66e5b5184256af470cc8f4537871 |
| Blaizzy/mlx-audio-swift | 8d86630ade569728aaea3dc1a29fc44e2efa719b |
Python: parakeet-mlx 0.5.3, MLX 0.32.3, moondream 2.6.1, kestrel 0.9.1 and sherpa-onnx 1.13.8. Native iPhone: vendored Sherpa-ONNX 1.13.2, mlx-swift 0.31.3 and mlx-swift-lm 3.31.3. Runtime-version differences are part of the tested stacks.
From the repository root, use a separate Python 3.12 environment:
benchmark_root=/private/tmp/parakeet-comparison-20261002
uv venv --python 3.12 "$benchmark_root/venv"
uv pip install --python "$benchmark_root/venv/bin/python" \
-r docs/ParakeetNotes/benchmarks/20261002/python-package-versions.txt
"$benchmark_root/venv/bin/hf" download moondream/parakeet-redux \
config.json tokenizer.json ternary.json model.safetensors README.md \
--revision 2bf128600aac4b16946f7ed8372e56117fe5e23b \
--local-dir "$benchmark_root/models/redux"
"$benchmark_root/venv/bin/hf" download kyr0/aidana-parakeet-tdt-0.6b-8bit \
config.json tokenizer.model vocab.txt model.safetensors README.md \
--revision af2f86c1a83a66e5b5184256af470cc8f4537871 \
--local-dir "$benchmark_root/models/aidana"
"$benchmark_root/venv/bin/python" scripts/benchmarks/parakeet_compare.py dataset \
--root "$benchmark_root"
"$benchmark_root/venv/bin/python" scripts/benchmarks/parakeet_compare.py run \
--root "$benchmark_root" --engine aidana-gpu --trial 1
Other engines: aidana-cpu, redux-cpu-2, redux-cpu-4, redux-mps,
onnx-1, onnx-2, onnx-4. --limit 3 is for reload trials; full accuracy
passes omit it. Compare regenerated corpus hashes with the captured manifest;
future dataset revisions may change row ordering. Keep runs serial per device.
The native helper requires Xcode and XcodeGen and creates a separate app,
com.blackstoreonline.ParakeetModelBenchmark:
python3 scripts/benchmarks/prepare_native_benchmark.py --root "$benchmark_root"
xcodebuild \
-project "$benchmark_root/native-benchmark/ParakeetModelBenchmark.xcodeproj" \
-scheme NativeParakeetiPhone -configuration Release \
-destination 'generic/platform=iOS' \
-derivedDataPath "$benchmark_root/NativePhoneDerivedData" -jobs 2 build
The actual build reused the production checkout’s SourcePackages via
-clonedSourcePackagesDirPath and disabled automatic resolution after initial
resolution. It retained iOS 18.0 and Swift 6. NativeParakeetMac builds the Mac
counterpart. Install with xcrun devicectl device install app --device <target>
and the built app path. Prepare a JSON case list, for example
[{"engine":"aidana-float16","trial":4,"limit":3,"warmup":true}], then run:
python3 scripts/benchmarks/run_native_iphone.py \
--root "$benchmark_root" --device '<unlocked-target>' \
--cases "$benchmark_root/iphone-cases.json"
"$benchmark_root/venv/bin/python" scripts/benchmarks/summarize_results.py \
--root "$benchmark_root" --output "$benchmark_root/summary.json"
The runner restarts only the QA bundle, copies completed JSON before advancing
and refuses to replace existing local evidence. Keep the phone unlocked and QA
foreground. The revised QA view disables its own idle timer while active, without
changing global lock settings. cold in filenames means no explicit warm-up,
not a powered-off phone or cleared caches. Logs, large checkpoints/WAVs, signing
outputs and phone connection metadata remain outside the repository in
/private/tmp/parakeet-comparison-20261002.
Production validation before the Aidana migration
Scheme ParakeetNotes. QA simulator SpokenKeep Recording QA 20261002,
iPhone 17 Pro, iOS 26.5,
382473D2-1902-4F8B-8578-FC6BC8CB8653.
26 tests passed, zero failures, including the regression test verifying that app initialization starts preparation before any screen task:
xcodebuild -project ParakeetNotes.xcodeproj -scheme ParakeetNotes \
-destination 'platform=iOS Simulator,id=382473D2-1902-4F8B-8578-FC6BC8CB8653' \
-derivedDataPath /private/tmp/ParakeetNotes-CrashQA-20261002 \
-clonedSourcePackagesDirPath WORKSTATION/SourcePackages \
-disableAutomaticPackageResolution \
-resultBundlePath /private/tmp/ParakeetNotes-AppPreloadPriorityTests-20261002.xcresult \
-parallel-testing-enabled NO CODE_SIGNING_ALLOWED=YES CODE_SIGN_IDENTITY=- \
-only-testing:ParakeetNotesTests test
The simulator app was installed, launched and showed readiness before recording. The updated production Release also built successfully:
xcodebuild -project ParakeetNotes.xcodeproj -scheme ParakeetNotes \
-configuration Release -destination 'generic/platform=iOS' \
-derivedDataPath /private/tmp/ParakeetNotes-PhonePreloadQA-20261002 \
-clonedSourcePackagesDirPath WORKSTATION/SourcePackages \
-disableAutomaticPackageResolution -jobs 2 build
Strict code-sign verification passed with the host trust store. SpokenKeep was installed and launched on YeahRight; the actual iPhone screen showed Speech recognition ready without starting recording. Both native benchmark Release builds, including the revised memory/idle-timer build, also succeeded.
The original recording crash remains unconfirmed: no matching Parakeet phone crash report was retrieved. The later CoreDevice/XPC failure is transport evidence, not proof of an app crash. Earlier simulator QA exposed and fixed an audio-tap actor-isolation regression, with a passing background-callback test. These on-phone public-file benchmarks do not establish a new controlled physical microphone recording test.