The user selected Aidana mixed 8-bit, float16 computation, MLX GPU as SpokenKeep’s default after the paired iPhone 15 Pro benchmark.
| Configuration | English errors / words | English WER | Spanish errors / words | Spanish WER | Model creation |
|---|---|---|---|---|---|
| Former ONNX, 2 CPU threads | 12 / 325 | 3.69% | 23 / 242 | 9.50% | 1.536 s |
| Selected Aidana mixed 8-bit, float16, MLX GPU | 7 / 325 | 2.15% | 13 / 242 | 5.37% | 1.421 s |
These accuracy results cover 24 public English and 24 public Spanish clips; creation is one observed launch per stack. The 47 matched valid timing clips showed 3.35× faster Aidana transcription. The previous first-use comparison did not establish the fastest possible whole-app startup. Full method, limits, raw predictions, source licenses and exclusions are in the comparison.
Exact default
- Checkpoint:
kyr0/aidana-parakeet-tdt-0.6b-8bit. - Revision:
af2f86c1a83a66e5b5184256af470cc8f4537871. - Weight bytes:
778033428. - Weight SHA-256:
0cb075a3c6d2ab65e44bae18ae35df3fcb3c4ab4f24ef94fe90da60a1668c014. - Config SHA-256:
3ee089787223b1c394898503a589bcb50c8c067fd8e7323e0d88c8a649ee0296. - Runtime:
Blaizzy/mlx-audio-swift, revision8d86630ade569728aaea3dc1a29fc44e2efa719b. - Native dependencies retain
mlx-swift 0.31.3andmlx-swift-lm 3.31.3. - The original checkpoint contains 221 U32 packed tensors and 918 BF16 tensors. The runtime’s verified quantized loader preserves the integer tensors, loads per-layer scales, and casts floating computation parameters to float16.
NotesViewModel injects ParakeetTranscriber by default; its actual recognition
implementation now calls ParakeetModel.fromDirectory(..., computeDType: .float16).
It uses MLX GPU evaluation and explicit synchronization, without an ONNX speech
fallback. Sherpa-ONNX remains for Silero VAD only. The app resource phase includes
Aidana, VAD and licenses; it excludes the former ONNX recognizer and Handy history.
Readiness and long recordings
App initialization starts preparation off the UI actor. The native loading bar tracks four actual stages: validation, model loading, VAD preparation and GPU warm-up. A one-second zero-valued audio buffer exercises inference without opening the microphone. Its output is discarded. The warm-up moves initial execution work into preparation; its net iPhone startup benefit still needs the new production phone trial. Cached preparation reuses the same model.
Recording starts and saves WAV audio independently of readiness. Pending transcriptions resume after preparation or reopening. Short clips use the benchmark’s greedy decoding path; clips longer than 30 seconds use the upstream runtime’s 30-second chunks with two-second overlap and token merging to bound GPU memory. The longest benchmark clip is 11.008 seconds, so chunking does not change the measured corpus’s decoding path. Long-recording accuracy is not established by the short-clip WER benchmark.
Reproduction and validation
python3 scripts/fetch_speech_model.py
python3 scripts/fetch_speech_model.py --verify-only
The installer and every app build verify the full weight/config hashes and audit the mixed tensor header. Launch checks the pinned manifest, size and small configuration hash without hashing the 778 MB weight file again.
The actual iPhone GPU tests cover public Spanish audio, silence, preload reuse, and three alternating trials with and without warm-up. Simulator tests verify the pinned bundle and recording/UI flows; MLX inference tests explicitly skip because iOS Simulator lacks the required Metal capabilities. A dedicated test checks that the unsupported GPU result is recoverable.
The signed optimized iPhone Release build succeeded, including the full-checkpoint build check. Strict deep code-sign verification passed. The built bundle’s 778,033,428-byte weight file was rehashed and matched the exact benchmark SHA-256. The bundle contains Aidana, Silero VAD and license notices; it contains no former ONNX recognition weights or private Handy history. Structured evidence is in the release identity record.
The updated Release was installed successfully on the iPhone 15 Pro. iOS reports that a passcode is required; a new production launch, microphone workflow and warm-up matrix require the owner to unlock the phone. The prior completed paired iPhone benchmark remains the evidence for the selected checkpoint’s accuracy; installation alone does not establish new production GPU execution.
The physical Device Hub crash-report list returned no matches for either
Parakeet or SpokenKeep on 3 October. Available Mac crash reports are simulator
reports: one audio-callback actor-isolation trap and two old ONNX initialization
faults. These are not a retrieved matching iPhone OTA crash. The isolated audio
callback fix is preserved, and production VAD initialization now scopes borrowed
C strings through the native constructor rather than using temporary NSString
pointers.
The current optimized simulator suite completed successfully: 54 tests executed, 47 passed, 7 skipped, zero failures. The simulator is SpokenKeep Monetization QA 20261002, iPhone 17 Pro, iOS 26.5 (23F77). Four skipped tests require the physical MLX GPU; three native StoreKit tests skip the documented iOS 26.5 configuration defect (Apple FB22237318). Recording independence, checkpoint identity, wrong-checkpoint rejection, scoped VAD configuration, transcript scrolling/highlights, light/dark/accessibility presentation, and organization during pending transcription passed. See the sanitized test summary.
Manual native simulator QA then granted the previously approved QA microphone permission, started recording while speech recognition reported unavailable, and saved a 30-second audio note with Done. The note showed a playback control and persisted the pending-transcription state without crashing. Timer and waveform were at the top, with the transcript card filling the remaining space. This proves the simulator capture/save flow under STT failure, not live physical-iPhone English/Spanish recognition.
Fresh native screenshots from the passing presentation tests: light, dark, and accessibility text. They use synthetic English/Spanish text and injected test speech/capture services in the production view; they demonstrate layout and highlighting, not recognition accuracy.
env TMPDIR=/private/tmp/spokenkeep-xcode-temp/ xcodebuild \
-project ParakeetNotes.xcodeproj -scheme ParakeetNotes \
-configuration Release \
-destination 'platform=iOS Simulator,id=2B1BE9FA-E440-4E3C-91D3-F3F5B79795CF' \
-derivedDataPath /private/tmp/SpokenKeep-Aidana-Release-20261003 \
-clonedSourcePackagesDirPath WORKSTATION/SourcePackages \
-disableAutomaticPackageResolution -parallel-testing-enabled NO -jobs 4 \
ONLY_ACTIVE_ARCH=YES ENABLE_TESTABILITY=YES \
CODE_SIGNING_ALLOWED=YES CODE_SIGN_IDENTITY=- \
-resultBundlePath /private/tmp/SpokenKeep-Aidana-Tests-Final-20261003.xcresult \
-only-testing:ParakeetNotesTests test
An initial simulator compilation was stopped before tests ran to correct the Release testability override and restrict the app to the active simulator architecture. The completed result above is the subsequent full test run. Earlier ONNX simulator results in the recording documents are historical and do not prove this migration.