BUILT FOR CODING IN SHARED SPACES

Voice your prompt.
Keep it to yourself.

Hold your iPhone close and press to talk. You can keep your voice down while your Mac turns it into text locally—right where your cursor already is.

A developer speaking quietly into an iPhone while text appears on a Mac in an open office

01 / THE SIMPLE IDEA

Your Mac is across the desk. Your iPhone can be inches away.

A distant laptop microphone picks up your keyboard, ventilation, and everyone nearby along with your voice. Move the microphone to your mouth and your voice reaches it first. That is what lets you speak more softly without asking the model to rescue poor audio.

IPHONE → MAC / HOW IT FEELS

Speak softly into your phone. Watch the words land in Codex.

The iPhone handles close-range capture. Encrypted audio goes to the Mac you paired, where it is transcribed locally and inserted at the active cursor.

Codex / Claude Codeslowmap / prototype
Building route preferencesScenicRouteRanker.swift
Build a walking-route map that ranks shade, noise, and slope instead of simply choosing the fastest path.
The undervoice speaking screen on an iPhone 17 Pro Max
AT YOUR DESK01
Speaking quietly into an iPhone at a desk

Prompts without the performance

Explain the change, the error, or the edge case in a low voice. The rest of the room does not need to join the conversation.

NEAR-FIELD AUDIO02
An iPhone held close to the speaker’s mouth

Give your voice the head start

Close-talk capture does not silence the office. It makes your voice stronger than the room before recognition even begins.

CODING FLOW03

Stay in the tools you already use

Keep focus in Codex, Claude Code, Xcode, Terminal, or a browser. The transcript arrives at the cursor.

02 / WHAT MATTERS AT A DESK

Useful numbers, not mystery metrics.

These are the facts that shape everyday use: how close the microphone is, where audio goes, and how much friction remains after you speak.

5–10 cm

A microphone you can bring close

Your iPhone can sit inches from your mouth while your MacBook stays 50–80 cm away on the desk.

14–24 dB

Less voice for the same input level

A distance-only estimate for 5–10 cm versus 50–80 cm. The room, device angle, and microphones affect real results.

0 cloud audio

No cloud transcription hop

Recognition runs on your Mac. Once the model is ready, local use does not need the public internet.

0 audio files

Nothing to clean up later

Audio is processed in memory and released. It is not left behind as a recording on either device.

0 copy / paste

The result goes where you were typing

No second app, clipboard shuffle, or context switch between speaking and coding.

1 hour

Short local history by default

Mac transcript history defaults to one hour and can be turned off. It never syncs to an undervoice server.

The boundary you can verify:paired devices only; no audio upload or audio files; insertion stops when focus changes; password and secure-input fields are excluded.

03 / MODELS AND LANGUAGES

Seven core languages. English technical terms included.

undervoice uses language-specific packs instead of pretending one model is equally good at everything. Chinese, Japanese, German, French, Spanish, and Korean packs accept English library names, variables, and commands; a dedicated English mode is included.

中文 + EnglishEnglish日本語 + EnglishDeutsch + EnglishFrançais + EnglishEspañol + English한국어 + English
ZH + ENLOCAL

Chinese with English in the sentence

A bilingual streaming model provides the live preview; Paraformer or the accuracy-first FireRedASR2 CTC produces the final text. Punctuation is restored locally.

Runtime
sherpa-onnx
Optimization
INT8 ONNX
DESIGNED FOR
fast preview, stronger final pass
EN / DE / FR / ESLOCAL

Four European language packs

NVIDIA Parakeet TDT 0.6B v3 powers the product-tuned English, German, French, and Spanish packs.

Runtime
sherpa-onnx
Optimization
INT8 ONNX
UPSTREAM BENCHMARK
6.34% average WER
JA + ENLOCAL

Japanese that keeps the English terms

Kotoba Whisper Bilingual v1.0 was trained for Japanese, English, and speech that moves between the two.

Runtime
whisper.cpp
Optimization
Q5_0
UPSTREAM BENCHMARK
9.3–16.8% Japanese CER
KO + ENLOCAL

Korean with names, commands, and acronyms

Whisper Large v3 Turbo covers Korean speech and the English technical language inside it while keeping the local final pass practical.

Runtime
whisper.cpp
Optimization
Q5_0
CHOSEN FOR
multilingual coverage and local speed
Read the accuracy figures carefully

WER and CER are lower-is-better. The published 6.34% and 9.3–16.8% figures come from upstream full-model test sets—not an overall undervoice score for close-talk iPhone audio, quantized models, or mixed-language speech. Different datasets and metrics do not make a fair leaderboard.

04 / IPHONE VS. MACBOOK MICROPHONE

Fix the input before judging the model.

Both routes can use the same local model on your Mac. The difference is where the audio begins: inches from your mouth or across the desk.

Shared-office useiPhone held closeMacBook microphone
Mouth-to-mic distance5–10 cm50–80 cm
Speaking quietlyYour voice reaches the mic firstVoice, keyboard, HVAC, and nearby speech arrive together
Voice level neededStay at a low voiceOften requires speaking up at desk distance
Where recognition runsLocally on your MacLocally on your Mac
Best fitOpen offices and shared desksPrivate rooms or conversations you do not mind sharing

The 14–24 dB estimate uses the free-field relationship 20 × log10(r₂/r₁) for 5–10 cm versus 50–80 cm. Reflections, placement, and speaking direction change real-world results.

05 / SECURITY AND NETWORKING

Your recording goes straight to the Mac you chose.

undervoice has no speech relay. On the same network, iPhone connects directly to Mac; across networks, you can use your own Tailscale setup. App-level authentication and encryption stay in place either way.

Installation and permission details
PAIRED

Only paired devices get in

Sharing Wi-Fi is not enough. A new iPhone must be explicitly paired or approved after a reset.

E2E

Encrypted before it leaves iPhone

Audio, transcripts, and control messages are authenticated and encrypted end to end.

ROTATE

A fresh key for every connection

Reconnect and a new session key is derived. Captured traffic from an old session cannot simply be replayed.

RAM

Audio never becomes a file

Recordings live in memory only. Diagnostics exclude audio and transcript text.

0 CLOUD

No undervoice speech cloud

On a local network, audio travels directly from iPhone to Mac. The model runs there.

REMOTE

Bring your own private network

For another Wi-Fi or cellular connection, use your own Tailscale network without dropping undervoice encryption.

06 / TYPE ANYWHERE

Wherever the cursor is, that is where the words go.

CodexClaude CodeXcodeTerminalAny text field

“Keep the public API, replace retries with exponential backoff, and add three boundary tests.”

listening

Your Mac remembers the active app and selection when dictation starts. If focus or selection changes, insertion pauses. Password fields and controls using Secure Keyboard Entry are never written to.

undervoice

TALK TO YOUR MAC, NOT THE ROOM

Speak softly. Code freely.

Install the model on Mac, scan the pairing code with iPhone, then hold to talk whenever an idea is faster to say than type.