Click the microphone and talk. Words appear in the compose box as you speak — each sentence transcribed at a pause and appended, never revised underneath you.
It knows your machines
The session’s own title, machine names and repository names are sent as bias, so
vm-03 and the name of your repo come back spelled properly instead of
phonetically. Spoken punctuation — “comma”, “full stop”, “new paragraph” —
becomes symbols, through a rules list an operator can edit.
It gets out of the way
Click into the box and the microphone pauses. Two seconds after you stop typing it resumes, appending to what you typed rather than overwriting it. It stops itself after silence, and after two minutes regardless.
It never sends on its own
This is the important property, and it is asserted by a test rather than left to good intentions.
You read what is in the box and press Send yourself. Transcription models occasionally invent a sentence out of silence, and a message nobody read should not reach a machine that can run commands.
The same reasoning runs deeper: a spoken answer can never authorise a destructive tool call. Dictation puts text in a box. It does not answer approvals.
Choosing a provider
An administrator can configure several speech providers at once, hold a standing fleet vocabulary — one term per line, with known mishearings — and then measure them against each other: record a sentence, type what you actually said, and every configured provider is scored on the same audio with a word error rate and a marked-up alignment.