For the people who will not type
Technicians, kitchen staff and field engineers send voice notes far more readily than they type. Hold the button, say what happened, carry on. The rest works exactly as it does for a typed message.
A voice note sent to the assistant is downloaded, transcribed and then treated as though it had been typed: it creates jobs, closes jobs, answers questions and is chased in exactly the same way. The reply opens with what the assistant heard, so a mis-heard registration or part number is visible immediately rather than turning into a wrong job nobody notices.
- What happens
- Transcribed, then handled as text
- The reply
- Opens with what it heard
- Language
- Set once, never guessed
- Cost
- About $0.002 for a 20-second note
- Needs
- An OpenAI key, even on Claude
One note, what happens
- They hold and talk
- 12 seconds
- Downloaded
- From WhatsApp
- Transcribed
- About a second
- Handled as
- An ordinary message
- Reply opens with
- What it heard
The last thing standing between the product and the people it is for
Every objection this product has ever had comes down to one thing: will your team actually use it. Removing the app removed most of that. What was left was typing - and a technician with dirty hands, standing under a car, is not going to type.
They will hold a button and talk, because they already do it a hundred times a week. That is the whole feature: the thing they already do, turned into a job list.
- Nothing new to learn - it is the same button they already use
- Works one-handed, walking, in a workshop, in a van
- The transcript becomes the message, so search and job matching see it
- Everything downstream - chasing, the briefing, the weekly review - works on it unchanged
It repeats back what it heard
Transcription mis-hears things, and the things it mis-hears are exactly the things that matter here: registrations, part numbers, names. A wrong job created silently from a wrong transcript is the one failure this feature could introduce.
So every reply to a voice note opens with the transcript, in quotes, before anything else. If it heard "Mrs Diaz" as "Mr Diaz" the sender sees it in the same second and can correct it. The note is also marked as spoken on the dashboard, so anybody reading the thread later knows those words were heard rather than typed.
- The reply starts: I heard: "…"
- The thread marks a spoken message, because a transcript is not the same as typed words
- A note it could not make out is refused, not guessed at
- A note that fails to download is answered, not silently dropped
- The language is set once in the admin and always sent, so the model never guesses it
Worth knowing before you switch it on
It needs an OpenAI key
Even if your assistant runs on Claude. Anthropic has no speech-to-text, so this one job goes to OpenAI on your own key. The admin says so rather than quietly doing nothing.
About a fifth of a cent
A 20-second note costs roughly $0.002 to transcribe. It appears in the same spend figures as everything else, so you can see it.
Your language
Set a two-letter code once. A model left to guess the language of a ten-second note gets it confidently wrong, so one is always sent.
Notes, not recordings
Anything over five minutes is refused with an explanation instead of being billed for. You can change the limit.
Off until you turn it on
With it off, a voice note gets a polite reply asking for text - the behaviour before this existed.
Same rules as everything
A number not on your roster is still refused. Voice changes what a message is made of, not who may send one.
More of the assistant
On this page
What if it mis-hears a part number?
You will see it straight away, because the reply opens with what it heard. Correct it in a follow-up message the way you would correct anything else. That is why the transcript is repeated back rather than quietly acted on.
Does it work in languages other than English?
Yes - set the language in the admin and it is used for every note. The assistant already replies in the language the person wrote in; this is the same idea for speech.
Why does it need an OpenAI key if we use Claude?
Because Anthropic has no speech-to-text API. Rather than pretend otherwise, the setting says so and lets you enter an OpenAI key used only for this. If you already run the assistant on OpenAI, it reuses that key and you enter nothing.
Where does the audio go?
From WhatsApp to your server, then to your own OpenAI account to be transcribed. We are not in that path - see <a href="/trust-and-security/">trust & security</a>. The audio is not kept; the transcript is, as the message.
Can it read photos too?
Not yet. A photo is stored on the conversation so it is not lost, but it is not understood. Tell us on the <a href="/contact/">contact form</a> if that is what your business needs next.