After every job, one question.
"Anything the next driver should know?" The driver taps once, speaks for up to thirty seconds, and submits. No forms, no typing, no menus. Hands-free by design.
- One tap to record, one tap to stop.
- Maximum recording time — 30 seconds.
- Runs entirely inside the Deliverd driver app.
- Not a chatbot. Not a voice assistant. An input method.
Anything the next driver should know?
Speech in. Structured operational fields out.
Every recording is parsed into the operational schema Deliverd already uses. Free-form language becomes typed values the API can act on.
Structured observations. Never raw speech.
Voice is a capture surface, not a storage format. The primary record is always the structured observation — the transcript is discarded once the fields are confirmed.
"Parked on the main road, no forklift here, use the rear loading bay — waited about fifteen minutes. Watch the stone steps, take a trolley."
One question. Only if we need it.
No forklift detected. Can you confirm?
The person who lives there knows the shortcut.
Customers can record short access notes before delivery. The same extraction pipeline turns their sentences into structured Arrival Passport fields.
"Our back gate is easier."
"The sofa won't fit through the front door."
"There are external stone steps."
One voice note. Four data products.
Access, entrance, stairs, doorway.
Loading bay, security, forklift, PPE.
Equipment and crew flags.
Waiting time, handball, dwell.
Never store unstructured notes where structured data can be extracted.
A single large microphone. Nothing else on the screen competes with it.
Voice Capture is a way to fill fields faster. It does not answer, chat or advise.