Which Request Reaches a Voice Agent's Tool?
In a financial-query benchmark, better transcription generally accompanied more correct tool calls, but reported errors could change the requested date range while added filler could leave the query intact.
The booking reference ended in M. In a developer’s posted March 2025 experiment, a voice assistant read it back ending in 8. The person corrected it. The next read-back ended in 9. After another correction, the assistant said it would look up a reference ending in F. Its posted function call contained F too.
The developer says the input transcript was correct each time. Without raw audio or independent reproduction, the account cannot locate the internal cause. But it puts the practical question in a precise place: which reference did the assistant actually send to the lookup?
For your own agent work, one useful evaluation question is whether the actual tool call retains the identifying values the person requested. This is a suggested check, not a tested fix. A transcript, spoken read-back and tool argument are different objects to compare. A correct argument would still leave the retrieved information and final answer to be checked.
NatWest’s new CavaBench preprint, submitted October 5, measures part of that journey for financial queries. Better transcription generally accompanied more correct tool calls. Selected errors also reveal why a merchant or interval needs attention even when most of a sentence survives. The study tests recorded requests, not naturally occurring banking sessions or financial losses.
Words that select the records
A spending query needs more than a recognizable topic. It needs the conditions that select the intended records: whose purchases, at which merchant, over which interval. A fluent answer about spending could concern a different set of purchases if one of those conditions changed.
The researchers retained 2,919 recordings from 90 participants reading financial prompts. Recognizers fed the same language model, using clean and artificially degraded audio across self-reported accent groups. A tool call passed only when all 5 fields matched its reference: measure, period, grouping, filter and comparison.
In clean audio, 75.1%–80.7% of calls passed, depending on configuration. The paper does not report how the downstream model performed with perfect transcripts. Some mismatches could therefore involve interpretation or matching rules as well as recognition. Nor does a strict mismatch establish that every resulting answer would be unusable.
Word error rate counts substitutions, omissions and insertions relative to a reference transcript. It applies a consistent rule, without giving a duration or merchant extra weight because that word determines a search. Across configurations, averaged over acoustic conditions, fewer word errors strongly tracked more correct calls, with a correlation of −0.93. Transcription quality was useful evidence for choosing a recognizer in this test.
The exceptions help explain what that aggregate relationship leaves open. The authors report requests for 60 days becoming 6 or 16 days, changing the query interval. In a simulated overlapping-speech condition, a chunked Qwen recognizer reached 49.8% word error rate while retaining 67% correct calls. The authors attribute much of its transcription error to inserted filler.
Filler can leave the requested conditions available to the model. A changed duration supplies a different condition. Both are transcription errors, but they ask different things of the next stage: disregard extra words, or recover an interval no longer stated correctly. A model can have enough information to form a query without having enough to reconstruct the intended one.
What else the model knows
An independent Korean question-answering experiment shows how much an answer can survive imperfect transcription in a different task. Under low-error conditions, models retained roughly 96%–99% of the answer-overlap score achieved with original question text.
Those questions were synthesized speech, and the models also received the original passage containing the answer. Questions containing digits were excluded. The experiment thus gave the model information beyond the damaged question. CavaBench’s model had to construct financial tool arguments from the transcript.
This difference suggests why recovery should be judged against the particular task and available information. Answering from a supplied passage does not establish that an agent can recover a missing merchant or date. The Korean experiment also found no consistent improvement from telling models to anticipate transcription mistakes. Being invited to infer intent was not a dependable repair in that test.
The booking-reference report marks another boundary. If its developer’s account of correct transcripts is accurate, better recognition alone would leave the reported downstream change unexplained. CavaBench and the posted experiment concern different systems; neither supplies the other’s internal cause.
Someone has to repair the request
People have several ways to keep going when voice fails them. In interviews with 30 voice-assistant shoppers in the UK and Nigeria, collected in 2022–23, one Nigerian participant described repeating words more clearly and still receiving a different search. A UK participant switched to typing after unsuccessful repetitions. Another handled calls and messages manually, taking care with particular names.
These are separate retrospective accounts from selectively recruited users, including people who felt companionship with their assistants. They do not measure today’s financial agents. They do document work the person takes on: say it again, enter it another way, or finish the task themselves.
Offering written input could give someone another route to the intended query. It could also impose a new difficulty. In the same interview study, a participant with dyslexia described asking Google to spell words. Voice helped with something they found difficult in written interaction.
That makes the choice of fallback part of the task, too. The person trying to retrieve a booking should not have to keep supplying corrections that never reach the lookup. The person asking for spelling may have chosen voice to get help with writing in the first place.
Sources
- Aadam Haq et al., Mind the Accent Gap: British Accent Robustness in Speech-Driven Financial Voice Assistants, October 5, 2026, arXiv v1; an ICASSP 2027 submission. Published under CC BY 4.0.
- ablestyle, Tips for improving accuracy of realtime voice input capture for licence plates and so forth?, March 17, 2025. Developer’s report and posted booking-reference exchange.
- Donghyuk Jung and Youngwon Choi, Analyzing Error Propagation in Korean Spoken QA with ASR–LLM Cascades, May 17, 2026, arXiv v1.
- Obinna Alo et al., The user experience of voice assistants in retailing: a qualitative comparative study, 2025. Interviews, recruitment and collection dates.