What to Keep From a Spoken Answer
In a 5-day study, people used notes and diagrams to ask follow-up questions and resume spoken AI conversations, but requests to change the existing material sometimes went unmet.
A learner studying Dutch asked a voice agent to write down its examples. Hearing them was not enough to remember them, the learner explained. In the recorded exchange, the agent made a note containing examples and meanings. The learner then asked for English equivalents, which appeared as annotations beside them.
The Dutch examples stayed in place while the English was added. A spoken explanation had become material the learner could ask the agent to work on: first keep these sentences, then add something to them.
The agent, VoCa, speaks while creating notes, diagrams and other material on a canvas. Its study was announced on arXiv on October 6. The researchers analyzed 18 young, experienced AI users speaking Chinese over 5 days. They chose their topics, but daily use was required and long conversations encouraged. This was an exploratory deployment, with no speech-only comparison or test of learning gains.
The study documents how people kept a conversation going with what the agent had put on the screen. They asked to retain particular examples, explain a drawn relationship or update an earlier plan. The next request could be small, but it changed what the visible material needed to do.
Keeping the example available
The Dutch learner needed selected examples, not a written copy of everything the agent said. Once those examples were on the canvas, a later request could concern the sentences themselves. The English equivalents belonged beside them, where the learner could consult both.
That sequence shows two jobs an explanation can do. The agent can answer aloud, and it can leave something to inspect after the answer has finished. Those jobs need not happen at the same pace.
Another participant described putting their phone aside during a reply, then consulting the mind map after the agent finished. Sometimes they skipped the speech and went straight to the canvas’s key points. A display meant to accompany speech had also become something to read separately.
For the Dutch learner, retaining examples made an addition possible. For someone inspecting a diagram, retaining a mark could raise a question about the original answer.
In a separate philosophy discussion, a participant asked what 2 arrows pointing toward existentialism meant. The agent explained and annotated the existing diagram. The question concerned a relationship already drawn, rather than another general account of the subject.
An arrow asserts a connection without necessarily explaining its kind. Giving that mark a visible place made it possible to ask the agent to account for it. The episode does not establish that the philosophy was correct; it shows an answer becoming the subject of further examination.
The column that stayed separate
A later request can ask for more than an explanation attached to an existing object. It can ask to change how that object works.
While developing a swimming plan, another participant requested a tutorial-link column in the existing stage table. They had appreciated the initial plan and annotations. But the agent supplied tutorials in a separate table and annotated how to combine the material, leaving the requested column absent. The participant reported that the subsequent tables increased the burden of following the plan.
A column places another kind of information beside each row’s existing contents. Supplying a second table can add the information while leaving the user to bring the arrangements together. Here, the requested location was part of the requested help.
VoCa was built to preserve committed contributions, expressing revisions through additions and annotations rather than overwriting or deleting. People could move whole objects, select, pan and zoom, but could not directly edit the contents or fine-tune the layout. The agent had ways to attach English to a Dutch example or explain an arrow. It did not insert the column into the table the participant wanted to keep using.
Keeping an earlier arrangement can still be useful. Another participant valued the visible strike-throughs left in a revised timetable. A record of what changed and a current plan with a new structure are both reasonable things to want. They require different operations.
An older study, Orality, documents another way to keep earlier work available. Its editable canvas organizes a person’s spoken thoughts. One participant refining interview responses repeatedly restored earlier states while trying different narrative structures for their work experience. The current arrangement could change while a previous one remained recoverable.
That is a different task from receiving a VoCa-style spoken explanation. Orality’s 12-person comparison with ChatGPT dictation found no statistically significant difference in rated clarity change or subjective workload. Its restoration episode documents a way back after trying a new structure, without establishing that the same design would have improved the swimming-plan interaction.
Coming back to the same list
Other updates fit comfortably beside an unchanged entry. A task list first names things to do; later, its entries can become places to report progress. Adding a mark can make the change the person wants while leaving the activity recognizable.
One VoCa participant dictated a list intending to keep the session open for several days. The recorded return came about 35 hours later. They reported completing 2 activities, and the agent attached completion badges to the original entries.
The activities were the participant’s report, not independently verified completions. The update to the canvas was concrete: beside the recruitment-poster item and dinner, the same list now carried completion marks.
Sources
- Yate Ge et al., VoCa: Designing Speech–Canvas Interaction for Voice-Based Conversational Agents, arXiv v1, submitted October 3, 2026; announced in the October 6 human-computer interaction group. Methods, prototype policy, deployment episodes and limitations inspected, including Appendix B.
- Wengxi Li, Jingze Tian and Can Liu, Orality: A Semantic Canvas for Externalizing and Clarifying Thoughts with Speech, CHI 2026; arXiv v1, March 3, 2026. Comparison methods, results and version-restoration account inspected.