Skip to content

Working Out What the Sources Can Say

In one experiment, writers given AI guidance before generation spent longer composing an argument and less time judging separate passages about the same sources than writers given immediate generation.

In a writing experiment, a participant asked an AI assistant how a Swedish study had measured sleep. Had there been a physical investigation, or were the findings assumptions?

The assistant pointed back to the supplied source summary. The participant found the answer: sleep and stress diaries. They then said they would have preferred a physical tracking device.

This exchange appears in the published conversation record of a person assigned to write about the four-day workweek. The writer had moved from a reported improvement to a question about how anyone knew it had happened. The assistant helped direct the inquiry, but did not turn it into a paragraph for the essay. At that stage, its instructions prohibited doing so.

The researchers were trying an unusual arrangement: AI could guide the writer from the beginning, but could not yet supply arguments or draft prose. Compared with immediate generation, this package led people to spend longer writing and less time judging separate passages about the same sources. It did not establish a net time saving or better checking overall.

For someone asking AI to produce a source-based draft, that pattern raises a practical question. When do you get to know the material well enough to judge what the prose says?

Before the draft arrived

Xiaotian Su and colleagues reported the experiment in an October 1 preprint. Their analysis included 398 participants who passed attention checks. Participants were asked to write at least 200 words using four summaries about the four-day workweek, including supportive and critical evidence.

In the first recruitment phase, participants were randomly assigned to write without AI, with immediate access to a generative chatbot, or with the restricted tool. The restricted version, called Engage-to-Unlock, began in guidance mode. Generation became available when a classifier detected a position and a supporting argument in the draft. A progress indicator displayed that development. Even after generation unlocked, the writer had to choose to switch modes.

This was guided work before generation. The assistant’s instructions prohibited supplying topic-specific arguments, evaluating their strength, proposing an outline or rewriting the writer’s prose. It could ask questions and teach general writing principles. The same model powered both modes.

After composing their essays, participants judged eight separately supplied AI-generated passages based on the same sources. Four were sound; four contained deliberately introduced errors in reporting evidence or drawing conclusions. The sources remained available. AI assistance did not.

They were not checking their own essays. The experiment tested subsequent judgments about familiar material, rather than the full process of getting one’s own generated document ready to publish.

The reported time pattern was striking. Engage-to-Unlock participants spent an average of 28.02 minutes writing, compared with 23.26 minutes for those with immediate generation. They then completed the passage judgments in 17.41 minutes, compared with 22.21 minutes. The researchers’ comparisons found significantly longer writing and shorter evaluation.

The combined averages were almost identical, about 45.4 minutes. No significant total-time difference was detected. That does not prove the conditions take equal time, and it is no demonstrated net saving. The clearest result concerns where the minutes went.

Finding a basis for judgment

The participant asking about sleep diaries shows one activity that could occupy the earlier minutes. A source’s conclusion becomes easier to assess when its method is visible. A finding based on reported sleep and one based on a tracking device rest on different measurements.

Locating the method does not settle what it is worth. The original Swedish trial acknowledged that the absence of objective measurements could introduce bias, while also noting prior evidence for the reliability of diaries. The participant’s preference for a device was a judgment to consider, not proof that the study lacked evidence.

This is a more specific kind of preparation than getting words onto the page. It gives a writer a basis for asking whether a later passage describes the evidence faithfully. The experiment did not establish that this understanding caused the faster judgments, however. More familiarity with the same sources could make evaluation faster without producing a skill that travels to new material.

Outside the experiment, the work after generation also contains jobs that preparation cannot finish in advance.

In a survey of 319 workers who used generative AI, one participant described editing a product-launch blog post to meet marketing guidelines and tone preferences, while making technical details accurate and understandable to its audience. This was a reported experience, not observed work or a test of how many errors the editing caught.

The account changes what extra editing time means. Familiarity with evidence might help a writer check a technical claim. It cannot by itself settle tone, audience comprehension or the fit with company guidelines. Some later work belongs to adapting the actual piece of writing for its destination. It need not be the same understanding postponed until after generation.

Nor did faster judgment in Su’s experiment become a demonstrated improvement in correctness. The study found no significant overall passage-accuracy advantage, or advantage in detecting and correctly classifying flaws across all flawed passages. Speed and successful checking need their own evidence.

The design leaves another question open: which part of the early assistance mattered? Guidance, the progress display, the mode default and the generation restriction came together. The randomized comparison tested the package. Longer task time and more prompts cannot tell us directly how much thought occurred.

The researchers tried to distinguish engagement-responsive access from waiting by giving a later cohort generation on schedules copied from Engage-to-Unlock participants. But a timer defect prevented 14 of those 94 participants from receiving scheduled access, and the cohort had been recruited separately. That comparison cannot cleanly establish that an adaptive gate works better than an equivalent delay.

What the writer needs next

Sometimes generation helps a writer discover what to write.

Alicia Guo and colleagues interviewed 18 creative writers already using AI and studied them working on their own projects. In the recorded writing sessions, one personal-essay writer used ChatGPT for editing, rewording and feedback. Another began a new personal essay by using it to research and explore themes and styles. Both kept the chatbot and writing document in separate tabs.

These selected, established AI users were not taking a source-fidelity test. Their activities show why requiring a position before generative help would fit some assignments more naturally than others. Exploring possibilities may be the purpose of opening the tool.

Time available can change the tradeoff, too. In a separate preregistered experiment with 393 participants, people wrote a document-based argument about a civic decision under 10- or 30-minute limits. With only 10 minutes, early or continuous AI access produced significantly higher essay scores than late access. Under 30 minutes, the essay-score pattern favored beginning independently, but those pairwise differences were not significant.

That experiment did not test Engage-to-Unlock. It does show why delay cannot be treated as a virtue independent of the deadline. An opportunity to read and formulate an argument has to fit within the work someone has time to finish.

The participant who questioned the sleep measurement eventually unlocked generation, selected that mode and submitted no prompts in it. The paper reports a final essay of 218 words. Their later passage judgments are not provided, so this sequence cannot explain the group result.

What the conversation preserves is the work behind one possible sentence. Before generation became available, the writer had found that the sleep claim came from diaries and had formed a preference about measurement. A later passage would still need to say what those diaries recorded, rather than quietly turn them into readings from a device.

Sources