Skip to content

Before an AI Interviewer Meets People

In a study of five AI interviewer designs on three topics, simulations adjusted using human data tracked differences in the information elicited and the interviewers' behavior, with weaker agreement on people's own ratings of comfort and experience.

Klaus had another point to make. In a Swedish university study published in May, the social scientist, identified by a pseudonym, was interviewed about using generative AI in research. He described sending an answer to the AI interviewer and wanting to add a second thought. The chatbot had already replied with its next question. As Klaus put it in later feedback, “I was intellectually still at the previous question.”

Could researchers rehearse an AI interviewer’s questioning before recruiting people? An October 8 preprint offers encouraging evidence. Its authors compared five interviewer designs across three topics, in simulated conversations and in studies involving 450 people. Simulated-study scores closely tracked human-study scores on several measures of the information elicited and the interviewers’ behavior. The simulations were adjusted using data from the human studies, so the result concerns those settings rather than a forecast made without human data.

For someone building an AI system that asks people about their work or experiences, a rehearsal could help compare ways of asking, following up and changing topics before a trial with people. The tested simulator gives artificial participants information to disclose, then evaluates how an interviewer draws it out. Its scores do not establish what an unstudied population actually thinks. Agreement was also weaker for people’s comfort and overall experience, judgments that still need the people answering.

Something to draw out

The new environment, InterviewPlayground, gives each simulated participant 200 generated memories. Most are background material. Just 6 are built around insights taken from previously published interview studies on the relevant topic. The interviewer gets a guide to the subjects it should explore, but has to elicit answers through conversation.

Participants differ in how much they know, how well they understand questions, how readily they reflect, remember, elaborate and disclose sensitive information. For each question, the simulator retrieves 5 memories and generates an answer using them, the participant’s assigned characteristics and the conversation so far. Access to the supplied material depends on the exchange, rather than every interviewer receiving the same complete response.

That makes a follow-up consequential: it can seek detail in a short answer, while a change of topic can explore another part of the guide. Different interviewers can reveal different amounts of information already placed in the simulation. Learning about new customer needs or employees’ actual experiences would require evidence from those people.

The tested topics covered weight management, Asian American identity and politics, and generative AI in knowledge work. All five interviewer designs used the same underlying model, GPT-5.4 mini. The comparison concerns different ways of directing that model’s interviewing, rather than a contest among foundation models.

Do the differences carry across?

Each interviewer conducted human studies with 30 participants per topic, for up to 30 minutes per interview. The simulated studies reused the same 30 artificial participants for each topic across the five designs. Researchers compared the human and simulated averages for each interviewer–topic combination: 15 paired study results for each measure.

Across 12 measures of responses, interviewer behavior, experience and conversation length, the reported average correlation was .86. Correlation describes how closely the study averages varied together, not the percentage of correct answers. Agreement was stronger for covering the interview guide, at .95, than for the amount of relevant material elicited, at .74. The latter counts response length weighted by relevance to the research question. Neither score independently establishes the truth or usefulness of the information.

A concrete questioning pattern appeared in both kinds of conversation. One design, LLMRoleplay, elicited the most relevant material and covered the most guide topics in both conditions. In sampled conversations, the authors found specific follow-ups, requests for missing detail and clear transitions. A human participant describing older and younger Asian Americans’ political views received a request for examples of politicians, parties or positions. A simulated participant received a similar request for a concrete difference in views.

The observations do not isolate which behavior caused the score advantage. They give the evaluator something to examine beyond a summary number: a way of seeking detail that appeared with both human and simulated participants. Rehearsal also exposed a recognizable failure. Among the authors’ selected lowest-rated interviews, repetitive questioning was the most common identified problem in both conditions. That selection cannot establish how often repetition occurred across all interviews.

The positive result comes with limits on what a team can carry into its next project. The simulated population’s traits and response timing were fitted using the human studies. A common AI judge scored the content and behavior of both kinds of transcript. The authors checked several measures against expert judgments, but shared scoring can still contribute to agreement.

There is also a difference between comparing all 15 study averages and choosing an interviewer for one topic. Scores pooled across topics do not establish equally reliable choices within each topic. For a new population or interviewing setup, the study supplies a reason to try rehearsal, with a human comparison still needed to learn whether its results carry across.

What the answer leaves out

The weaker correlations concerned the encounter itself. Simulated comfort and overall experience were judged from transcripts, while people rated their own experience afterward. Correlations were .60 for comfort and .67 for overall experience, with lower simulated ratings on both. The estimates were imprecise: the reported 95% confidence interval for comfort ran from .13 to .85.

In the authors’ selected low-comfort interviews, objections reached the interviewer differently. People often used a report button or later survey notes to express discomfort. Simulated participants more directly said they did not want to discuss something, giving the interviewer a cue to move on. This analysis concerns a selected subset, not a finding that people generally conceal discomfort. A conversation can leave out something consequential about how it felt.

The older Swedish study also contained a different response from Klaus’s. Thomas, another pseudonymous participant, said that knowing he was answering a chatbot made honesty easier. These accounts concern a text interviewer built with Gemini 1.5 Flash, not the five designs in the new study. They document variation in answering questions, without independently validating InterviewPlayground.

For your own interviewer, a useful question is: which judgment about this conversation would still need to come from the person answering? That is a suggested evaluation check, not a remedy tested by either study. Rehearsal can supply conversations to inspect and questioning strategies to compare. A human trial can also let someone explain that the previous answer was unfinished.

Sources