Start by asking experts what makes a study feel untrustworthy. Then turn those specific problems into failure modes you can monitor over time.
The hardest part of evaluating an AI user research agent like atypica is not finding a stronger model to grade its work.
It is knowing what to evaluate in the first place.
In one beverage brand study, atypica recruited eight participants: a software engineer early in his career, a café owner, a fitness coach, a young mother, and others with different ages, jobs, and lifestyles.
Yet when asked about the brand, they all gave essentially the same answer:
“It feels like something my dad’s generation would drink.”
In another study, the final report said:
“One 26-year-old participant said she would reward herself with a drink after working late.”
It sounded natural. It even looked like a useful insight. The client highlighted it.
But when we went back through the interview transcripts, we found that nobody had said it.
Neither failure was on our original testing checklist. We only discovered them after the agent had completed real studies and experts had reviewed the results.
So instead of trying to design an exhaustive evaluation framework upfront, we began with a simpler question:
What, specifically, makes you distrust this study?
We then turned recurring answers into defined failure modes and asked an LLM to monitor those known problems at scale.
That became the starting point for Agent Evals at atypica.
An AI user research project is more than its final report.
At atypica, a study moves through several stages. Each stage can appear reasonable on its own while still contributing to an unreliable result.
| Research stage | What can go wrong |
|---|---|
| Understanding the brief | The agent mistakes an example for the actual research subject |
| Planning the study | The research scope does not match the user’s real question |
| Finding participants | The selected audience lacks meaningful diversity |
| Building personas | The personas have different profiles but identical perspectives and reasoning |
| Conducting interviews | The interviewer asks about attitudes but fails to probe behavior and motivation |
| Synthesizing findings | A minority opinion is presented as a common need |
| Writing the report | Evidence is exaggerated, or quotations are invented |
If we assign only one score to the final report, we may learn that the result is poor. We still will not know where the problem began or what the team should change.
Evaluating a complex agent therefore requires more than checking its final output. Feedback needs to point back to the specific agent step where the failure occurred.
There is also an important boundary to clarify.
Whether companies should trust synthetic consumers involves questions about data sources, simulation fidelity, differences from general-purpose language models, and data security. We address those questions in the Synthetic Consumer Trust Map.
This article focuses on the next layer: even if the underlying approach is sound, an individual study can still fail during execution. Agent Evals help us detect those failures and prevent known problems from returning after a product change.
We added a simple feedback entry point beside every agent step in the research replay.
When experts noticed a problem, they could describe it in their own words:
The eight participants have different jobs, but their opinions, reasoning, and language are almost identical.
We did not begin with a complete list of predefined issues. We also did not ask experts to rate each step from 1 to 5.
A score of 2 for interview quality tells us that something went wrong. It does not tell the team what went wrong.
The following feedback is much more actionable:
At this stage, experts do not need to use consistent terminology. They only need to answer two questions:
If the feedback is vague—“the personas do not feel real” or “the report is not professional enough”—ask one more question:
If you could point to only one specific problem, what would it be?
The second answer is usually much more useful for evaluation.
Experts may describe the same problem in several ways:
These observations all describe the same underlying failure.
We call a problem that can be identified repeatedly across different studies a failure mode.
For example:
Participant response convergence: Participants with different backgrounds show highly similar opinions, reasoning, experiences, or language, without reflecting the differences implied by their backgrounds.
An LLM can help organize the feedback, but it should not make the final decision:
Each failure mode should answer at least three questions:
| What to define | Example |
|---|---|
| What counts as a failure? | Multiple participants reach the same conclusion for the same reason |
| What does not count? | Participants reach the same conclusion through clearly different experiences and decision processes |
| What evidence is required? | Compare direct excerpts from at least two participants |
Suppose eight participants all dislike the same product. That does not necessarily mean the personas have converged.
One person may find it too expensive. Another may worry about sugar. A third may be influenced by family habits, while a fourth cannot imagine a relevant consumption occasion.
Their conclusions are the same, but their experiences and reasoning are different. They still sound like different people.
The real warning sign is when the conclusion, reasoning, experience, and even wording are interchangeable.
After our first round of analysis, we had failure modes such as these:
| Failure mode | What it means |
|---|---|
| Misunderstood research intent | The agent treats the user’s example as the research subject |
| Participant response convergence | Different personas give the same opinions and reasoning |
| Shallow interviewing | The interviewer asks about attitudes but not behavior or motivation |
| Evidence amplification | The report turns mild feedback into a strong claim |
| Fabricated quotation | The report includes a quotation that never appeared in the research |
| Overstated coverage | The agent claims to have completed research it did not actually perform |
Unlike labels such as “poor research quality,” each failure mode points to a specific behavior in a specific part of the workflow. The team knows where to investigate and can discuss what needs to change.
If you do not define every failure mode upfront, an obvious question follows:
How long should you keep collecting feedback?
We track the number of new failure modes discovered each day, not the cumulative total.
When atypica first reviewed studies intensively, experts identified 14 previously unseen failure modes on the first day. They found 4 on the second day, 4 on the third, and 3 on the fourth. By the end of the month, new expert feedback rarely produced a new failure mode.
The cumulative number can only increase. What matters is whether the number of new failure modes is declining.
When most new feedback fits an existing category, you have probably seen the main problems in that batch of studies. You can begin automating checks for those known failures.
This does not mean no new problems will ever appear. It simply means that the current discovery round can pause.
Experts are good at noticing new problems, but they cannot review every study the agent runs.
During atypica’s first evaluation round, experts submitted fewer than 300 pieces of feedback. Over the following two months, an LLM used the resulting failure-mode list to check more than 1,500 studies.
The division of labor is straightforward:
People discover and define problems. LLMs check known problems at scale.
Do not ask the LLM:
Was this study good?
The question is too broad. The model can easily produce a plausible evaluation that is difficult to verify.
Ask a narrower question instead:
Do the participant responses in this study show convergence?
Or:
Does the report present anything that was never collected during the study as a direct participant quotation?
Each evaluation task should check one failure mode and preserve the supporting evidence.
For participant response convergence, a useful result might look like this:
| Field | Example |
|---|---|
| Decision | Failure detected |
| Observation | Several participants describe the brand as something for their parents’ generation and give similar age-related reasons |
| Evidence 1 | P02: “It feels like something my dad’s generation would drink.” |
| Evidence 2 | P05: “This is something my parents’ generation would drink.” |
| Human review needed | Confirm whether the target audience genuinely shares a strong age-related perception of the brand |
This is less useful:
The personas lack sufficient differentiation. Consider making them more distinctive.
It gives a conclusion without evidence, so a PM cannot verify whether the Judge understood the study correctly.
Here is a simplified prompt you can adapt:
One boundary is essential:
The Judge monitors existing failure modes. It does not discover new ones.
If the Judge checks participant convergence every day, it will naturally produce many records about participant convergence. Feeding those records back into failure discovery would only amplify the system’s own prior assumptions.
New failure modes should still come from expert feedback. They may also surface through user behavior or unusual system signals, but a person needs to return to the study and confirm what actually went wrong.
A Judge that can run at scale is not necessarily a Judge you can trust.
Before putting its results on a dashboard, prepare a set of study records that experts have already labeled. Then compare the Judge’s decisions with the expert decisions.
This is usually called a golden test set.
It should contain both positive and negative examples.
For participant response convergence, that might include:
Do not look only at overall accuracy. Separate at least two kinds of mistakes:
The acceptable trade-off depends on the failure mode.
Missing a fabricated quotation may allow false evidence to reach a customer. But an overly sensitive convergence check may mistake genuine audience consensus for a persona failure.
A higher overall score is not automatically better. PMs need to understand where the Judge fails and what those mistakes mean for the product.
When the Judge performs poorly, do not immediately add more prompt rules. First ask:
Sometimes a weak Judge reveals that the team itself has not agreed on what counts as a failure.
The golden test set validates the Judge, not the agent.
Historical studies have already happened. Changing the agent’s prompt or workflow will not change their outputs.
To determine whether an agent update improved the product, save a separate set of fixed research inputs. Run the same inputs through the old and new agent versions, then use a validated Judge to compare their outputs.
We compare versions using the occurrence rate:
Occurrence rate = number of studies containing the failure ÷ total number of studies run during the period
Suppose we run the same research briefs through Agent v1 and Agent v2:
| Failure mode | Agent v1 | Agent v2 |
|---|---|---|
| Participant response convergence | 40% | 15% |
| Shallow interviewing | 25% | 15% |
| Evidence amplification | 10% | 10% |
| Fabricated quotation | 5% | 15% |
Agent v2 improved persona differentiation and interview depth, but fabricated quotations became more common.
We should not average all four dimensions into one score and declare v2 better.
A single change can fix one problem while making another worse. Regression evaluation exists to expose those trade-offs before release.
The distinction between the two datasets is simple:
The golden test set validates the Judge. Fixed inputs compare the agent.
| Dataset | What it evaluates | Question it answers |
|---|---|---|
| Judge golden test set | LLM Judge | Does its judgment align with expert judgment? |
| Agent fixed-input set | Agent | Is the new version better than the previous one? |
After release, the evaluation system needs two parallel feedback loops:
LLMs continuously monitor known failures. People continue discovering new ones.
For established failure modes, track occurrence rates and trends. If one rises for several days, open the relevant studies, inspect the evidence cited by the Judge, and trace the problem back to the earliest agent step where it appeared.
But do not remove the expert feedback entry point.
After atypica completed its first round of failure-mode discovery, new problems continued to appear. One later failure mode was overstated coverage: the agent promised to research eight platforms but analyzed only three.
An existing Judge could not have discovered it because the failure had never been defined.
Agent Evals are therefore not a permanent checklist. They are a continuing loop:
Real runs expose problems
→ People describe them clearly
→ The team defines failure modes
→ The Judge monitors them
→ Product changes reveal new problems
Not every failure mode can be fixed by changing a prompt.
Suppose the user has not provided enough detail. Should the agent make a reasonable assumption and continue, or pause and ask another question?
Continuing is faster, but it may lead the study in the wrong direction. Asking is safer, but it may make the product feel slow and repetitive.
If the team has not decided how proactive the agent should be, it cannot consistently label either behavior as correct or incorrect. The Judge will not have a stable standard either.
Evals sometimes reveal a bug. Other times, they reveal a product decision the team has not yet made.
That is why PMs need to participate in Agent Evals. Evaluation does not only test model performance. It forces the team to define how the product should behave.
You do not need a complete evaluation platform to begin.
Choose one problem that genuinely damages user trust, such as participant response convergence, and take it through a small end-to-end loop:
You can use sample data to understand the process or build a quick demo.
But to discover the failures that matter in your own product, you eventually need real runs. Demo data is usually too clean. The most important failures often appear when users interact with the product in ways the team did not anticipate.
Do not begin with a composite score, dozens of failure modes, or a complex dashboard.
Start by defining one meaningful problem clearly enough that both experts and the Judge can identify it consistently. That is already a useful foundation for Agent Evals.
The first two problems atypica discovered were simple to describe: eight participants sounded like the same person, and a report included a quotation that nobody had actually said.
Before experts pointed them out, we did not even know they belonged on the evaluation checklist.
Now these problems have names, boundaries, and evidence requirements. An LLM can monitor them continuously. When the agent changes, we can rerun the same inputs and see whether a problem has actually decreased—or simply moved to another part of the workflow.
That is why Agent Evals should not begin with an exhaustive list of metrics.
A more useful starting point is to open a real run, find the part that makes you hesitate, and ask:
What exactly went wrong here?
How can we catch it earlier next time?
Review the following agent study for one specified failure mode.Failure mode: Participant response convergenceDefinition:Participants with different backgrounds show highly similar opinions,reasoning, experiences, or language without reflecting meaningfuldifferences in their backgrounds.Requirements:- Check only this failure mode. Do not evaluate other aspects of the study.- Describe what you observed before giving a decision.- If you detect a failure, quote at least two participants as evidence.- Similar conclusions do not count as a failure when the underlying experiences and reasoning are clearly different.- If the evidence is insufficient, return "No failure detected."- Do not invent new failure modes.Return:1. Failure detected / No failure detected2. Observation3. Direct evidence4. Anything that requires human reviewStudy record:[Paste the study record here]