Agentic AI changes the role of the user — and UX Research decides who survives
Agentic AI is changing something that sounds simple but carries enormous weight: the role of the user.
For decades, we built tools that waited for commands. You clicked, typed, directed every step. Now, systems act on their own — they assess a goal, break it into steps, and execute without asking permission.
This shift transforms the user from operator to delegator. And delegating is a completely different cognitive process from doing. It’s not just a technical change — it’s a change in the relationship between people and systems, and it touches something much deeper: trust.
The numbers already show the movement is real. According to the McKinsey Global Survey on AI 2025, 62% of organizations are already experimenting with agents at some level. But Gartner projects that more than 40% of those projects will be canceled by 2027 — and the reason is almost never technical. 😬
The problem lies somewhere else: people don’t understand what the agent did, don’t trust the result, or simply feel like they’ve lost control of their own work. We have plenty of benchmarks to measure model accuracy, but almost nothing that measures why real users abandon agents. That data gap is already, in itself, part of the problem that UX Research needs to solve — not as a product detail, but as the piece that can decide whether a system survives or ends up in the trash.
What changes when the system acts on its own
When a user clicks a button, they know what’s going to happen. When an agent makes a sequence of decisions on its own, that implicit contract of predictability disappears. The system may have done everything right — but if the user can’t follow the reasoning behind the actions, the feeling left behind is one of unease, not efficiency.
This has a name in cognitive psychology: it’s a broken mental model. And when the mental model fails, trust goes with it.
This phenomenon is even more intense in professional contexts. Imagine an agent that reorganizes your calendar, responds to emails on your behalf, and prioritizes tasks without you having asked for each of those steps individually. Technically, it might be flawless. But the emotional experience of seeing something act in your place — especially on decisions you’ve always considered yours — triggers a discomfort that isn’t irrational. It’s human.
Email, by the way, is a great example for understanding this dynamic. Most users are perfectly fine letting an agent organize their inbox. But ask if they’d accept the agent sending a reply on their behalf, and the reaction changes completely. That threshold varies for every person, every task, and every industry. And it’s exactly the kind of thing that UX Research needs to investigate before the product hits the market.
The central question isn’t whether the agent is capable. It’s whether the user believes it’s capable — and more than that, whether the user feels safe letting the agent act. These are questions no technical metric can answer. They require qualitative research, in-depth interviews, analysis of real behavior, and active listening about what people feel when they interact with systems that take initiative.
The challenge of designing for predictable unpredictability
Agentic AI challenges a basic UX principle: these systems are, in theory, non-deterministic. Give an agent the same task on Monday and then on Friday, and it might solve it differently both times. For professionals trained to expect that identical inputs produce identical outputs, this variation creates real design tension.
The NN/g research agenda for 2025 on generative AI frames the question well: how do you evaluate a system that changes over time?
Since it’s not possible to guarantee consistency of action, the way forward is to guarantee consistency of intent. This reframing has practical implications for measurement. Instead of tracking whether the agent completed a task the same way twice, the focus shifts to studying whether users believe the agent understood what they wanted. In practice, this can be operationalized through:
- Post-task trust assessments on goal alignment
- Think-aloud protocols, where users narrate expected versus actual outcomes
- Comparative studies between different agent paths, measuring perceived alignment
The metric becomes perceived alignment, not behavioral consistency.
In an actual evaluation, 10 participants were given the same four tasks to complete with a voice agent. Researchers tracked how the agent’s tone, language, and level of detail varied across sessions. Some users received concise and direct responses, while others heard longer and more conversational explanations for the same question. In one case, the agent hallucinated a detail — and the user’s trust dropped immediately, not just for that task, but for all subsequent ones.
What was being measured wasn’t whether the agent said the same words every time. It was whether users believed it understood their goal — and how quickly that belief crumbled when something went wrong. 🧠
Trust isn’t a feature, it’s a construct
One of the most common mistakes in developing Agentic AI systems is treating trust as something the user will naturally develop over time with use. That logic works for simple tools. For autonomous agents, it’s dangerous. Trust in systems that act independently needs to be actively built — through transparency, clear feedback, and mechanisms that give the user back the feeling that they’re still in control, even when delegating.
Recent research in Human-Computer Interaction shows that users develop trust in agents when three elements are present: predictability — they can anticipate what the agent will do; explainability — they understand why the agent did what it did; and reversibility — they know they can undo or adjust any action. When all three elements are absent, even a technically flawless agent generates rejection.
And rejection, in this context, rarely comes with constructive feedback — it comes in the form of silent abandonment.
The role of UX Research in this scenario is precisely to map where these elements are failing in the real experience. Not the idealized experience imagined by the product team, but the lived experience of people who use the system every day, under pressure, in a hurry, within a context no lab can fully replicate. This requires going into the field, observing, asking, and above all, listening to what’s between the lines — because users rarely say they don’t trust a system. They simply stop using it.
From isolated interactions to long-term relationships
Most digital products are transactional. You search for a flight, buy it, move on with your life. Agentic AI is more like onboarding a new team member. The agent adapts to your preferences. You learn its tendencies. The relationship evolves over weeks.
Traditional research methods weren’t designed for this time scale. Anyone who’s studied automation in aviation recognizes what’s happening here: monitoring fatigue. The more reliable a system becomes, the worse humans get at catching its errors. Agentic AI adds an extra layer to this phenomenon. Unlike autopilot, agents don’t follow the same procedure every time. Users aren’t just monitoring — they’re supervising something unpredictable while deciding how much supervision it actually needs.
In research with real users, it’s common to see participants swing between over-trusting and anxiously double-checking everything within the same session. It’s a genuine dilemma: you can’t blindly accept every output, but if you’re reviewing every action the agent takes, it’s not saving you any time.
The fundamental research question becomes: what kinds of transparency signals — like citing sources or explaining reasoning — make users comfortable enough to reduce oversight? Evidence from studies using trust frameworks with voice agents shows that the patterns hold regardless of modality: context awareness, clear status indicators, and graceful error recovery are what separate agents people use once from agents they actually rely on.
Delegation as a design experience
Delegating to a system isn’t the same as clicking confirm. When you delegate a task to a person, you consider their track record, the trust you’ve built over time, the ability to ask questions, and the option to adjust course. With an agent, this process needs to be intentionally replicated within the interface and the system’s behavior. It doesn’t happen by accident — it’s the result of very specific design decisions, guided by real research on how people make delegation decisions.
One of the most relevant discoveries in this area is that users tend to delegate more when they perceive the agent communicates uncertainty honestly. In other words, an agent that says something like I’m not sure about this step, would you prefer I continue or wait for your confirmation? generates more trust than an agent that acts with full autonomy even in ambiguous situations. This perception of system humility — the ability to acknowledge limits — is a far more powerful trust factor than a raw display of technical capability.
This has direct implications for designing interactions in agentic systems. Confirmation screens, action summaries, activity logs, deviation alerts — none of this is interface bureaucracy. It’s the language the system uses to build or destroy the relationship with the user over time. And each of these decisions needs to be validated through research, tested with real users, and refined based on observed behavior, not assumptions about what would be more convenient or visually elegant.
Four research methods that work in the agentic world
Four approaches have consistently proven their value in this new era. And there’s a tension running through all of them: Agentic AI teams ship fast. The agent might get updated three times during a four-week study. Every method needs to generate actionable signal at speed.
Discovery research for building from scratch
This stage happens before any line of code — and it’s the one teams skip the most. The question isn’t how should this agent work. The question is: should this agent exist? Who’s the real audience? What are these people struggling with? Do they even want an AI handling this task?
We’ve seen what happens when teams skip this phase. One team built a capable agent for a workflow where users didn’t want any autonomy — they wanted better tools. Contextual inquiry and concept validation interviews would have revealed that. Two weeks of research can save months of unnecessary engineering.
Research-driven AI evaluations
Most evaluation frameworks test accuracy — whether the agent got the answer right. That question matters, but it’s a terrible predictor of retention. Without human-centered criteria, agents can produce what’s already being called agent slop — low-quality output at scale, generated by systems without adequate guardrails.
Research-driven evaluations add the human dimension. Was the tone appropriate? Did the explanation make sense to a non-technical user? Did the agent push forward when it should have paused for confirmation? These evaluations need to run continuously, with user feedback feeding into model tuning and guardrails. Without that loop, you end up with a green dashboard and a product nobody comes back to.
Longitudinal diary studies
A usability test captures a moment. A diary study captures the full arc: the initial excitement, followed by the realization that the agent gets things wrong sometimes — and gets them wrong confidently. Daily entries reveal precise inflection points: when trust eroded, what triggered it, and whether recovery was possible.
Co-design workshops
This approach brings users and product teams together to define authority boundaries as a team. Where can the agent act independently? Where does it need explicit permission? One framework being tested is the defer, suggest, lead model. It’s still evolving, but the pattern holds: for high-risk decisions, the agent presents data and waits. For medium risk, it offers ranked suggestions. For routine tasks, it executes and reports afterward. What makes this effective isn’t the framework itself — it’s that users participated in drawing those lines.
The human benchmark — and the accessibility question
Engineering teams prioritize MMLU scores and accuracy rates. These indicators matter for model performance, but they’re terrible predictors of whether someone keeps using an agent after the first week. We’ve seen agents with strong benchmarks get abandoned because the interaction felt opaque or presumptuous.
There’s also a dimension that most teams haven’t even begun to address: accessibility. How does a screen reader user supervise an agent making autonomous actions in real time? How does someone with a cognitive disability manage the delegation decisions described above? These aren’t edge cases. If we’re not researching these questions now, we’re building technology that works for some people and excludes others by design.
According to Gartner, 15% of daily work decisions will be made autonomously by agents by 2028, up from zero in 2024. The technical infrastructure is moving fast. But none of this works if people don’t trust the system, and none of it is fair if it only works for some of them.
What UX Research needs to investigate right now
With the rapid growth of Agentic AI systems, UX Research needs to broaden its scope of investigation in very practical ways. Understanding how a user navigates an interface isn’t enough anymore — it’s necessary to understand how they think about delegation, where their comfort limits are, which types of decisions they would never delegate to a system, and why. These answers vary widely depending on context, user profile, and the industry where the agent operates, making contextual research absolutely essential.
Some of the topics gaining traction in research agendas of product teams working with agents include:
- Mental model mapping: how the user imagines the agent works, even if that picture is imprecise or incorrect.
- Delegation boundaries: which tasks the user is willing to delegate and which they feel must remain under their direct control.
- Reactions to autonomous errors: how the user responds when the agent does something wrong without asking permission — and how much that affects future trust.
- Transparency perception: whether the system’s explanation mechanisms are actually understood or simply ignored.
- Supervision fatigue: how much attention the user can sustain over the agent’s actions before they start trusting blindly — which is also a risk.
Each of these points represents an area where the product can fail silently and progressively. And most importantly: these are areas that only surface through quality research, conducted with rigor and with genuine openness to what the user has to say — even when what they say contradicts the product team’s hypotheses.
The moment we’re living through with Agentic AI is rare. It’s one of those windows where technology is advancing faster than human understanding of how to use it well. The agents that will thrive won’t necessarily be the most capable ones. They’ll be the ones people trust enough to rely on. And it’s exactly in moments like these that UX Research has the greatest impact — not as validation of what’s already been decided, but as the voice that puts the human at the center before the problems become irreversible. 🚀
