Designing for Agents
May 25, 2026
PUBLISHED
TOPICS

Most of the conversation about designing for agents focuses on the interface. How should a person talk to an agent? What happens when the agent gets confused? How do you handle errors, handoffs, confirmation steps? These are real questions and they matter. They are also, in the framework I’ve been building in this series, Recognition-level problems. Taste and judgment applied to a new interaction surface.

The harder problem sits above that. When agents are doing the building, the generating, the optimizing, the publishing, the question changes. You stop designing what agents look like and start designing what agents believe. What they prioritize. What they refuse to do. How they maintain coherence with each other and with your brand when nobody is watching.

That’s a Commitment problem. And almost nobody is working on it yet.

The interface is the easy part

I don’t mean easy in the dismissive sense. Agent interaction design has its own craft. The conversational patterns, the way you surface and preview changes before applying them, the trust signals that let someone know whether to accept or override the agent’s suggestion: these require real skill to get right.

But they’re solvable with the tools we already have. They’re extensions of interface design patterns we’ve been refining for decades: progressive disclosure, confirmation dialogs, undo, transparency about what’s happening under the hood. The inputs are different (natural language instead of clicks), but the underlying principles of helping someone feel in control of a system are the same ones we’ve always used.

The unsolved problem is what happens when the human steps away. When the agent is generating a page section at 2am based on your design system. When it’s writing content that will publish without a review step. When it’s deciding which personalization variant to serve to which audience segment. When multiple agents are operating across different parts of the same product, each optimizing for their own objective, and nobody is asking whether the total experience still makes sense.

Recognition becomes the evaluation layer

In the earlier posts I described Recognition as the ability to see quality and make good decisions in the moment. In an agent context, Recognition translates to a systems problem: can your infrastructure assess whether agent output meets your standard?

Today, most agent workflows rely on a human reviewing the output before it ships. Someone looks at the generated page, the suggested copy, the proposed layout, and decides whether it’s good enough. That works at low volume. It falls apart when agents are producing at the speed and scale they’re designed for.

The question becomes whether you can encode your Recognition into something that operates without a person in the loop. Can your design system evaluate whether a generated section is consistent with your patterns? Can your brand guidelines flag when generated copy drifts from your voice? Can your interaction standards catch when an agent’s suggestion breaks a convention that matters?

Some of this is already happening. Design systems that constrain generation to a defined set of components and styles. Brand voice parameters that shape what the AI produces. Content governance rules that prevent certain categories of output from publishing without review.

But most of it is primitive. The design system constrains visual consistency. It doesn’t evaluate whether the experience is coherent. The brand voice parameters keep the tone in range. They don’t catch when the message contradicts something said on another page. The governance rules catch obvious violations. They miss the subtle ones, the kind where the output is technically within bounds but feels off in a way that erodes trust over time.

The gap between what automated evaluation can catch and what a skilled human would notice is where the real work is. Closing that gap requires encoding Recognition into systems with more depth than most teams have attempted. It means going beyond “is this on-brand?” to “is this true to what we’re trying to create?” That’s a much harder question to formalize, and the teams that figure it out will have a meaningful advantage.

Commitment becomes the governance layer

If Recognition is about evaluating individual outputs, Commitment is about the stance that governs all of them. The durable point of view about what your product should be, applied to systems that operate without constant human oversight.

In practice, this breaks into questions that most organizations haven’t answered yet.

What do your agents optimize for? Every agent has an objective function, whether or not anyone made it explicit. A personalization agent might optimize for conversion. A content agent might optimize for engagement. A support agent might optimize for resolution speed. Each of those is reasonable in isolation. The problem shows up when several agents operate across the same product, each pushing toward their own objective, and the cumulative effect is an experience that feels incoherent or extractive. A site that simultaneously tries to convert you, engage you, and resolve your issues can feel like talking to three people with different agendas.

Commitment means defining what the total experience should feel like and constraining each agent’s behavior accordingly, even when that means accepting lower performance on any individual metric. This is the same tradeoff design leaders have always navigated: the tension between local optimization and global coherence. Agents make it harder because the volume is higher and the feedback loops are faster.

What do your agents refuse to do? This is the question that matters most and gets the least attention. Every product has implicit lines it won’t cross: content it won’t generate, actions it won’t take, experiences it won’t serve. In a human-driven workflow, those lines are maintained by the people doing the work. They know the norms. They feel the discomfort when something crosses a boundary. They self-correct.

Agents don’t feel discomfort. They need to be told where the lines are, in ways specific enough to be enforceable and flexible enough to handle situations nobody anticipated. This is governance in the deepest sense: deciding what your product stands for and making that stance hold even when the system is operating autonomously.

Most agent governance right now is about preventing obvious harms: blocking offensive content, avoiding legal liability, filtering for brand safety at a surface level. The real governance question is subtler: will your agent recommend a practice you disagree with? Will it sacrifice long-term user trust for short-term conversion? Will it produce work that is technically within your guidelines but philosophically at odds with what you’re trying to build?

The organizations that take this seriously will be the ones that can articulate their product philosophy clearly enough to operationalize it as agent behavior. The ones that can’t will discover, gradually, that their agents have developed a de facto philosophy on their own, assembled from whatever optimization signals were loudest. And that de facto philosophy will probably look like the ambient mediocrity discussed in Post 0, scaled up and running on autopilot.

Coherence across agents

There’s a layer beyond individual agent governance that very few people are thinking about yet: how multiple agents maintain coherence with each other.

Consider a site with an AI-powered content engine, a personalization system, a localization workflow, and an SEO optimization agent. Each operates on the same underlying content. Each has access to the same design system. Each has its own parameters and objectives. None of them is responsible for the experience as a whole.

A person visiting that site encounters a single experience. They don’t know or care that four systems produced it. They just know whether it feels considered or scattered, whether the site seems to know what it is or seems confused about what it’s trying to do.

This is a sensibility problem at system scale. The coherence that used to come from a small team with shared instincts about the product now has to come from something more durable: an explicit product philosophy encoded into the infrastructure that all agents operate within. Design systems are part of that infrastructure. Brand guidelines are part of it. But neither is sufficient on its own, because coherence across agents requires agreement about behavior, priorities, and tradeoffs, not just visual consistency.

What this means for design teams

The job is expanding again.

For the last decade, the scope of design leadership grew from “what it looks like” to “how it works” to “how the whole system behaves.” Now it’s growing to include how the system behaves when humans aren’t directly controlling it.

That means design leaders need to think about agent behavior as a design surface. The parameters you give an agent, the constraints you set, the objectives you choose, the lines you draw: these are design decisions with the same weight as choosing a layout or an interaction pattern. Possibly more, because they operate at higher volume and lower visibility.

It also means the infrastructure investments that felt optional before, the design system, the documented brand philosophy, the interaction standards, the governance frameworks, become load-bearing. When humans were the bottleneck, you could get by with informal shared understanding. When agents are producing at scale, informal understanding breaks down. You need the explicit version.

The organizations that built conviction into infrastructure, that articulated what they stood for before they automated, are the ones that will navigate this transition with their product’s character intact. The ones that didn’t will spend the next several years trying to wrest coherence back from systems that were never given a philosophy to follow.

The design problem underneath

I’ve been arguing throughout this series that the principles of design didn’t change with AI. Purpose. Form and function. Saying no. Building systems that hold. The agent era doesn’t change those principles either. It applies them to a context where the stakes are higher and the feedback loops are faster.

The products that will feel alive in five years, the ones that reward attention and earn trust and hold together across every touchpoint, will be the ones where someone decided what the product should believe before they gave agents the power to build it. Where the design philosophy was clear enough and deep enough that it could govern systems operating at a speed and scale no human team could manage alone.

The question hasn’t changed. It’s just gotten louder: What do we actually believe about how this should work, and are we willing to hold that line even when the builders aren’t human?

Why I'm still optimistic

The people I know who care about this work, who have opinions about how products should feel and the patience to hold those opinions under pressure, have spent most of their careers constrained by execution capacity. They could see the right answer and couldn't get there in time. They knew what the product needed and watched it get scoped out. They held standards that the calendar made impossible to maintain.

Those constraints are lifting. The tools are finally catching up to the ambition. And the people who did the hard, unglamorous work of building depth, developing taste through years of attention, earning judgment through real decisions with real consequences, encoding their sensibility into systems that hold, committing to a direction and staying with it, those people are about to find out what they can do without the old limitations.

That's not a guarantee. The tools amplify whatever's already there, including emptiness. But for the people and teams who built something real underneath, this is the most interesting moment to be doing this work that I can remember.

The principles haven't changed. The opportunity to live up to them has.