When AI writes the interface, what exactly are designers handing off?
The interface is no longer only what we draw. It is also the behavior we leave behind.
You are specifying
for a model.
For most of my career, design handoff meant something fairly tangible: a Figma file, components, states, interaction notes. But when the system itself starts assembling the answer, the artifact changes.
For most of my career, design handoff meant something fairly tangible.
A Figma file. Components. States. Interaction notes. Maybe a prototype showing how everything should behave.
But while working on AI-native experiences at Nayya, I started running into a different kind of design problem.
The interface wasn't always something we could fully design ahead of time.
Sometimes, the model was creating it.
It was deciding how to explain something. What information to emphasize. How much context to provide. When to ask another question. When to acknowledge uncertainty. And, eventually, when it was appropriate to help someone take action.
That changed what “designing the experience” meant.
The design artifact wasn't always a screen anymore.
Sometimes it was Markdown.
Sometimes it was a response template.
Sometimes it was a set of rules describing what the model should — and should not — say.
And increasingly, I started thinking about those rules as product UI.
The interface became a system that writes
A traditional interface gives designers a lot of control.
We decide what appears first. We establish hierarchy. We determine what happens when someone clicks a button. We can inspect almost every state before shipping it.
Generated interfaces behave differently.
You can design the container, but the experience inside that container may be assembled dynamically.
So the questions I found myself asking started changing.
What should this screen look like?
- How should the answer begin?
- What information should the model use?
- What should it never assume?
- When should it explain its reasoning?
- When should it stop and ask the user instead?
- What evidence should appear with the answer?
Those aren't just content questions.
They're interaction design questions.
Four behaviors started showing up repeatedly
As we worked through benefits guidance experiences, I began separating generated interactions into different levels of agency.
Not because every AI interaction needs a formal taxonomy, but because the distinction forced us to be explicit about what the system was actually allowed to do.
At the lowest level, the system might simply suggest something worth looking at.
For example, it could notice that someone's priorities may have changed since their last enrollment and suggest comparing their existing choice with this year's options.
The important boundary is that the system surfaces an opportunity. It doesn't make the decision.
A second behavior is to ask.
AI systems are very good at filling gaps, which can also make them very good at confidently assuming things we never actually learned.
Sometimes the better experience is simply: What matters most this year?
Before comparing options, clarify whether the person cares more about predictable monthly costs or their expected healthcare needs.
The interaction isn't impressive because the AI knew the answer. It's useful because it knew that it didn't.
Then there's explanation.
In benefits, users often aren't looking for another recommendation. They're trying to understand why two choices that look similar might lead to very different outcomes.
A lower premium, for example, doesn't automatically mean lower overall cost.
The experience needs to make the underlying tradeoff visible: deductible, expected usage, coverage and other relevant factors.
The model can explain the comparison without pretending that the future is certain.
Finally, there is action.
Once an AI system can actually change something on someone's behalf, the design problem changes again.
The proposed action needs to be visible. The destination needs to be visible. The consequences need to be visible.
And consequential actions should have a separate confirmation.
At that point, a confirmation screen isn't merely UI polish. It's part of the system's permission model.
The answer needed its own design file
One of the more practical changes to my workflow was realizing that the Figma file wasn't enough.
We still needed it for layouts, components, states and the surrounding product experience.
But generated experiences needed another artifact.
I started thinking of it as the answer file.
The artifact defined things like hierarchy, response formats, tone, use of personal context, caveats, citations, guardrails and golden examples.
In our case, much of this lived in Notion and Markdown rather than Figma.
And the ordering mattered.
Design file
Layout, components, states. Still necessary for surfaces.
Answer file
Hierarchy, tone, caveats, citations, guardrails, golden examples. Necessary for generation.
Through playtesting, one pattern became particularly clear:
Because technically correct answers could still fail.
A model could provide everything a user needed and still produce a wall of text nobody wanted to read.
It could give the correct recommendation while sounding overly certain.
It could cite the right evidence but bury the actual answer.
The generated output needed information architecture just as much as a screen did.
Brand stopped being something around the answer
This became especially important in a B2B2C product.
When an employer buys a benefits experience for its employees, the generated response isn't simply “AI copy.”
It's part of the product that company chose to put in front of its people.
If the system suddenly sounds generic, salesy or excessively confident, the problem isn't only tone.
The product is behaving differently from the brand that was sold.
That made me think differently about brand systems for AI products.
Traditionally, we might define typography, color, spacing, illustration, voice and tone.
But if the product can continuously generate new language after the designer has left the file, the brand system also needs to define behavior.
How certain are we allowed to sound? How do we communicate risk? When do we cite evidence? When do we ask rather than assume? What kinds of claims should the system never make?
Guardrails aren't something added after the experience is designed.
They are part of the experience.
There's still a piece I haven't solved.
There is one part of this workflow that still feels unresolved to me: design systems.
We've gotten much better at connecting AI tools to tokens, components and code. But giving a model access to a design system and getting it to consistently behave according to that system are two different things.
I've experimented with ways of making components, rules and patterns available to generation. The tooling is getting better quickly.
But I don't think we've reached a point where connecting a plugin, MCP server or component library suddenly means the model understands the design system.
A design system isn't only a collection of reusable components. It's also the judgment behind when and why those components should be used.
That distinction becomes much more visible when the thing consuming your system is a model.
I haven't cracked that problem. But I've encountered the gap enough times that I don't want to pretend the tooling has solved it.
Maybe handoff is the wrong word
The more I work this way, the less “handoff” feels like the right mental model.
A Figma handoff describes an experience we've already designed.
An AI system needs something different.
It needs boundaries within which it can keep designing after we've stopped.
That means our artifacts increasingly describe: what good looks like, what context matters, what the system can decide, what requires the user, and where the system must stop.
The designer isn't specifying every answer.
We're specifying the space of acceptable answers.
And that might be one of the bigger shifts in designing AI products:
It's also the behavior we leave behind.