Staff Designer, Generative AI Lead · Interaction design & systems thinking · ServiceNow Otto · 2023–2026 · Shipped

I designed AI Steps to make agentic work legible

Translating backend processing into task-level progress people could understand and trust

At ServiceNow, I led the design strategy for AI Steps, the product pattern that helped Otto show what it was doing while it worked.

From

Show that the system is working.

To

Show what kind of work the AI is doing, and whether it matches the request.

Illustrative example
What I Did

I designed AI Steps as a product pattern and a set of rules

Otto was becoming more capable. It could check records, compare options, call tools, prepare actions, and synthesize results. But the experience still treated that work like loading.

On the surface, AI Steps was a user-facing component for following agentic work. Underneath, it was a translation layer that turned backend events into task-level progress language.

Transparency is not more information. It is the right level of detail at the right moment.

What the AI is doing

ProblemMulti-step AI work was either hidden behind generic loaders or exposed through raw technical labels.
SignalUsers tolerated longer waits when the work made sense, but escalations showed that raw processing details felt noisy and confusing.
ShiftI defined processing state as an orientation layer: task language, not architecture.
OutcomeAI Steps shipped in Otto. Customer escalations asking to disable processing updates dropped to zero.
My Role

I led the strategy and stayed close to implementation

I led the strategy, wrote the rules, shaped the component, and stayed close to implementation.

I partnered with UX Research to understand how people interpreted latency, output quality, and visible AI work. I pushed the team away from exposing backend events directly and toward a generated language layer that described recognizable task work.

Three layers of the work

I wrote the processing-state specification for engineering, defined the language rules, and shaped the AI Steps interaction pattern so the product surface and behavior system worked together.

AI Steps had to work at three levels: the product pattern people saw, the system rules that made it scalable, and the model behavior underneath.

That meant I was not just designing a component. I was deciding what the AI should reveal, what it should hide, and how the product should help people follow the work without taking over the experience.

This is the kind of work I think modern AI design requires: designing the surface, the system, and the behavior together.

As product design, I had to decide:

  • Where the pattern lived.
  • How much it showed by default.
  • How it behaved while the answer streamed.
  • How it stayed useful after the response completed.
  • How it supported people who wanted to glance, monitor, or inspect.

As system design, I had to decide:

  • How backend events became user-recognizable task steps.
  • Which internal labels should never surface.
  • What vocabulary generated steps should use.
  • How the pattern could scale across agentic workflows.

As model behavior design, I had to decide:

  • What the assistant should reveal.
  • What it should suppress.
  • What it should avoid implying.
  • When specificity was earned.
  • When transparency became noise.

That stack is why this case belongs beside disambiguation and evaluation. It shows the product surface and the behavior system underneath it.

01 / The Orientation Problem

I reframed waiting as an orientation problem

As Otto became more capable, the pause could mean the assistant was retrieving a record, checking a policy, comparing options, calling a tool, generating an answer, or preparing to act.

That changed the design question. During the wait, people were really asking:

Did it understand me?
Is it making progress?
Does the work make sense?
Can I trust the result?

Stage 01 · 2023
Three-dot loader

User can infer: something is happening.

Stage 02 · 2024
Skeleton loader

User can infer: a response is being prepared.

Stage 03 · 2025
Technical processing labels

User can infer: the system is doing work, but not what that work means.

Stage 04 · 2026
AI Steps

User can infer: the assistant is doing recognizable task work.

The component changed because the meaning of the wait changed.

The component changed because the meaning of the wait changed.

Recreation · the four stages rebuilt in CSS; labels illustrative
What people need to know to trust the system

The old problem looked simple: the assistant needed time to respond. A spinner could reassure people that the interface had not stalled.

The user was no longer waiting for a page to load. They were waiting while an AI system did work on their behalf.

Processing state had to answer a sequence of user questions, not just fill time.

Did it understand me?
At the beginning, people need evidence that the system interpreted their request correctly.

Is it making progress?
During the wait, people need to know the assistant has not stalled.

Does the work make sense?
As steps appear, people judge whether the assistant is doing relevant work.

Why did it make that choice?
When the output includes a recommendation, prioritization, or routing decision, people need a trace of the evidence.

Can I trust the result?
When the answer arrives, people need enough context to decide whether to rely on it.

Am I still in control?
When the AI is preparing to act, people need awareness, agency, and the right chance to intervene.

This moved the work from “make waiting feel shorter” to “make AI work legible while trust is being formed.”

02 / The Wrong Layer

I pushed back on transparency that exposed the wrong layer

The first attempt treated transparency as exposure. Every agent, tool, and system operation could produce a visible line.

The original Now Assist processing panel: an expanded 'View AI Steps' list showing raw backend names like 'Used the tool WPS Fetch User Location' and 'Used the tool WPS Knowledge Graph,' next to the live chat thread mid-response.
What actually shipped before the redesign: real tool and agent names, straight from the backend, with no translation layer between them and the user.
Actual product · before the redesign

From an engineering standpoint, that looked transparent. The system was showing its work.

But for users, it was mostly the machine explaining itself in machine terms.

“Executing agent XRS-142” does not build trust.

“Checking your open incidents” does, because it tells the user the system understood the task.

What users were reacting to

The language was technically accurate. It was also mostly meaningless to the people reading it.

Usability research found people stumbling on terms like “Payload ID” and “Rationalizing.” Customers escalated. Some asked us to reduce the detail. Others asked for processing updates to be turned off entirely.

That contradiction mattered:

We added detail to increase transparency. Users asked us to hide it.

That told me visibility and understanding were not the same design problem.

The real question was not:

How much of the system can we expose?

It was:

What actually helps someone follow the work?

What the system exposed

  • Agent names
  • Payload IDs
  • Internal tools
  • Backend tables
  • Routing engines

What users needed

  • What task is happening
  • Whether it matches my request
  • Whether progress is being made
  • Whether I still have control
  • Whether the result deserves trust
The argument that changed direction

The early read from leadership was that this visibility was a feature: show which agent ran, which tools fired, and what the system was doing underneath.

I pushed back because technical names did not build trust. They just moved the confusion into the interface.

There were two arguments I had to win.

1. Suppress technical names

Leadership saw technical transparency as proof of work. I argued that this was not user transparency. It was internal observability leaking into the interface.

A technical label can be true and still fail the user.

The principle I pushed was:

Name the work, not the machinery.

2. Generate better processing language

Once technical names were suppressed, the harder problem was what to show instead.

Hard-coded strings would not scale across agentic workflows. The assistant needed to generate processing language from the request, the task context, and the work being performed.

That was the real design work: defining the translation layer.

03 / Research

I used research to find the level people could actually use

The organization’s first instinct was to treat this as a latency problem: agentic tasks take longer, so make the wait feel shorter.

I partnered with UX Research to test that assumption, and it did not hold.

Users tolerated longer waits when the result felt worth it and they understood what was happening. A shorter wait was not automatically better if the user felt lost.

Not time to response. Time to outcome.

It took a while, but if I knew this is the amount of information it would give me, it’s totally fine.Participant, latency and output-quality study

We ran a card sort using Otto’s actual processing steps, testing what people could understand, what they wanted to see, and which backend operations felt like one piece of work.

Too technicalExecuting ‘relationship_rules_engine’
Exposes architecture, not meaning.
Right level“Checking related incidents.”
Names recognizable task work.
Too generic“Working on your request.”
Reassures without orienting.
The pipeline mapping

Participants did not produce a simplified version of our architecture. They grouped work around what the AI was accomplishing, not which component was executing.

I used the card sort to separate what the system did from what users could meaningfully follow.

The card sort helped me map the system model to the user model.

System model vs. user model

classifierAPIdatabaserules enginerouting agent
understandinvestigatedecideact

The point was not to pretend the technical work did not exist.

The point was to translate it into the level at which a user could judge whether the AI was doing the right thing.

A classifier, API call, database lookup, rules engine, and routing agent might be separate system events. To a user, those could be one recognizable task: “checking related incidents” or “finding the right team.”

The design decision was to show the user model, not the system model.

04 / The Specification

I wrote the specification that turned backend work into task language

Once we stopped exposing technical names, the hard part was deciding what Otto should say instead.

I wrote the engineering-facing specification that turned this from a design preference into a buildable behavior standard.

System languageUser language
Executing AI Agent ‘fetching_user_details’ Finding your contact information
Executing agent XRS-142 Checking your open incidents
Initializing case handler / Writing HRSD table Creating your HR case
Running routing policy evaluator Finding the right team

Every line named a task rather than a component.

The 12-page engineering document

It needed to be specific enough to build confidence, plain enough to understand, brief enough not to distract, and accurate enough not to imply capabilities the system did not have.

The rules were simple, but strict.

Rule 01
Task, not technology

Name what is being accomplished, not the mechanism.

Rule 02
Active and present tense

The user is watching work happen, not reading a log of it.

Rule 03
Specific when it earns it

Detail is how you show real context was found.

Rule 04
No internal terminology

Agent IDs, API paths, model names and tables never leak.

Rule 05
Product vocabulary

Say incident, request, change. The words users already have.

Rule 06
Glanceable

Processing is read in peripheral attention, or not at all.

The engineering-facing specification was the piece that moved the work from argument to implementation.

It explained:

  • why raw technical transparency was not the same as user understanding
  • which backend labels should never surface directly
  • how to generate task-level descriptions from request context
  • how to keep processing language short without making it generic
  • when specificity was earned
  • how to avoid overstating what the system had actually done
  • how to evaluate whether a step helped the user stay oriented

The key principle was:

Processing state is not loading UI. It is how users confirm the AI understood their request and is acting on it.

This helped shift the conversation from “which labels should we show?” to “what does the user need to know to follow the work?”

When transparency becomes noise

I treated processing language as product behavior, not a decorative status layer.

That meant ruling out labels that created the appearance of transparency without adding useful understanding.

Patterns I ruled out:

  • Raw agent names
    They describe implementation, not work.
  • Payload IDs and backend objects
    They make the interface feel technical without making it interpretable.
  • Model or tool names as proof of work
    They ask the user to trust architecture instead of giving them task evidence.
  • Generic reassurance
    “Working on your request” reduces anxiety for a moment, but does not help the user assess direction.
  • Over-specific claims
    “Checking your 2 open incidents” is useful only if the system actually found two open incidents.
  • Long activity logs
    Too much detail competes with the answer and turns the user into an operator.

The standard became:

Show evidence of the task, not exhaust from the system.

05 / The Shipped Pattern

I designed AI Steps around attention and shipped it as a reusable pattern

The final pattern was not a verbose activity log. It supported three modes of attention, and any one person might only ever need the first.

The decision was collapsed AI Steps: current step first, liveness near the answer, full list available on demand.

AI Steps in the shipped product: glance, monitor, inspect.
Actual product · screen recording, Otto

The pattern created a reusable language system. Teams shipping new agentic capabilities inherited a way to describe their work instead of inventing a new transparency pattern each time.

Getting there took two separate calls: how to show the work, and where it should sit.

Shipped · howA spinner for liveness, a status line for the current step, the full trace one click away.

Shipped · whereSplit across the message: the named step stays put, liveness follows the reading.

Five directions for how, and three for where, are below with what worked and what each one cost.

Product choices and placement

I had to decide where the pattern lived, how much it showed by default, how it behaved while the answer streamed, and how it stayed useful after the response completed.

AI Steps had to work as a product pattern, not just a list of generated labels.

The core product decisions were:

  • Collapsed by default
    The pattern stayed quiet unless the user needed more.
  • Current step first
    The visible state answered the most immediate question: what is happening now?
  • Full sequence on demand
    Users could inspect the work without forcing everyone to read the full trace.
  • Placed near the answer
    The processing state stayed connected to the output it helped produce.
  • Answer remained primary
    Transparency supported the response. It did not compete with it.
  • No progress bar
    The system could not reliably estimate completion, so a bar would have performed false precision.
  • No full workflow card by default
    A large panel made transparency the main event and pulled attention away from the result.
  • No pure spinner
    A spinner showed activity but not meaningful progress.

The product pattern worked because it matched the attention model: glance, monitor, inspect.

The second design question was where the processing pattern should sit.

Placement changed what the pattern meant.

Placement A

Above the message

  • WorksSeen first, and it holds still while the answer streams in underneath it.
  • CostsPuts the waiting above the result, so waiting becomes the main event.
Placement B

Below the message

  • WorksSits with the newest content, which is exactly where someone following a growing answer is looking.
  • CostsOnce the answer runs long, the record of the work is somewhere off the bottom of the screen.
Placement C

Split, part at the top and part at the bottom

  • WorksSplits the two jobs. The named step stays where it can be found; liveness follows the bottom, where the reading is.
  • CostsTwo places to look instead of one, and both have to stay quiet enough to ignore.

The tracked revision was:

Pick one direction → Answer both: option 03, split across the message.

The pattern needed to do two jobs:

  • orient the user while the work was happening
  • stay connected to the answer after the work finished

The final placement supported both.

Directions I explored

The first design question was how to show the work.

I explored five directions:

Option 01

Status line that cycles through the steps

  • WorksOne line, always current. It moves, so it never reads as a system that has stalled.
  • CostsEvery step it already did is gone. You can see the moment, never check the work.
Option 02

Thinking spinner that expands on demand

  • WorksNearly silent by default, with the whole trace one click away for anyone who wants it.
  • CostsClosed, it says no more than a spinner, and nothing signals that opening it is worth the click.
Option 03

Spinner and a changing status line

  • WorksCombines 01 and 02: liveness and the current step at a glance, the full list still on demand.
  • CostsTwo moving things in one row. The motion has to stay very quiet or it competes with the answer.
Option 04

Progress card

  • WorksA count and a bar answer “how much is left?” directly, which no other option does.
  • CostsA bar promises an estimate. The system could not reliably produce one.
Option 05

Expanded workflow card

  • WorksThe right amount of evidence when the AI is doing something consequential.
  • CostsFar too heavy as a default. A one-line answer arrives buried under its own receipts.

The final pattern took pieces from several directions: a glanceable current state, a sequence when needed, and a quiet default that did not compete with the final response.

Strengths, limits, and reactions

What worked

  • The language became understandable.
    Processing states described recognizable task work instead of backend components.
  • The pattern respected attention.
    People could glance, monitor, or inspect depending on how much they cared about the result.
  • The system scaled.
    Teams had a reusable pattern and language standard for agentic processing.
  • Escalations dropped to zero.
    Customers stopped asking to disable processing updates after launch.

What it did not solve

Many readable steps can still become overwhelming when the system is doing genuinely complex work.

That limitation became the starting point for the next version.

In an informal before/after side-by-side, users independently described the new labels as clearer and more trustworthy.

That mattered because the redesign did not win by adding detail. It won by choosing better detail.

The lesson was not that users always want the AI to show more.

They want the AI to show the part of the work that helps them decide whether the system understood them and whether the result deserves confidence.

06 / Blueprint Concept · not shipped

I used AI Steps to explore when AI should show more

AI Steps solved the first layer of the problem: task-level progress while the AI works.

But as agentic systems take on more consequential work, progress alone is not always enough.

I started defining the next model as a four-part framework:

Progress

What is the AI working on? Purpose: orientation.

Evidence

What information is it using? Purpose: verification.

Judgment

Where did it make a meaningful choice? Purpose: explanation.

Action

What did it actually do? Purpose: agency.

The goal was calibrated visibility: show more when the system is doing work with higher consequence, and stay quiet when the work is low-risk or obvious.

Risk ladder for calibrated visibility

Not every task needs the same amount of transparency.

Low consequence, retrieval: progress only

Example: checking a policy or finding a record.

Moderate consequence, recommendation: progress + evidence

Example: comparing options or recommending a next step.

Higher consequence, judgment: evidence + explanation

Example: prioritizing an incident or deciding which policy applies.

Irreversible action, acting on someone’s behalf: explanation + explicit agency

Example: submitting, escalating, approving, canceling, or changing access.

This operationalized calibrated visibility.

A low-risk retrieval task does not need a full evidence panel.

A high-consequence action needs more than a status line.

A stronger future-state scenario used incident escalation rather than low-risk troubleshooting: higher stakes, and a real decision to make.

Otto
What’s causing INC0098734, and who should I escalate to?
INC0098734 most likely traces to change CHG0456892, a connection pool update to the orders database that deployed 40 minutes before the first alert. Sarah Chen implemented it and is on call for the database team. I’ve drafted the escalation with the evidence attached. Should I send it?
Investigating INC0098734
Understood the incident
Read INC0098734: orders database unavailable, priority 1
Checked monitoring alerts: connection errors began at 09:38
Checked related changes
Found three changes to the orders service in the last 24 hours
Compared each deployment with the first alert
whyCHG0456892 deployed at 08:58, 40 minutes before errors began. The other two finished earlier with no errors after them
Identified the likely cause
Matched the error logs to CHG0456892: connection pool limit lowered
whyPool exhaustion errors start right after the deployment and nowhere before it
Checked the known issue database: no matching open problem
Found the escalation owner
Sarah Chen implemented CHG0456892 and is on call for the database team
Prepared a recommendation
Drafted the escalation with the change, the timing, and the error logs attached
Holding the escalation until you confirm
whyEscalating pages an on-call engineer, so the decision stays with you
Working…
Concept mockup · illustrative scenario, not shipped
Why this scenario worked
  • The stakes were higher than VPN troubleshooting.
  • The user needed to know not just what the AI found, but why it connected the outage to a specific change.
  • The system needed to expose evidence without turning the user into an operator.
  • The user needed agency before escalation or action.

This scenario made the next frontier clearer: transparency has to scale with consequence.

07 / What Changed

Customers stopped asking us to hide the work

The work changed how ServiceNow communicated agentic AI work.

Otto in production: a user asks how Mother’s Day is trending compared to last year, and the AI Steps panel reports the agent’s work in task language. Two steps are complete (reviewed surge plan activity at Memphis Hub, checked seasonal worker requisition status), one is in progress, and two are still pending.
The shipped pattern in production: every step named as the work it does, with complete, active and pending states readable at a glance.
Actual product · shipped
BeforeAfter
Loading as latency.Waiting as an orientation moment.
Backend events exposed to users.Task-level progress in human language.
Technical names used as proof of work.Recognizable task language used as evidence of understanding.
One-off processing messages.A reusable AI Steps pattern for agentic work.
A UI component carried the whole burden.The component, language rules, and model behavior worked together.
Transparency meant showing more.Transparency meant choosing the right level of detail.

The pattern did not make processing faster.

It made the work understandable enough that customers stopped asking us to hide it.

How the work evolved

Latency research
Research showed that users did not judge the wait only by its length. They judged whether the work and result felt worth it.

Card sort
Users grouped backend processing around recognizable work, not system architecture.

Translation layer
I defined rules for turning system events into task-level processing language.

Product pattern exploration
I explored shape, placement, hierarchy, disclosure, and attention level before landing on collapsed AI Steps.

Shipped AI Steps
The product shipped a pattern that supported glance, monitor, and inspect modes.

Blueprint
The next framework extended transparency from progress into evidence, judgment, and action for higher-consequence AI work.

Closing

I design how AI behavior becomes understandable to people

When AI is acting and not just answering, people need a way to stay oriented without becoming operators of the machinery underneath it.

Product design and model behavior design are not separate layers.