Internal Enablement · Systems Design · Thought Leadership · ServiceNow

Creating a Conversation Design
Discipline at ServiceNow

Almost overnight, a design org of nearly a thousand people was building conversational AI with no training, no framework, and no shared definition of good. I named conversation design as a discipline and built the scaffolding for it: the principles, the training that reached about 2,000 people, and the handoff tool that turns design intent into something engineering can build. The standards are now a required baseline in more than eight business units.

600+designers directly trained (~2,000 reached)
8+business units use the standards as required baseline
84%of attendees felt more equipped after the training

I'm a word nerd who's been advocating for better AI communication for over a decade. This was my moment. I originated the framing and the standards, led the training program with two content-design colleagues, and built the design brief tool myself.

Role
Staff Designer, Conversation Design Lead
Company
ServiceNow
Scope
Conversation design discipline, standards, enablement, infrastructure
Timeline
2024 – Present
“Alea is a major reason I became excited about conversational design. Her work is motivating and sets a high bar for quality and strategic thinking.”Design Coworker · ServiceNow
01 / My Philosophy

Conversation is the most intuitive interface ever built

Often, teams still treat conversation quality as a writing problem. Better copy, clearer tone, more natural language. These things help, but they don't address the real issue: conversation design wasn't treated as a discipline at all.

At ServiceNow, it was treated as decoration on top of a system. A way to make technical output feel friendlier. But that's not what conversation is.

The principle

Conversation design is a discipline. It requires linguistic patterns. It requires ethics about how a system should behave. It requires behavioral principles about when to act versus ask, how to handle uncertainty, what builds trust. It requires actual design thinking, not just better words.

When you get this right, conversation is easier to interact with than any UI. Faster. More natural. Less cognitive load.

When you get it wrong, the damage is worse than a bad UI. Because bad UI is frustrating. Bad conversation feels like betrayal. You trusted the system to understand you, and it didn't.

ServiceNow's design org needed this discipline, so I built it.

The discipline, built in four deliverables

01 Principles
01
02
03
04
05
06
Published in Horizon
02 Training Curriculum
Part 01
Part 02
Part 03
600+ designers directly trained
3 sessions across the org
03 Design Brief
01
02
03
04
05
06
AI tool for designing patterns in non-deterministic experiences
04 Engineering Handoff
system_prompt
entry_prompt
decision_logic
golden_datasets
guardrails
Spec for engineering to build from
02 / How I Framed It

Defining what good looked like

As I began to explore the landscape of skills my colleagues had, I quickly realized that one of the major issues was there was no clear definition of what good meant, both in conversation in general and especially for the ServiceNow users and use cases. This was something I could fix.

I started with principles.

Conversation design at ServiceNow is different from consumer AI because the stakes are different. When someone asks the system to fulfill an access request or route a ticket, it isn't assisting, it's acting. And the user is responsible for what happens next.

Three realities shape every design choice:

First: users are doing high stakes tasks in a time-constrained environment. A service desk agent has 90 seconds to triage. They don't need personality. They need a system that already understands their role, has inferred the context they don't have time to explain, and surfaces the next move without confusion.

Second: automation demands transparency. As the system automates more, users need to understand why it made a decision, not just trust that it can. It's not about showing your work, it's about building confidence in the system's reasoning.

Third: enterprise users know what they want, even if they don't say it clearly. They already have intent, but they may not have time to articulate it perfectly. The system's job is to infer that intent from context and respond intelligently. And because edge cases are what actually ships, design for those first.

These constraints produced the principles below. They're not universal conversation design. They're grounded in the specific work the system is automating.

Principle 01
Know who you're designing for

Design for a specific person in a specific moment. Account for what the system already knows, and design for the variations, because the variations are what ship.

Principle 02
Reduce every ounce of effort

Every interaction should remove work, not add it. Use automation to remove burden.

Principle 03
Design for human agency

Give users real control, not the illusion of it.

Principle 04
Sound like a partner, not a product

No corporate voice, no sycophancy. Just competence. The goal is to collaborate, not just assist.

Principle 05
Design for uncertainty

The system will sometimes be unsure. Signal that clearly. Recovery matters more than prevention.

Principle 06
Design for trust

As users automate more of their tasks, they need to understand what the system is doing on their behalf. Show the evidence that lets them trust the outcome, not a log of everything that happened.

The conversation design principles published in Horizon, ServiceNow's design system
The conversation design principles as published in Horizon, ServiceNow's design system, living as an org-wide baseline, not just a doc.
03 / The Training

Building the training

With principles defined, I needed to teach them to the organization. I brought in two content design colleagues to help me design and run a set of trainings. We spent roughly a month planning and creating the sessions. It ended up being a great collaboration and mind-meld. The three of us together created something stronger than any one of us could have alone. The training ran across three sections, giving designers the fundamentals and showing them how to use that to build the best products.

Part 01
Conversation fundamentals

The base layer: what conversation design is and why it's different from the rest of the craft.

Part 02
Writing for ServiceNow

How to write good conversations for real ServiceNow use cases.

Part 03 · I led this
How & when to use conversation

Conversation is a unique mode of interaction. It can't just replace a UI; it has to earn its place as its own good experience.

Training slide on agentic AI differences Training slide listing the session takeaways
Key concept slides from the training delivered across the design org.

The third part (how and when to use conversation) was the one I led. The core idea was that conversation is a unique mode of interaction. It can't just replace a UI. It needs careful consideration to be its own good and unique experience. Understanding that distinction changed how people approached the work.

The response was immediate

We ran three live sessions and directly trained 600+ designers. The enablement materials and recordings have since reached close to 2,000 people across design, PM, and engineering. Reach isn't the same as impact, so I'll be precise: what these numbers prove is distribution and demand, the material spread well beyond the rooms it was delivered in, and people kept pulling it in. The proof of behavior change comes later, in the standards teams adopted and the work they went on to produce.

One leading indicator worth naming honestly: a month after the training, 84% of attendees reported feeling more equipped to approach conversation design. That’s confidence, not competence, but it’s the precondition for an org actually adopting a new way of working, and adoption is what followed.

“She spent time sharing her knowledge and helping me upskill, which made a huge difference in my understanding of the space.”Content Design Coworker · ServiceNow
04 / Delivering in the Age of AI

How to deliver a modern conversation design

With training complete, I began to ideate on how conversation design functions in the modern age of AI. When I started my career in conversation design, I created flow diagrams and scripts for every conversation that would be supported. That was no longer needed. The AI could write the words itself. So what does a conversation designer actually need to do?

I started by looking at what was happening inside the company, and across the tech industry as a whole.

At the company

I brought in a junior colleague to help figure out exactly what was going on. They sat in on sessions designers were quietly running for their own teams, because the official guidance wasn't working. They interviewed engineers across the organization. They dove deep to get a thorough understanding of the landscape.

Designers
Drafting a single happy-path flow in a Figma doc
Frustrated the product looked nothing like their design
Following training best practices, but under-delivering
No clear way to hand work to engineering
Engineers
Given a single script covering one use case
Specific responses defined, but no AI behaviors
Hard-coding conversations, defeating the point of AI
Filling the gaps with unvalidated assumptions

What they found: nobody knew how to deliver a conversation, and most were under-delivering. Designers were handing off a single happy-path flow and getting back a product that looked nothing like it.

“I wrote a script and gave it to my engineering team. Why doesn't it look like what I designed?”Designer

Engineering had the mirror-image problem: a single script that defined responses but no behaviors, forcing them to either hard-code the flow (defeating the point of AI) or fill the gaps with assumptions design never validated.

“I can't implement this. The AI will never produce those exact responses.”Engineer

The reality was that designers weren't handing off the right deliverables, and were under-defining the things that needed to be clarified. The model is non-deterministic. You can't script every response. You have to design the constraints and let the model reason within them. That wasn't a gap in anyone's experience. It was a shift in what the problem required.

How the industry is handling it

The gap I identified at ServiceNow, designers handing off scripts and engineers making assumptions, isn't unique. Across tech, conversational AI teams are struggling with the same problem of defining and handing off specific AI behaviors. As researchers at Carnegie Mellon University put it:

“Prompt engineering practices exacerbate underspecification. Prompts are developed with the expectation that not everything needs to be specified, ideally, LLMs should behave like a human and fill in the gaps with commonsense. This expectation encourages a ‘minimal-specification’ prompt engineering practice: developers begin with an initial prompt, observe violations of expected behavior, and iteratively revise through adding more instructions… This trial-and-error process lacks the rigor of traditional requirements engineering, and exposes only a narrow slice of possible behaviors.”Carnegie Mellon University · arXiv 2505.13360

I can't know how things function inside the major frontier labs. All I have access to is the products they ship, the articles they publish, and, in some cases, their published system prompts. But from those signals, one thing is clear: they've baked intentional behavioral decisions directly into their prompts. How a system reasons, when it signals uncertainty, how it carries itself, those aren't technical defaults. Someone designed them.

This is the same conclusion, reached from the other direction.

The takeaway

The behaviors and experience that emerge from an AI system have to be designed into the constraints. That work requires both engineering rigor and design intent.

05 / How My Thinking Produced the Solution

The shift from writing to designing

The most important realization came when I accepted what the real problem was. The first goal was to get designers to write better conversations. We tried a number of techniques. But I kept running into the same wall: you can't expect an entire design organization, including people who aren't writers, people for whom English is a second language, people who'd never written conversational copy, to suddenly become conversation writers. That was an unrealistic ask.

The pivot

So I stopped trying to teach them to write. I started teaching them to design for conversation.

The question became: what if designers didn't write conversations at all? What if instead, they designed the system that produces conversations?

The old ask
Write the conversation
A scripted ideal exchange, handed to engineering as the target.
My dashboard isn't working

I can help with that. Is it the Sales dashboard or the Ops dashboard?

Sales

Got it. Checking your permissions now.

It looks like your access expired last week. Want me to request it again?

Yes, please

Done. I've sent the request to your admin and you'll get an email when it's approved.

The new ask
Design the system
The constraints and reasoning the system behaves from.
Jobs to be DoneRestore access to the dashboard the user is trying to reach, with minimal back-and-forth.
Data AvailableUser role, permission history, recent dashboards, system status
Edge CasesPermissions issue · data sync delay · browser cache · dashboard deleted or moved
Behavioral RulesIf more than one dashboard matches, confirm which one before acting. Never change a user's access level without explicit confirmation.

The distinction: output vs. behavioral intent

Building a tool to create a better design and a better handoff

The team was handing off conversation flows that didn't work. They were misaligned with platform patterns. They didn't translate into something an engineer could reliably recreate in a nondeterministic system.

So I am building the Conversation Design Brief. A tool that doesn't ask designers to write response copy, it asks them to do systems thinking. It asks them to define constraints and guardrails. What data flows at what moment. What the system is never permitted to do. What requires explicit user confirmation. How recovery works when things go wrong.

The playbook walks designers through six steps:

The Conversation Design Brief playbook landing page, showing its six steps
The Conversation Design Brief: the landing view, showing the six steps it walks a designer through.
Step 01
Who the user is and what they're trying to do

The designer describes the user and use case.

Step 02
What starts the conversation

Example prompts the user might say to trigger this specific flow.

Step 03
How data is used

Which data sources contextualize the conversation, and what guardrails prevent misuse.

Step 04
What shape the conversation takes

Pre-defined jobs or tasks that become the building blocks of the flow.

Step 05
How visuals are used

What lives in images versus text in the AI response.

Step 06
Edges, escalation, and guardrails

What can go wrong, how to recover, what's off-limits.

A section of the Conversation Design Brief capturing edge cases, escalation, and guardrails
One step in the brief: capturing edge cases, escalation, and guardrails before anything ships.

At the end, the tool uses those responses plus a detailed system prompt I wrote to generate three sample conversations: the happy path and two edge cases. This serves two purposes: it helps non-writers produce a strong script and proves the system will generate what they want. If not, they adjust and regenerate.

Once the designers approve the scripts, the tool produces an engineering handoff: system prompt, data model, decision logic, golden datasets, and guardrails. Everything an engineer needs to implement it.

The start of a system prompt produced by the handoff document The entry prompt section of the engineering handoff document
The engineering handoff: system prompt, data model, decision logic, goldens, and guardrails. Everything an engineer needs to implement the flow.

The tool is still in development, but early feedback from designers is clear:

“I'm not a writer, but I can do this.”Designer · after using the brief

Next is internal usability testing with designers at different levels, engineering validation that they can use the handoff, then release and training. This is a fundamental shift in how design and engineering collaborate on conversational AI. It deserves to be done right.

06 / The Evidence

What the work actually moved

Here is what is already true: the standards I created are a required baseline in more than eight business units, and the discipline changed how the whole org frames its problems. That is adoption, not just reach, and the two aren't the same thing. (The design brief is still in active development, so I won't claim outcomes it hasn't earned yet, that story is still being written.)

Adoption: the standards became the baseline

The clearest signal isn't reach, it's that the frameworks stopped being my documents and became the org's. (How the behavioral standards became the Otto PRD's source of truth is its own story, told in the advocacy case study.) The aggregate picture belongs here: 8+ business units now use the standards as a required baseline, not guidance a team could route around. The three shipped specs, processing state, disambiguation, and response taxonomy, are referenced across PRDs and cross-functional working groups rather than living in a design folder. The conversation design brief is still earning that track record; it's a fourth framework in progress, not counted among these yet.

PRD
Individual frameworks cited as required input across the org's PRDs
8+
Business units using the standards as a required baseline input
4
Core specs referenced across PRDs, not confined to design

Systemic footprint: the thinking outlived the documents

The point of building shared foundations is that other people build on them. Across the conversational AI ecosystem, 25+ identified downstream artifacts, PRDs, framework libraries, training materials, disambiguation standards, and quality programs, trace back to these frameworks and the thinking behind them.

How I measure this

The metric that matters isn't how many things I shipped. It's how many other people's work got better because they had stronger thinking to build on. The shift from individual-team fixes to a shared architectural foundation is visible in how the organization now frames conversation failures: not as model limitations, but as design and architecture problems that can be addressed systematically.

07 / What This Means for the Field

Conversation design is systems design

When I say I built a discipline, I mean the infrastructure that forces an organization to treat conversation as a behavioral system rather than a writing task: principles that define behavior at the system level, training that teaches people to reason about it, and a handoff that turns design intent into something engineers can build from and validate against. Underneath all of it is one idea, designers don't need to write the conversation, they need to design the system that produces it.

Why this matters beyond one company

Most organizations still treat conversation design as a writing problem:

Make the responses shorter Make the tone warmer Sound more natural

That approach breaks down immediately with LLMs. You can't deterministically specify what a model will say. You can specify how it should reason, define confidence thresholds, describe which signals matter to the user, and set behavioral boundaries. The hard problems in conversational AI aren't at the output layer, they're at the reasoning layer. How does the system decide what to do? How does it know when it's uncertain, and signal that? How does it recover from error? When does it ask instead of act?

Those are design problems, and most organizations don't yet have the vocabulary to talk about them. The evidence that the field is converging here is visible from the outside: the major AI labs have baked deliberate behavioral intent directly into their system prompts, not technical defaults, but designed behavior. Someone decided how those systems should reason and carry themselves. That's the same work, at a different scale.

The throughline

Unless you are the systems designer defining behavior at the global level, the way you deliver a conversation is through the building blocks, guardrails, and constraints that let an intelligent model do what you intend. The gap between an assistant people tolerate and one they reach for isn't a model gap, it's a design gap. Closing gaps like that, repeatably, is the work.

References

  1. Chenyang Yang, Yike Shi, Qianou Ma, Michael Xieyang Liu, Christian Kästner, and Tongshuang Wu. “What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts.” Carnegie Mellon University, 2025. arXiv:2505.13360