Internal Enablement · Systems Design · Thought Leadership · ServiceNow

Creating a Conversation Design
Discipline at ServiceNow

Turning conversation design from copywriting into AI behavior design

Almost overnight, a design organization of nearly a thousand people was expected to design conversational AI. Designers were drafting flows, engineers were implementing them, product was asking for AI features, but the org had no shared definition of good conversational AI, no common language for AI behavior, and no reliable handoff between design intent and model implementation. I built the discipline that closed that gap.

What this produced

600+designers directly trained. Training, recordings, and enablement materials reached nearly 2,000 people across design, product, and engineering.
84%of attendees felt more equipped to approach conversation design, one month after training.
8+business units require the standards as baseline. 25+ downstream artifacts trace back to the frameworks and thinking.

I defined conversation design as a systems discipline, created the principles that became an org-wide baseline, co-led training that directly reached 600+ designers and indirectly reached nearly 2,000 people, and built the Conversation Design Brief: a tool for specifying AI behavior instead of writing brittle scripts.

Role
Staff Designer, Conversation Design Lead
Company
ServiceNow
Scope
Conversation design discipline, AI behavior standards, training, design-to-engineering handoff
Timeline
2024 – Present
01 / The Problem

Teams were designing scripts for systems that do not follow scripts

ServiceNow’s design teams were being asked to create conversational AI experiences, but most of the available design artifacts came from a pre-LLM world: happy-path flows, scripted dialogue, and polished example responses.

That created a predictable failure. Designers thought they were handing off the experience. Engineers were receiving one example of the experience. The model then produced something different, because nondeterministic systems do not reliably follow scripts. The real missing artifact was not better copy. It was a behavioral specification.

What the deliverable looked like

A polished script: one ideal exchange, handed off as the target.

What it needed to be

A behavioral specification: the constraints and reasoning the system decides from.

What the internal research surfaced

I worked with a junior colleague to understand what was happening across teams. They observed sessions designers were already running for their own product areas, interviewed engineers, and helped map where the handoff was breaking down.

Designers
Created one ideal flow and expected the shipped experience to match it
Frustrated the product looked nothing like the design
Following training best practices, but under-delivering
No clear way to hand work to engineering
Engineers
Received scripts that didn’t say how the AI should behave across real variation
Specific responses defined, but no AI behaviors
Hard-coding conversations, defeating the point of AI
Filling the gaps with unvalidated assumptions

The pattern was consistent. When implementation differed from the design, both sides were frustrated.

“I wrote a script and gave it to my engineering team. Why doesn’t it look like what I designed?”Designer
“I can’t implement this. The AI will never produce those exact responses.”Engineer

The issue was not collaboration. It was that the deliverable was wrong for the medium.

02 / Where Conversational AI Was Headed

Conversation design is model behavior design

I spent a long stretch working out where conversational AI was going once LLMs took over the response layer, and I concluded that the discipline had to change with it. AI conversation quality does not live in the words. Tone matters. Clarity matters. Good writing matters. But in an AI system, the response is the surface expression of deeper behavioral decisions: what the system inferred, what context it used, what it decided to ask, what it chose not to do, and how it handled uncertainty.

The reframe

Designers do not need to write every conversation. They need to design the system that produces good conversations. That meant defining how the AI should reason, when it should act, when it should ask, what boundaries it should respect, and what evidence users need in order to trust it.

The five layers conversation design actually covers
01
Surface response

People assumed: just the words

Tone, clarity and phrasing. The only layer most people picture, and the one that can least often fix a bad answer.

02
Interaction pattern

People assumed: a flow diagram

What shape the exchange takes: ask, confirm, act, show. Decided long before anything is written.

03
Behavioral rules

People assumed: engineering’s call

What the system infers, decides, or defers to the user. Design’s work, once someone specifies it.

04
Context & data use

People assumed: a privacy review

What the system knows and what it’s allowed to use. A design decision with a trust consequence.

05
Guardrails & evaluation

People assumed: a QA step at the end

What it must never do, and how that gets verified. Written with the standard, not after it.

Conversation design lives at every layer of this stack, not just the top one.

Why this matters for enterprise AI

At ServiceNow, conversational AI was not just answering casual questions. It was helping people complete enterprise work: triaging issues, fulfilling requests, troubleshooting problems, summarizing information, and taking action inside business systems.

That raised the stakes. A consumer assistant can often recover from a vague or overly broad answer. Enterprise AI has less room for that. Users need the system to understand role, context, permissions, task state, and risk. They need to know when it is acting on their behalf, when it is asking for confirmation, and why a recommendation is safe to trust.

That is why I treated conversation as a behavioral system, not a tone layer.

03 / The Foundation

I gave the organization a shared definition of “good”

Before teams could design better conversational AI, they needed a shared quality bar. I created a set of conversation design principles for ServiceNow and published them in Horizon, the company’s design system. The goal was not to create generic writing guidance. It was to define the behavioral expectations for enterprise AI experiences. They are published in Horizon, ServiceNow’s design system.

Know who you’re designing for

Design for a specific person in a specific moment. Account for role, permissions, task, context, and what the system already knows.

Reduce every ounce of effort

Every interaction should remove work, not add it. Use context, automation, and inference to reduce the user’s burden.

Design for human agency

Give users real control, not the illusion of it. Be clear about what the system can do, what it’s about to do, and where the user must confirm.

Sound like a partner, not a product

No corporate voice. No sycophancy. No fake personality. Just competence, clarity, and collaboration.

Design for uncertainty

The system will sometimes be unsure. Handle that intentionally, through clarification, caveats, escalation, or recovery.

Design for trust

Trust comes from useful evidence. Users need to understand what the system did, why it matters, and what they can do next.

The conversation design principles published in Horizon, ServiceNow's design system
The conversation design principles as published in Horizon. An org-wide baseline, not just a doc. Read them in Horizon →
04 / The Training

I taught designers to reason about AI behavior, not just improve AI copy

Once the principles existed, I helped turn them into an organization-wide training program. I partnered with two content design colleagues to design and deliver three sessions across the design org. Together, we taught the fundamentals of conversation design, how to write for ServiceNow’s product context, and, in the section I led, how to decide when conversation should be used as an interaction model at all.

Part 01
Conversation fundamentals

What conversation design is and why it differs from general UX writing.

Part 02
Writing for ServiceNow

How to write clear, useful, enterprise-appropriate conversations for real product use cases.

Part 03 I led this
How & when to use conversation

Conversation as a unique mode of interaction, not a replacement for UI.

Training slide on agentic AI differences Training slide listing the session takeaways
Key concept slides from the training delivered across the design org.

The most important shift I taught was getting designers to ask better product questions: Should this be conversational? What does conversation make easier? Where does the user need structure instead of chat? When should the AI infer, ask, act, confirm, or escalate? What needs to be visual instead of verbal?

What the training actually changed

A month after each session, 84% of attendees said they felt more equipped to approach conversation design. That was measured at one month rather than on the day, because the day-of number measures the session and the month-later number measures whether anything stuck.

The material later became a hands-on workshop at ServiceNow Knowledge 2026, where I taught customers and practitioners how to identify and fix common AI conversation failures, written up in Why AI Conversations Go Wrong.

Alea Abrams presenting the conversation design workshop at ServiceNow Knowledge 2026, beside a slide listing the session agenda
Running the conversation design workshop at ServiceNow Knowledge 2026.
“She spent time sharing her knowledge and helping me upskill, which made a huge difference in my understanding of the space.”Content Design Coworker · ServiceNow
Training structure and evidence

The core message across all three parts was consistent: conversation should not be used because AI is available. Conversation should be used when it creates a better path for the user.

Reach isn’t the same as impact, so I’ll be precise about what these numbers prove: distribution and demand. The material spread well beyond the rooms it was delivered in, and people kept pulling it in. The 84% figure is confidence, not competence, but it’s the precondition for an org actually adopting a new way of working, and adoption is what followed (see the outcome).

The Knowledge 2026 version took the internal training outside the company, teaching the same anti-patterns to a room of customers and practitioners rather than ServiceNow designers.

05 / Creating an Implementation-Ready Artifact

Engineering needed the behavioral logic, not a better script

Training created shared language, but teams still needed a better way to deliver conversation design. The existing handoff model asked designers to produce polished examples. But engineering needed something else: the behavioral logic behind those examples.

A useful AI handoff needed to answer:

What is the user trying to do? What context should the AI use? What decisions can it make on its own? What requires confirmation? What should never happen? How should it recover when it’s wrong? What examples prove the behavior works?

This was the shift from output design to system design. Researchers at Carnegie Mellon University describe the same failure mode from the engineering side:

“Prompt engineering practices exacerbate underspecification. Prompts are developed with the expectation that not everything needs to be specified… This trial-and-error process lacks the rigor of traditional requirements engineering, and exposes only a narrow slice of possible behaviors.”Carnegie Mellon University · arXiv 2505.13360
Why scripts failed as AI handoff

A scripted flow can communicate tone and intent, but it cannot fully define a nondeterministic AI experience. For example, a script might say:

My dashboard isn’t working.

I can help. Is it the Sales dashboard or the Ops dashboard?

But implementation needs to know much more:

Should the AI check recent dashboards before asking? Should it infer from the user’s role? Should it check permissions first? What if the dashboard moved? What if multiple dashboards match? What if the user lacks access? Can the AI submit a new access request? Does that require confirmation?

The problem was not that scripts were useless. It was that they were incomplete. They showed what the AI might say, but not how it should decide.

I built the Conversation Design Brief

To solve the handoff problem, I created the Conversation Design Brief: a structured tool that helps designers define AI behavior instead of writing brittle scripts. The brief asks designers to specify the system underneath the conversation: the user goal, trigger prompts, available data, interaction shape, visual response patterns, edge cases, escalation paths, and guardrails.

01Designer intent
02Structured brief
03Generated examples
04Revised behavior
05Engineering handoff

The tool uses those inputs, along with a detailed system prompt I wrote, to generate sample conversations across the happy path and edge cases. Designers can inspect whether the output reflects their intent, revise the behavioral inputs, and regenerate. The brief walks designers through six steps:

The Conversation Design Brief playbook landing page, showing its six steps
The Conversation Design Brief: the landing view, showing the six steps it walks a designer through.
Step 01
Who the user is and what they’re trying to do

Define the user, their context, and the job they need the AI to help complete.

Step 02
What starts the conversation

Capture the kinds of prompts that might trigger the experience, including vague or ambiguous inputs.

Step 03
How data is used

Specify what data the AI can use, when it should use it, and what boundaries prevent misuse.

Step 04
What shape the conversation takes

Define the interaction pattern, jobs, and tasks that structure the experience.

Step 05
How visuals are used

Decide what should be communicated through text versus UI, cards, tables, or citations.

Step 06
Edges, escalation, and guardrails

Define what can go wrong, what requires confirmation, what’s off-limits, and how the system recovers.

A section of the Conversation Design Brief capturing edge cases, escalation, and guardrails
One step in the brief: capturing edge cases, escalation, and guardrails before anything ships.

Once the examples are approved, the tool produces an engineering handoff with the system prompt, entry prompt, data model, decision logic, golden datasets, and guardrails.

“I’m not a writer, but I can do this.”Designer · after using the brief
See the engineering handoff artifacts and implementation details
The start of a system prompt produced by the handoff document The entry prompt section of the engineering handoff document
The engineering handoff: system prompt, data model, decision logic, goldens, and guardrails.

The handoff produced by the brief includes: system prompt, entry prompt, data model, decision logic, golden datasets, guardrails, escalation rules, and sample conversations.

This matters because it gives engineering a behavioral target, not just a sample output. The sample conversations are useful, but they are not the source of truth. The source of truth is the underlying behavioral specification: what the system should know, infer, ask, confirm, refuse, show, and recover from.

The tool is still in development. Next is internal usability testing with designers at different levels, engineering validation that they can use the handoff, then release and training.

06 / The Outcome

The work became organizational infrastructure

The strongest signal of impact is that the work stopped being mine and became part of how the organization defined conversational AI quality. The standards are now used as a required baseline across more than eight business units. The frameworks show up in PRDs, cross-functional working groups, training materials, and downstream design artifacts. Across the conversational AI ecosystem, 25+ identified artifacts trace back to the principles, frameworks, and thinking.

PRD
Individual frameworks cited as required input across the org’s PRDs
8+
Business units using the standards as a required baseline input
25+
Downstream artifacts tracing back to these frameworks

That is the measure of a discipline: other people’s work gets better because the foundation exists.

How I measure the impact

I separate reach from adoption. Reach means people saw the material: the 600+ designers trained and nearly 2,000 people reached through recordings and enablement. Adoption means the standards became part of how teams worked: used as required baseline inputs, referenced in PRDs, and carried into downstream artifacts.

The design brief is still in development, so I would not claim shipped outcomes for it yet. Its role in the story is different: it shows how I translated the discipline into the next layer of implementation infrastructure.

What I personally did

My contribution was to make an emerging discipline operational. I originated the framing that conversation design is a systems discipline, not a writing task. I created the principles that defined the quality bar. I partnered with content design colleagues to scale the training. I led the section that taught designers how to decide when conversation is the right interaction model. I identified the handoff gap between design and engineering. I designed the Conversation Design Brief, wrote the system prompt behind it, and shaped the handoff around what engineers need to implement nondeterministic AI behavior.

Framed the discipline Defined the principles Co-led training Identified the handoff failure Designed the brief Wrote the system prompt Specified the engineering handoff
07 / Why This Matters for AI Design

The next generation of AI design happens below the interface

As AI systems become more capable, the most important design work moves below the surface. The question is no longer only what the user sees. It is how the system decides what to do.

What should it infer? When should it ask? When should it act? What requires confirmation? What should it explain? What should it refuse? How should it recover? How does it preserve user agency while still being useful?

This case study shows that I can work at that layer. I can design the interface around the model, but more importantly, I can design the behavioral system the model operates within. That is what I bring to a senior AI interaction or model-behavior design role: the ability to define what good AI behavior means, turn it into practical systems, and help an organization build toward it repeatedly.

The throughline

Model capability+product context+behavioral constraints+user agency=trustworthy AI experience

References

  1. Chenyang Yang, Yike Shi, Qianou Ma, Michael Xieyang Liu, Christian Kästner, and Tongshuang Wu. “What Prompts Don’t Say: Understanding and Managing Underspecification in LLM Prompts.” Carnegie Mellon University, 2025. arXiv:2505.13360