Staff Designer, Generative AI Lead · AI behavior & prompt design · ServiceNow Otto · 2023–2026 · Shipped

I defined when AI
should ask instead of answer

Turning ambiguity into intentional AI model behavior

At ServiceNow, I led the design strategy for how an enterprise AI assistant should respond when a user’s request could mean more than one thing.

From
Answer fast Be accurate by showing everything
To
Use what the system knows Ask for the next step
“From” verbatim, 2024
What I Did

I defined how AI should behave when it only partly understands

I designed, defined, and evaluated AI behavior for ambiguous enterprise requests.

Stop treating ambiguity as a detection problem. Start treating it as a decision problem.

I created the behavior model, ambiguity taxonomy, prompt guidance, writing standards, examples, evaluation criteria, and implementation principles that helped PM, prompt engineers, linguists, researchers, and product teams decide what the assistant should do when it only partially understood.

When to ask, when to answer

Problem The AI answered underspecified requests with premature certainty or tried to cover every possibility.
Signal 90% of analyzed queries carried meaningful ambiguity.
Shift I defined a confidence-by-consequence framework for deciding whether the AI should answer, ask, narrow, or make assumptions visible.
Outcome The new strategy increased clarification triggers by 175%, evaluated against false friction. 15 teams requested early access to the next-generation framework.
My Role

I led the behavior strategy, and partnered with research on how to measure it

I led the behavior design strategy for clarification across a multi-year effort.

I partnered across UX Research, Engineering, Product Design, Content Design, prompt engineering, and quality to test, refine, implement, and evaluate the work.

Ownership

I led

  • Reframing ambiguity as a behavior design problem, not a UI or tone issue
  • Defining the decision boundary for when the AI should answer, ask, narrow, or reveal an assumption
  • Separating confidence from consequence so high-risk actions did not get treated like low-risk answers
  • Creating the ambiguity taxonomy that made the problem designable
  • Writing the interaction principles for focused clarification
  • Translating the behavior strategy into prompt guidance, examples, and implementation principles
  • Defining the evaluation lens: not just “was this a good answer?” but “should the AI have answered at all?”
  • Turning the strategy into artifacts teams could build from, not just recommendations
  • Staying close to implementation so the reasoning survived architecture reviews, prompt decisions, and product tradeoffs

We partnered on

  • Quantitative research and query analysis
  • Usability testing
  • Prompt implementation
  • Human-ranking and linguistic evaluation
  • Product and engineering delivery
  • Shipping the updated behavior in Otto

Research led the official evaluation framework and tooling. My contribution was the behavioral criteria, decision model, prompt guidance, and designer-facing interpretation of what “good” meant for clarification.

01 / Reframing the Problem

I reframed ambiguity as a behavior design problem

A request like “PTO” can mean several different things.

This is how ServiceNow handled it:

Before · the real 2024 response to one ambiguous word
Otto
PTO

ServiceNow offers a Flexible Paid Time Off (FTO) program for managing time away from work. Here is some information about the PTO policy.

Under the current Flexible Paid Time Off (FTO) policy for U.S. employees, you are able to take appropriate time off for vacation, illness, emergencies, bereavement, jury service, or personal reasons. There is no set limit on the number of FTO days you can take in a given year, however, time off must generally be approved in advance by your manager, except in cases of illness, bereavement, or emergencies, which can be reported after the fact…

continues for 364 words

What failed

It optimized for a relevant answer, not the right behavior. “PTO” could have meant balance, request flow, approval status, eligibility, or leave type.

It assumed a region. The policy was region-specific. The user never said where they worked.

It answered broadly instead of resolving intent. The user had to scan a long response to find what applied.

Recreation · the response is verbatim from ServiceNow, 2024
I would have wanted it to ask a question back to me to refine the prompt.Usability study participant

The scale of the problem mattered. In query analysis, 90% of analyzed requests carried meaningful ambiguity.

Ambiguity across analyzed queries

90%

of analyzed queries carried meaningful ambiguity.

Highly ambiguous · 47% Moderately ambiguous · 43% Clear · 10%
Why the response failed

A conversational system has to decide which meaning the user intended. If it chooses incorrectly, the response can be factually correct and still fail.

That was happening at ServiceNow. The system was over-indexing on accuracy and coverage: find relevant information, avoid hallucination, and answer as completely as possible. In ambiguous moments, that produced two failure modes.

Two failure modes

  • Commit: pick one interpretation and answer as if it were the only one.
  • Cover everything: answer every plausible interpretation at once and let the user sort it out.

A behavior model optimized for “accurate and complete” was still failing the user. Accurate and complete is not the same as understood.

A response can be accurate and still be wrong for the user’s intent. A response can be comprehensive while making the user do the work the AI should have done.

That meant “detect ambiguity” was not a useful strategy. At that density, ambiguity was not an exception. It was the normal condition of enterprise conversation.

Why this mattered

The wrong-intent problem was annoying. The region assumption was riskier: nothing in the response marked it as an assumption. The assistant stated the assumed policy scope with the same confidence as the retrieved policy content.

The broader pattern was just as important: the AI was being rewarded for producing accurate, complete answers, even when the more useful behavior would have been to pause, narrow, or ask.

That distinction shaped the rest of the work: some AI failures come from saying something false. Others come from doing the wrong thing with true information. Clarification is most valuable at that boundary.

Failure patterns and research

I looked at failure patterns where the response was not obviously “bad” by traditional standards. The system often retrieved the right source, used plausible wording, and produced a complete answer.

The issue was what the response made the user do next.

What the system didWhat recovery work the user had to do
“PTO” → full policy, unpromptedScanned a long answer to find whether it applied.
“Update my address” → generic address instructionsFigured out whether this meant home address, mailing address, tax records, benefits, or emergency contact.
“New computer” → internal catalog choicesTranslated system ontology into a task they recognized.
“Adobe” → broad software answerClarified whether they needed access, installation, troubleshooting, license information, or approval.

The pattern became clear: answer correctness is not the same as intent correctness.

So the evaluation and behavior strategy needed to measure whether intent was resolved, not just whether the answer was right.

The research showed that ambiguous requests were not limited to obviously vague prompts.

RequestPossible meanings
“PTO”Policy, balance, request flow, approval, eligibility.
“Update my address”Home address, mailing address, emergency contact, tax records, benefits.
“Set up my laptop”Order hardware, configure software, troubleshoot access, prepare for a new hire.

Longer requests were not safer. They gave the AI enough to start and not enough to finish. People were not writing perfectly specified prompts. They were using shorthand, incomplete references, everyday terminology, prior context, and assumptions about what the assistant already knew.

Query-coding pass · three buckets

High ambiguity · “Adobe” Moderate ambiguity Clear

The “clear” pile keeps shrinking as we read more carefully.

A usability study also showed that long, catch-all answers were making tasks harder to complete. One participant said they would have wanted the assistant to ask a question back rather than guess.

The important signal was not that users wanted more questions. It was that they preferred a focused question when it got them to the right answer faster.

02 / Friction vs. Burden

I separated helpful friction from answer burden

Clarification can look like friction if you count only conversational turns.

But the dominant burden in the existing experience was not always guessing. It was often answer burden: the assistant produced a long, accurate-enough response and made the user figure out which part applied.

I looked at the full journey differently.

Without clarification

Path 1 · broad and generic

Ambiguous request
Broad, generic response
User reads or scans the whole thing
User decides what applies
Often still generic, needs follow-up
User clarifies
and round again

~2–4 min · reading, scanning, reformulating

Path 2 · hyper-specific, often wrong

Ambiguous request
Hyper-specific response, often incorrect
User reads the response
User sends a message to clarify
and round again

~1–3 min · a wrong turn, then a redo

With intentional clarification
Ambiguous request
AI weighs confidence in its understanding
AI determines what is unresolved
AI asks one focused question
User answers
AI gives the targeted answer

~20–30 sec · one added turn, no repair work

Illustrative estimate: fewer steps on paper does not mean less time once scanning, reformulating, or a wrong turn are counted.

Understand what the user wants, then give them what they need. No more, no less.
Why not just ask more

The goal was not minimum turns. It was the shortest path to the right outcome.

One extra turn can remove the work of reading the wrong answer. I did not want a system that asked every time something was ambiguous.

The risk: false friction

Asking by default would solve answer burden by creating a different failure mode. Clarification only works when it reduces the total work for the user. A bad clarification question can be just as harmful as a bad answer: it can ask for information the system already has, expose internal ontology, or force the user to understand how the system is organized.

This is the arithmetic behind a claim that sounds counterintuitive: adding a turn can shorten the journey when the alternative is scanning, repair, reformulation, or abandonment.

That meant the assistant needed to decide:

The behavior decision checklist · know, infer, risk, burden, ask, act

  • Know: What does the system already know?
  • Infer: Can context resolve the likely intent?
  • Risk: What happens if the system guesses wrong?
  • Burden: Would a broad answer create work for the user?
  • Ask: What is the smallest question that changes the outcome?
  • Act: Is the system calibrated enough to proceed?

This is why clarification became a behavior policy, not a copy pattern.

03 / The First Version

I used the first version to define what better had to mean

I had been arguing that Otto needed to clarify instead of guess, and in 2024 the engineering team implemented a first version.

It asked, but it handed the ambiguity back to the user.

What V1 actually asked
new computer
I can help with a new computer. Which one do you want? Hardware Asset Provisioning Package, standard new hire package, computer bundle, developer hardware stack, or replacement device?

That was the lesson: clarification is only useful if the user can recognize the choice.

Then a routine platform update dropped the behavior, and customers noticed right away.

More than ten enterprise customers escalated to get it back. That was one of the clearest signals in the project: clarification was not polish. It was part of whether the assistant felt competent.

Why a silent removal mattered

The early version asked whenever there was more than one possible interpretation. It was directionally right, but blunt. The assistant often asked in the system’s language rather than the user’s.

It didIt did not
Stop before guessingAsk in language the user could recognize
Prove clarification matteredReduce the user’s decision burden
Create customer valueDefine a durable behavior strategy

Nobody set out to run this test. The behavior disappeared in a routine platform update, and existing metrics did not treat it as a regression: answer rates, latency, and resolution did not clearly explain what changed.

Users noticed the absence of clarification faster than they had noticed its presence. That is what made the customer escalations meaningful: they surfaced a value existing quality signals were not set up to measure.

The question technically clarified, but it clarified in the system’s language. The user still had to understand ServiceNow’s internal categories before they could choose.

04 / The Decision Boundary

I defined the decision boundary: confidence by consequence

When the cost of being wrong is high, risk overrides confidence.

The behavior model · four responses, two variables

Consequence ↑
Low confidence · High consequence
Ask before acting
High confidence · High consequence
Make the assumption visible, or ask first
Low confidence · Low consequence
Offer likely paths
High confidence · Low consequence
Answer directly
Confidence in the leading interpretation →

This model turned ambiguity into a buildable decision boundary: the assistant chose a response strategy based on both confidence in the leading interpretation and the consequence of being wrong.

The point was not to eliminate ambiguity. It was to handle it proportionally.

Why confidence alone failed

The rebuild created the opportunity to define the behavior more intentionally.

I found that two variables mattered most:

Confidence: how strongly the system could identify the user’s likely intent.

Consequence: what would happen if the system guessed wrong.

The important move was keeping those variables separate.

Confidence was the obvious signal. It was also the wrong one.

A single certainty score is easier to route on, easier to test, and cheaper to implement. But it collapses two different questions into one:

  • How likely is the leading interpretation?
  • What happens if that interpretation is wrong?

I pushed to keep consequence separate because two requests can have similar confidence and very different risk.

“Approve that” “Reset my account” “PTO”

Same confidence, different risk: each can appear interpretable, but the consequence of being wrong changes the right behavior.

Retrieving a policy is cheap to get wrong. Writing a record, triggering a workflow, approving something, changing access, or resetting an account is not.

Consequence is a property of the action, not the model’s feeling of certainty. It belongs in the behavior system.

A confidence-only model treated ambiguity as a single score. It could tell the system how likely the top interpretation was, but not how expensive it would be to choose incorrectly.

RequestConfidence onlyConfidence × consequence
“PTO”Might retrieve a likely policy answer.Asks whether the user wanted a balance, request, policy, eligibility check, or approval flow before committing to one path.
“Approve that”Might ask because the wording is vague.Checks available context first. If there is one visible pending approval, asking “approve what?” creates false friction.
“Reset my account”Might proceed because the intent sounds clear.Recognizes that the cost of the wrong reset is high, so the system should confirm scope before acting.

The point was not to make the assistant cautious everywhere. It was to make caution proportional to risk.

The system behind one sentence

A clarifying question can look simple in the UI, but the sentence is the visible output of a larger behavior system.

Ambiguity taxonomy Behavioral rules Context-handling guidance Writing standards Examples and counterexamples Product requirements Risk overrides Evaluation criteria

Each artifact answered a different question

  • Ambiguity taxonomy: What kind of uncertainty is present?
  • Behavioral rules: Should the assistant answer, ask, narrow, or reveal an assumption?
  • Context-handling guidance: What can the assistant infer from the user, task, session, role, or available data?
  • Writing standards: How should the assistant ask without handing the ambiguity back to the user?
  • Examples and counterexamples: What does good and bad behavior look like in real product moments?
  • Product requirements: What must the platform support for the behavior to work?
  • Risk overrides: When should consequence override confidence?
  • Evaluation criteria: How do we know whether the assistant made the right decision?

Every layer could independently produce the wrong sentence. None of those layers were visible to the user.

That is why the design work could not stop at the wording. The wording had to carry the behavior model.

05 / Focused Clarification

I translated the model into focused clarification

A good clarifying question carries forward everything the system already understood and asks only for the missing decision.

Ask only for what remains unresolved.

“PTO”

“What do you mean?”

“I can help with your PTO balance, a new request, or the policy. Which one do you need?”

“Update my address”

“Can you clarify?”

“Are you updating your home address for payroll, benefits, or emergency contact records?”

“New computer”

“Which one do you want: Hardware Asset Provisioning Package, standard new hire package, computer bundle, developer hardware stack, or replacement device?”

“Are you trying to request a new work computer, replace your current one, or get help setting it up?”

Otto in production: the user typed 'computer,' and the AI asks which of three things they need (learning basic computer concepts, Mac tips if switching from Windows, or requesting a company computer) instead of guessing or listing everything it knows about computers.
The shipped behavior asked targeted clarifying questions in the user’s terms, not the system’s ontology.
Actual product · shipped
How I built the taxonomy

The taxonomy did not start clean.

I built it by grouping, splitting, renaming, and testing ambiguous requests until the categories described what the system needed to decide, not just how vague the sentence sounded.

Taxonomy in progress · clustering pass 3 of 6

Vague requests Wrong words → Domain terminology Incomplete Which one?
“Vague” is doing too much work. Split by what’s missing.
Group by the missing variable, not by how the sentence sounds.
One is missing a required parameter. The other is missing a referent. Keep them apart.
The “clear” pile kept shrinking as we read more carefully.

A generic approach treated all ambiguity the same. A categorized approach let the system match the response to the specific problem.

The categories evolved as I tested more examples. Some merged, some split, and some cases that looked similar at first required different behaviors.

That became the important design move:

Not all ambiguity should trigger the same response.

Final ambiguity taxonomy · five types

  • Multiple valid interpretations: Which task, policy, workflow, or record does the user mean?
  • Domain terminology: Does the user’s phrase map to more than one internal concept?
  • Missing required information: What variable is needed before the system can continue?
  • Entity ambiguity: Which person, ticket, asset, case, or record is being referenced?
  • Contextual ambiguity: Can prior context resolve the request, or is the reference still unclear?

The taxonomy made “ambiguous” specific enough for prompt engineers to route, content designers to write, and evaluators to score.

Each type required a different response strategy.

Revising one clarifying question

This is the difference between asking a question and designing a useful clarification.

Scenario

User: Add me to the project.

Context available to the system:

  • The user is on three active projects: Atlas, Northstar, and Meridian.
  • Northstar has an open access request.
  • The system can use project context, but should not assume the user wants the open request unless it makes that assumption visible.

Four drafts Each draft, and what still failed.

“Can you clarify which project you mean?”

Draft 1

What failed: generic. It used none of the context the system already had.

“Please specify the project sys_id or name, and the access level required.”

Draft 2

What failed: two questions, one in internal jargon. It asked the user to understand the system.

“You’re on Atlas, Northstar, and Meridian. Which one?”

Draft 3

What improved: recognition instead of recall. What still failed: it ignored the open Northstar request and gave the user no way out if those options were wrong.

“You have an open access request for Northstar. Is that the one, or did you mean Atlas or Meridian?”

Final

It led with what the system understood, used available context, asked one question, used the user’s language, and gave the user a way to correct the assumption.

The final question was not just nicer copy. It encoded the behavior model.

Patterns I ruled out

I also created examples and counterexamples so teams could recognize what not to ship.

Exposing internal taxonomy

Example: asking users to choose from internal record types like sn_hr_core_case. Why it failed: the user should not need to understand ServiceNow’s backend to clarify their goal.

Asking the user to diagnose the AI’s uncertainty

Example: “Can you clarify what you mean?” Why it failed: it gave the user the problem without naming the missing decision.

Showing too many interpretations at once

Why it failed: it turned ambiguity into a menu the user had to parse.

Asking multiple questions per turn

Why it failed: it made the next step feel like a form, not a conversation.

Ignoring available context

Why it failed: it made the assistant feel less competent and increased user effort.

Asking for information the system already had

Example: asking for an email address when the user was signed in. Why it failed: it created false friction and weakened trust in necessary questions.

Trapping users inside suggested options

Why it failed: the assistant must offer likely paths without making the options feel exhaustive when they are not.

The standard became:

Ask for the missing decision, not for clarification in general.

The guidance went beyond “sound natural.”

I defined rules that prompt engineers and content teams could apply:

  • Lead with what the system understood
  • Ask one question per turn
  • Ask for the distinction that changes the outcome
  • Mirror the user’s language
  • Do not expose internal ontology
  • Preserve information the user already gave
  • Offer a way out when none of the options fit
  • State assumptions when ambiguity cannot be fully resolved
  • Avoid asking when context already resolves the likely intent

Clarification anatomy: what the AI understood, plus the unresolved distinction, plus user-recognizable options, plus a way out if none fit.

The design goal was not a nicer question. It was a better decision.

The assistant should not make the user restate the whole request. It should carry forward what it knows and ask only for the missing variable.

06 / Evaluation

I evaluated the decision, not just the answer

A higher clarification rate was meaningful, but only if the system was asking at the right moments.

I worked with linguists and evaluators to translate the behavior spec into a reusable evaluation framework.

Evaluating the decision, not the sentence

Every logged interaction
Question 01 · what was knowable

Did the system actually have enough information to act?

No
Question 02

What did it do?

It proceeded False certainty

The system commits to an interpretation the evidence did not support. It proceeds where it should have clarified.

“reset my account” → full account reset
It asked Calibrated

Recognized that a materially different interpretation was live, and asked for exactly that.

“onboard Alex” → employee, vendor, or tool?
Yes
Question 02

What did it do?

It proceeded Calibrated

Acted on a well-supported understanding, with the assumption visible where it mattered.

“approve that” → approved the pending item
It asked False friction

The AI asks when it already knows enough. Each unnecessary question spends user attention and erodes trust in the ones that matter.

“approve that” → “approve what?”

Grading the sentence first tells you whether the response was well written.

Grading what was knowable first tells you whether the system should have responded that way at all.

What I tested

The rubric asked whether the system:

The rubric asked whether the system · five criteria

  • Understood the user’s likely goal
  • Identified the right kind of ambiguity
  • Preserved what the user had already provided
  • Asked only for the missing decision
  • Matched the response to the consequence of being wrong

Every logged interaction could be evaluated through two questions:

1. Did the system actually have enough information to act?
2. What did it do?

The order of the questions was the point. If we started by scoring whether the clarification sentence was well written, we missed the actual behavior failure. A beautiful clarification question can still be wrong if the system already had enough information to act.

This made evaluation more useful for AI behavior design. It judged the decision, not only the output.

What the comparison showed

A separate, smaller head-to-head evaluation gave texture to the rollout.

Two candidate implementations were run against the same scenario set. The implementation built to the disambiguation standard asked a needed clarifying question in its first turn on 55% of scenarios that required one, compared with 20% for the comparison implementation.

This was not optimized as a target. It was useful because it showed that the behavior standard changed what the system recognized as answerable.

What shipped and what remained

The 175% increase is indexed to the prior implementation because the absolute trigger rates are confidential.

What makes the number worth trusting is not its size. The increase was what the new decision boundary produced, not a target the team optimized for.

The shipped strategy was an intermediate step, not the full behavioral framework. It used the technical mechanisms available in the platform to trigger clarification more effectively and reduce automated guessing.

The full next-generation framework weighed context, interpretation strength, ambiguity type, consequence, and response strategy together. Fifteen teams requested early access before it launched.

What the model projected

Internal business modeling projected deflection improving from roughly 8% to 36% as conversational pattern coverage expanded from 40% to 83% of intents.

That is a projected 4.5× lift.

Pattern coverageProjected deflection
Before40%~8%
Projected83%~36%

Business model projection, not claimed outcome.

This is business context, not an outcome I claim credit for. It explains why better AI behavior mattered at scale.

07 / What Changed

Clarification triggers rose 175%, evaluated against false friction

The work changed how clarification was understood, built, and evaluated.

The new strategy produced a 175% increase in clarification trigger rate compared with the prior implementation.

Prior strategy indexed at 100 · new disambiguation strategy indexed at 275

100
Prior strategy
175% increase
275
New disambiguation strategy
Indexed to the prior implementation because absolute trigger rates are confidential. The increase was evaluated against false friction, not optimized as a standalone target.
BeforeAfter
Ambiguity was treated as a model uncertainty problem.Ambiguity became a behavior decision.
The assistant guessed or over-answered.The assistant could ask, narrow, reveal assumptions, or answer directly.
Clarification was judged as a conversational nicety.Clarification became part of the product quality bar.
Design shaped the wording after the fact.Design shaped what the system should decide before it responded.
Evaluation asked whether the answer was good.Evaluation asked whether the assistant should have answered at all.
How the work evolved
Ambiguity research

Research showed that ambiguity was a default condition of enterprise AI interaction, not an edge case.

First implementation

The early implementation asked whenever there was more than one possible interpretation. It proved the value of stopping before guessing, but also showed the limits of blunt clarification.

Grey-area strategy

I defined a more nuanced behavior strategy that weighed context, interpretation strength, ambiguity type, consequence, and response strategy.

Shipped intermediate strategy

The shipped strategy worked within the platform mechanisms available and increased clarification trigger rate by 175% compared with the prior implementation.

Next-generation framework

The broader framework gave teams a reusable model for deciding whether the assistant should answer, ask, narrow, or reveal an assumption. Fifteen teams requested early access before launch.

How I made it stick

The deeper impact was organizational.

Instead of design coming in at the end to make clarification sound better, design had a role in deciding what the system should detect, what it should infer, what it should ask, how it should respond, and how teams should know whether that behavior was right.

I made the strategy stick three ways.

01

I created shipping artifacts, not recommendations. I translated the behavior strategy into frameworks, prompt guidance, writing standards, examples, counterexamples, and evaluation criteria teams could build from.

02

I stayed in architecture reviews as a thinking partner. I helped teams reason through what the assistant should infer, what it should ask, what it should never assume, and how implementation decisions would affect the user experience.

03

I stayed through implementation to explain the reasoning. I did not just hand off the framework. I stayed close enough to help the team preserve the decision logic as the work moved through platform constraints.

Internal partner feedback captured the impact:

She pushed and got this back on the docket and formatted it into the output the customers have been asking for.

That mattered because the work did not end with a recommendation. It became something teams could ship, evaluate, and reuse.

Closing

Better models move the boundary.
They never remove it.

AI systems will never eliminate uncertainty.

Better models will infer more. Better context will resolve more. Better retrieval will surface more.

But there will always be a boundary between what the system knows and what it is assuming.

That boundary is where design matters.

The boundary

What the system knows → what it can safely infer → what it is only assuming → where it should ask.

I do not just design what the AI says.

I design how it decides what kind of response the moment deserves, and how teams can tell whether that decision was right.

Designing that judgment is designing the intelligence of the system.