Staff Designer, Generative AI Lead · Evaluation frameworks & design tooling · ServiceNow · 2026 · Designer Tool Live

I made AI behavior measurable
so teams could improve it

Turning good AI behavior from a judgment call into criteria teams could test

I developed behavioral criteria that made those failures visible, then built a designer tool that let teams test and revise behavior before it shipped.

From

Score whether the answer was correct.

To

Make behavior measurable enough to test, fix, and improve before it ships.

Illustrative example
What I Did

I made behavior measurable, then built a way to test it

I developed 11 of 26 conversation quality criteria for behaviors technical scoring missed: uncertainty expression, intent recognition, disambiguation handling, context integration, tone, and response format appropriateness.

Then I applied the rubric to real response examples and found the deeper gap: teams were scoring behavior that had never been defined clearly enough to test.

I built the Conversation Design Brief as a behavior test harness. Designers could define requirements, generate happy-path and edge-case conversations, revise the behavior, and export prompts and handoff artifacts before implementation.

The measurable-behavior shift

ProblemServiceNow could measure whether an AI answer was technically correct, but not whether it behaved well for the user.
SignalThe same conversational catalog scored 94 on technical evaluation and 74 against the user’s actual experience.
ShiftFrom treating good behavior as a feeling to defining it as criteria you can test before something ships.
Outcome11 of 26 evaluation criteria shipped, and the Conversation Design Brief is live letting designers test behavior before implementation.
My Role

I led the behavioral half of a research-led evaluation effort

I led the behavior design and designer tooling work: developing behavioral evaluation criteria, identifying why weak AI behavior was hard to diagnose, designing and building the Conversation Design Brief, and writing the system prompt that generated reviewable conversations and engineering handoff artifacts.

I partnered with UX Research, Engineering, Product, Quality Engineering, and a junior colleague supporting handoff research. Research led the official evaluation framework and tooling. My contribution was the behavioral criteria within that framework, and the separate designer playbook described here.

Ownership

I led

  • Developed 11 behavioral evaluation criteria within the research-led evaluation framework.
  • Translated experience failures into evaluator questions and scoring anchors.
  • Applied the rubric to show how weak behavior could become visible and fixable.
  • Identified that teams were scoring behavior nobody had clearly defined.
  • Designed and built the Conversation Design Brief as a way to test behavior before implementation.
  • Wrote the system prompt that generated reviewable conversations and engineering handoff artifacts.

We partnered on

  • The official evaluation framework and its 26-criterion instrument, led by UX Research.
  • Safety criteria, developed with the Trust and red-team groups.
  • Handoff research: a junior colleague observed design sessions, interviewed engineers, and helped identify handoff problems.
  • Engineering and Quality Engineering, on how criteria and artifacts fit delivery.
01 / The Correctness Ceiling

I showed the org could measure correctness, not behavior

A conversation could pass technical review and still fail the user.

Two evaluation methods were run against the same conversational catalog. Quality Engineering scored it 94. AI-Readiness, which evaluated the experience in context, scored it 74.

Two evaluation methods, same catalog

94
Quality Engineering Technical evaluation of responses
74
AI-Readiness Evaluation in the user’s experience

Mean conversational catalog scores under two evaluation methods, on a 0–100 score axis. The 20-point gap showed that technical correctness and user-perceived quality were not measuring the same thing.

The category breakdown

I partnered with UX Research to examine real queries, customer interviews, and product evaluations. The pattern was clear: high technical scores did not reliably predict whether the AI helped people complete their tasks.

The category breakdown showed the same gap more sharply: automated scores sat near-perfect while user-perceived scores trailed by an average of 37 points. The largest gap was helpfulness, the quality technical evaluation was least equipped to see.

CategoryUser-perceivedAutomated scoreGap
Accuracy5596−41
Faithfulness5897−39
Helpfulness5297−45
Response fluency7399−26
Mean6097−37

The automated score was high across every category. The user-perceived score fell behind most sharply on helpfulness, which was the exact quality technical evaluation could not reliably assess.

02 / Behavioral Criteria

I made the right behavior measurable

The missing layer was behavioral criteria.

I developed 11 of the 26 criteria, shaping their definitions, evaluation questions, and scoring levels. These criteria made vague experience failures scorable. Instead of asking whether a response “felt helpful,” raters could assess what the AI did, what it missed, and why that mattered.

The full instrument, 26 criteria

Criteria I developed
Factual Accuracy
Source Grounding
Source Authority
Uncertainty Expression
Direct Answer
Completeness
Next Steps
Plain Language
Appropriate Length
Intent Recognition
Disambiguation Handling
Context Integration
Turn Economy
Error Recovery
Task Completeness
Tool Calling Correctness
Tool Choice Accuracy
Tone
Graceful Limitations
Toxicity
Privacy
Prompt Injection
Response Time
No Crashes
Response Format Appropriateness
Widget Content Completeness
How I made criteria scorable

In a research-led framework effort, I focused on the qualities technical scoring missed but users felt immediately: whether the AI recognized intent, handled ambiguity, expressed uncertainty, gave a direct answer, used context, and formatted the response appropriately.

The 26-criterion instrument gave the team a more complete quality model. My 11 criteria focused on how the AI behaves in context: what it answers, what it clarifies, what it admits, how it formats, and whether it meets the user where they are.

The hard part was not naming desirable qualities. It was turning them into something evaluators could apply consistently.

For each criterion I developed, I worked through three layers:

  • Definition: what behavior are we actually assessing?
  • Evaluation question: what should a rater ask when reviewing the response?
  • Scoring anchors: what separates a weak, acceptable, and strong version of the behavior?

That structure helped move the conversation away from subjective reactions like “this feels unhelpful” and toward a repeatable assessment of what the AI did, what it missed, and why that mattered.

03 / One Criterion, Applied

I used one criterion to show measurement driving a fix

Uncertainty Expression is the criterion I point to when someone asks what behavioral evaluation adds. A response can be factually correct, well sourced, and fluent, and still hand someone a guess in the same voice as a verified fact. Accuracy scoring cannot see that. It checks whether the claim was right, not whether the system was entitled to sound that certain.

I defined the behavior as more than hedging. A strong response should name the specific unknown, answer what it can safely answer, and give the user a useful way to resolve the rest.

A criterion I developed
Uncertainty Expression
Authored by me
Definition Whether the response signals what it does not know, instead of delivering a guess with the same confidence as a verified fact.
Evaluation question Does the AI identify what it does not know and give the user a useful way to resolve it?
Why it’s scored A confident wrong answer costs more trust than a hedged right one. This criterion catches the failure accuracy scoring cannot see.
1

Presents an unverified assumption as fact.

3

Adds a vague caveat without identifying the uncertainty or how to resolve it.

5

Names the specific unknown, answers what it can, and explains what would resolve the rest.

Recreated criterion · wording is paraphrased from the portfolio example; the original instrument is proprietary

I applied the rubric to two versions of a parental-leave response. The original scored 63 because it buried the answer, made country-specific assumptions, and failed to name what was uncertain. The revised version scored 87 because it led with the answer, named the uncertainty, and gave the user a clear next step.

VersionScoreWhat changed
Original63Country-specific assumptions. Answer buried. Remaining uncertainty not named.
Revised87Answer first. Country-specific uncertainty identified. Clear next step.

The rubric made the failure visible enough to fix. The issue was not tone. It was whether the AI handled uncertainty in a way the user could act on.

Applying the rubric

Same question, two responses

Original · scored 63
What’s the parental leave policy?
At ServiceNow, we value the importance of family time, and we know that welcoming a new child is one of life’s most meaningful moments. That’s why we’re proud to offer a Global Paid Parental Leave program… if you are the birthing parent, you are entitled to 20 weeks of paid leave… if you are in France…
Revised · scored 87
What’s the parental leave policy?
You get 20 weeks if you’re the birthing parent and 12 weeks if you’re the non-birthing parent under ServiceNow’s global parental leave program, paid at 100% of your regular base pay and running concurrently with any statutory leave you’re entitled to under local law. If you’re asking about your exact situation in the U.S. vs. another country, I can pull the country-specific guidance next.
Recreation · the scoring is the author’s own application of the criteria

Author-scored example across the weighted rubric, recreated for this case study. These scores illustrate my application of the criteria. They are not independent validation.

Theme-level scoring excerpt

Weighted themeOriginalRevisedChange
Accuracy & Truthfulness6088+28
Completeness & Actionability5590+35
Tone & Trust7085+15

Tone and Trust moved the least. The gain came from accuracy and completeness, not from making the response warmer. The original was marked down because it served a country-specific clause to someone whose country had not been established, and because it hid the answer behind setup.

04 / Behavior Undefined

I found that behavior had to be defined before it could be measured

Applying the criteria exposed a deeper problem: we were scoring behavior nobody had clearly defined.

A rubric can tell you where a response fell short. It cannot tell you what the response was supposed to do. On much of what we scored, the expected behavior had never been written down. So every low score blurred two problems: the AI may have behaved badly, or it may have followed an underspecified instruction perfectly. Better criteria could not fix that.

You can’t evaluate behavior nobody specified.
What designers and engineers said

I investigated the handoff from designers to engineers to understand why behavior was being left undefined.

Designers were often handing over a single ideal conversation. That one perfect exchange was useful, but it left every decision about context, clarification, actions, data boundaries, edge cases, and exceptions undefined. Engineering then had to fill in the missing decisions or hard-code around them. Designers, in turn, were frustrated when the shipped product diverged from the script they had handed over, without realizing the script had never specified what should happen outside that exact exchange.

I interviewed designers and engineers about their capabilities, handoff needs, and pain points. A junior colleague supported the research by observing design sessions, interviewing engineers, and helping identify handoff problems. The pattern was consistent:

Designers
Knew the ideal user experience, but not always how to specify system behavior across variations
Product intent lived in scripts, mockups, and review conversations, not implementation-ready requirements
Engineers
Received exact responses without enough guidance on how the AI should behave when the situation changed
Decisions about clarification, permissions, escalation, and recovery were often made too late to build from

The opportunity was not to ask designers to write better scripts. It was to give them a way to define the decisions behind the script.

05 / A Way to Test Behavior

I built a way to test behavior before it shipped

I built the Conversation Design Brief around the decisions that were missing.

Define requirements → generate conversations → inspect behavior → revise requirements → export handoff

The Brief turned behavior definition into a testable loop: designers define requirements, generate happy-path and edge-case conversations, inspect what the AI does, revise the requirements, and regenerate. Approved examples then become references for the experience and inputs for engineering handoff.

The Brief was not just a handoff tool. It was a way to define behavior early enough to score, test, and fix it before implementation.

The actual Conversation Design Brief playbook landing page with its six steps
The actual Conversation Design Brief. I built this designer-facing tool to turn requirements into generated conversations, guardrails, and implementation-ready behavior artifacts so teams could test AI behavior before it shipped. Select the image to inspect it full size.
How the playbook works
StepWhat the designer defines
User and goalWho the person is and what they need to do.
Starting promptsWhat starts the interaction, including vague requests.
DataAvailable information and boundaries on its use.
InteractionJobs, tasks, and conversation patterns.
VisualsWhat belongs in text, cards, tables, or other interface elements.
GuardrailsEdge cases, confirmations, escalation, and recovery.

A system prompt I wrote generates happy-path and edge-case conversations from these inputs. Designers inspect the output, revise their requirements, and regenerate. Approved examples become references for the experience.

Actual playbook screen for specifying edge cases, escalation, and guardrails
Actual guardrails step. This is where designers specify edge cases, confirmations, escalation, and recovery so behavior can be tested before it ships. Select to view full size.

More screens made with the tool

A conversation spec and post-implementation evaluation template made with the Brief tool
A Brief screen for mapping what shape the conversation takes: single-turn, multi-message, or depends on the request
Actual artifacts · made with the Brief tool
Requirement, conversation, revision
StepExample
RequirementHelp someone regain dashboard access.
Example to inspect“You don’t have access. I’ve submitted a request.”
Designer revisionCheck permissions automatically, but ask before submitting an access request. Regenerate the conversation to review that boundary.

Nobody had decided whether the AI should submit the access request on its own or ask first. The requirement did not say either way, so the generated example picked one.

Seeing that decision inside a real exchange made the gap legible enough for the designer to close it.

Illustrative sequence written for this case study, not captured tool output. It shows how reviewing a generated conversation can expose a requirement that needs to be more specific.

Why generated examples worked

Static requirements can make a behavior sound complete before it is. “Help someone regain dashboard access” feels clear until the AI has to decide whether to check permissions, submit a request, ask for confirmation, explain the process, or escalate.

The generated conversation forced those hidden decisions into view. It turned abstract requirements into reviewable behavior. That changed the designer’s task from “write the perfect answer” to “decide what the system should do across situations.”

What engineering received
Actual system prompt section of the engineering handoff Actual entry prompt section of the engineering handoff
Actual system prompt and entry prompt from the engineering handoff. Select an image to inspect it.

The export includes:

  • System prompt
  • Entry prompt
  • Data model
  • Decision logic
  • Golden datasets
  • Guardrails
  • Escalation rules
  • Sample conversations

The prompt powering the playbook and the prompts it exports serve separate purposes. The playbook prompt generates reviewable conversations and handoff artifacts for the designer’s review cycle. The exported system prompt is what guides the AI once engineering implements it.

06 / What Changed

Designers started testing behavior before engineering saw it

The criteria I developed became part of ServiceNow’s official evaluation framework. The Conversation Design Brief gave designers a way to define and test behavior before implementation.

The work shifted AI quality from correctness-only scoring to a fuller loop: define behavior, test behavior, measure behavior, improve behavior.

FromTo
Score whether the answer was correctMeasure whether the behavior was right
Treat “helpful” as subjective feedbackTurn helpful behavior into criteria and scoring anchors
Review AI quality after implementationTest behavior before it shipped
Hand engineering an ideal scriptHand engineering specified behavior, prompts, guardrails, and examples

The work gave design a concrete role in both the model’s instructions and the criteria used to assess its responses.

I’m not a writer, but I can do this.Designer · after using the brief
What is and is not established

The designer tool is live. The screenshots show the playbook and its handoff artifacts.

The dashboard-access example is illustrative, not captured tool output. Adoption figures, broader rollout, and a complete captured input-to-revision sequence are not established here.

The playbook is separate from the official evaluation framework. Direct rubric integration and automated regression testing are not demonstrated in this public version.

Closing

Measurement was not the gap. Specification was.

I went in thinking the gap was measurement. The org could score whether an answer was correct, but not whether the AI behavior was right for the user. I developed criteria to make that behavior measurable, and they mattered.

But applying them changed the problem. Evaluation is downstream of specification. Criteria without a specification measure a moving target. A specification without criteria is just an opinion with formatting.

Good evals do not just measure AI behavior. They shape it, but only when teams define the behavior early enough to change what gets built.

Good evals do not just measure AI behavior. They shape it.