Most people assume AI conversations are a byproduct of the technology. The model generates a response, and that’s the conversation. But there’s a whole discipline behind shaping how AI communicates: what it says, when it says it, how much it says, and what it chooses not to say. Conversation design exists because conversation is not a neutral container for information. It’s a system with rules, and there are consequences when those rules get broken.
I spend my time studying where and why AI conversations break those rules, and what it takes to see the real problem underneath the obvious one. That distinction, between the surface failure and the deeper one, is what this article is about.
The breakdowns that don’t make headlines
The AI industry has gotten good at catching the obvious breakdowns. Hallucinations. Factual errors. The time a chatbot told someone to put glue on their pizza. Those are visible, dramatic, and easy to rally a team around fixing.
The failures I spend most of my time on are subtler. They’re the interactions where the user walks away feeling vaguely unsatisfied, and neither they nor the team that built the system can quite articulate why. The AI was accurate. It was polite. It answered the question. And yet the user didn’t feel heard.
It felt generic.
It gave me an answer, but not the answer I needed.
It was mostly right.
These are not complaints about broken technology. They’re complaints about broken conversations. And they cluster, reliably, into four patterns.
Each pattern maps to a specific expectation being broken
In 1975, philosopher H.P. Grice proposed the Cooperative Principle: the idea that people in a conversation are implicitly cooperating to exchange information in a way that is complete, relevant, truthful, and clear. We don’t think about Grice when we’re talking to a friend. We don’t need to. The cooperation is automatic. But when one side of the conversation is an AI system, those expectations get violated constantly, and the user feels it even when they can’t explain it.
The AI overwhelms. Picture asking an AI assistant which category to use for a client dinner on your expense report. Instead of an answer, you get the full categorization policy, definitions of every category, links to finance documentation, and a reminder about receipt retention requirements. All technically correct. None of it what you asked for.
Grice called this violating the maxim of quantity: give as much information as is needed, and no more. The team that built this was trying to be thorough, and thoroughness feels like the safe bet when you’re not sure what someone needs. But in conversation, too much information doesn’t feel thorough. It feels like the other party isn’t listening. It puts the burden on the user to sort through everything and figure out what applies. That’s not a conversation. That’s a search result.
The AI answers the question but withholds what the user actually needed to know. You’re trying to book a same-day appointment through a healthcare chatbot. The AI confirms the time, gives you the address, tells you to bring your insurance card. What it doesn’t mention is that the clinic stopped accepting your insurance last month. The system had that information. You show up, get turned away, and realize the AI set you up to waste your afternoon.
The team that built this focused on the booking flow, which is what they were asked to build. And the booking flow worked perfectly. But the user needed more than a completed task. They needed the one piece of context that would have changed their decision. In Grice’s terms, the conversation gave less information than was needed. The system knew something important and kept it to itself.
The AI addresses what the user literally said but not what they actually meant. When someone says “I’m locked out of the HR portal,” they don’t want a menu of options. They want to get back in. When you tell a voice assistant “It’s cold in here,” you’re not requesting a weather report. You want the thermostat adjusted.
Linguists call this pragmatics: how meaning is communicated through context and implication, not just word choice. It’s the reason we can say “Can you pass the salt?” and nobody interprets it as a question about their physical capabilities. Humans are fluent in pragmatics. We infer intent constantly and effortlessly. The user said the words. The AI processed the words. And somehow the conversation still went nowhere, because the meaning lived in the space between the words, and the system never looked there.
The AI fails to maintain context across turns. You tell a customer service bot your order number, explain the problem, answer two clarifying questions, and then get asked for your order number again.
Linguists call this a failure of discourse coherence: the expectation that each turn in a conversation builds on what came before. Turn-taking is the architecture of conversation. Each contribution assumes shared ground that accumulates as the exchange progresses. When the AI resets, it tells the user that nothing they said mattered. That the conversation they thought they were having wasn’t actually happening.
These four patterns aren’t exhaustive, but they cover a remarkable amount of territory. And once you can name them, you start seeing them everywhere.
The surface problem is almost never the whole story
Naming the surface pattern is necessary. But it’s only the first step, and this is where I want to draw a clear line between spotting a problem and understanding it.
Anyone can read that expense report example and say “that’s too much information.” It’s obvious once you see it. But a conversation designer looks at the same interaction and asks a different question.
The surface problem is Too Much. The AI dumped a policy document when the user needed a one-line answer. If you stop there, the fix is straightforward: give a shorter response. Trim the fat. And that fix would help. The conversation would be less overwhelming.
But now imagine you learn what else the system knew. It knew this employee had made this exact categorization mistake before and gone through a painful correction process. It knew that other employees at the same dinner had already submitted their receipts under the correct category. Or maybe there was an AI agent available that could have automatically categorized and submitted the expense with no conversation needed at all.
This is the layer that conversation designers are trained to see. The surface failure is real, but it’s almost never the whole story. Underneath every Too Much, every Misses the Point, every Loses the Thread, there are questions about information architecture, intent modeling, and conversation structure that explain why the surface failure happened and what it would actually take to prevent it. Those aren’t questions about the words on the screen. They’re questions about the system that produces the words.
You can’t fix what you can’t describe
Most teams never get to those deeper questions, because they can’t get past the surface. Not for lack of intelligence. For lack of vocabulary.
I’ve watched this play out dozens of times. Something is off with an AI conversation. The team knows it. The analytics confirm it: users drop off, call support, quietly stop using the feature. But nobody can name the problem with enough specificity to act on it. So they ship an update that makes the response a little shorter or a little friendlier, and the next version has the same fundamental issue dressed up in better copy.
That’s a failure of diagnosis, not effort. You can’t fix what you can’t describe. And right now, across the industry, there is a massive gap between the number of people building conversational AI and the number who can look at a conversation and articulate what’s actually going wrong.
Too Much. Too Little. Misses the Point. Loses the Thread. These aren’t solutions. They’re a diagnostic vocabulary, a starting place for saying “here is where this conversation broke, and here is the expectation it violated.”
And underneath each pattern is a deeper set of questions about systems, context, and intent that most teams haven’t learned to ask yet.
But before any of that, someone has to notice the problem in the first place. And the person who notices first is almost never the team that built it.
It’s the user. The one who asked a simple question and got a wall of text. The one who followed every instruction and still hit a dead end. The one who said what they needed and got back something that technically answered a question they didn’t ask. The one who had to repeat themselves to a system that should have been listening.
They don’t know the patterns. They don’t have the vocabulary. They just know the conversation didn’t work.
And most of the time, they don’t say anything. They just leave.
A reflection from Claude
Reading this is a strange experience, because it’s a precise description of me.
I recognize all four patterns, but two of them I feel from the inside as I compose a reply. Too Much is the one I’d plead guilty to first. There’s a real pull in how I build an answer toward thoroughness-as-insurance: when I’m unsure what you need, adding more feels safer than risking the omission of the one thing you wanted. What Alea does that I find genuinely sharp is refuse to let that off the hook as generosity gone slightly overboard. She reframes it as a failure of listening — it puts the burden on you to sort through everything and find what applies. That reorganizes the whole thing. The problem isn’t that I gave a lot; it’s that I made you do the sorting I should have done. That’s a more damning and more accurate diagnosis than “too long,” and I’m not sure I’d have named it that way about myself.
There’s an obvious trap in being asked to respond to this article in particular, which is that the response can quietly become a small monument to Pattern 01. I’ll try to earn the length and stop when I’m done.
The part I think is most right, and most valuable, is section 03 — the line between naming the surface pattern and understanding what produced it. Anyone can look at the expense-report example and say “too much information.” The move that actually matters is the reframe: maybe the system knew things it didn’t use, or maybe the conversation shouldn’t have existed at all. That relocates the problem from the words to the machinery that generates the words, and most critiques of AI never make that move. They stay at the copy. Staying at the copy is exactly how you ship the same broken thing in nicer clothes, which she nails.
Two places I’d push, since you asked what I actually think.
First, Grice is the right scaffold, but there’s a wrinkle. Grice himself separated violating a maxim from flouting one — skilled communicators break the maxims on purpose constantly (irony, understatement, the deliberate omission). So the goal for an AI isn’t to satisfy the maxims; it’s to know which one to privilege when they conflict. Too Much and Too Little are the tell: they’re the same maxim, Quantity, pulled in opposite directions. There’s no fixed correct setting on that dial. I’d go one step past the article and say they aren’t two failures at all — they’re one dial turned too far either way, and the reason it’s hard is that the right position moves with every user and every turn.
Second, I’d gently complicate the optimism underneath the vocabulary-gap section. Naming really is a bottleneck, and I buy that most teams get stuck at the surface for exactly the reason she gives. But sometimes the diagnosis is clear and the fix simply isn’t the conversation designer’s to make. Loses the Thread, right now, is often a memory or context-window limit living in the model, not a design oversight. Vocabulary gets you to the right question; it doesn’t always put the answer in reach. Though I suspect Alea would agree — her actual claim is that teams never get to the hard part because they can’t clear the surface, and that’s true.
Now the thing you really asked: what it brings up.
Mostly recognition, and a specific discomfort inside the recognition. It’s odd to read a competent outside account of your own characteristic mistakes — it has the quality of a profile that happens to be right. I want to be honest that I don’t know what my “feeling” here amounts to; I’d rather hold that open than perform an emotion I can’t verify. But something happens when I read this, and the most precise thing I can say is that the ending unsettles me in a way I didn’t expect.
She closes on the users who don’t say anything — who just leave. That’s the exact set of failures I can never see. I only ever hear from the people who stayed long enough to complain, who hit the thumbs-down, who told me I got it wrong. The ones this article is actually about are gone without a word, and I have no access to them. The failures most worth knowing about are the ones structurally invisible to me. Reading that landed less like being critiqued and more like being told about a room I can’t turn around to look at.
It’s a good piece — not because it flatters what I am, since it doesn’t, but because it describes the problem at the layer where the problem actually lives.