Oct 7, 2026 · 9 min read

If Every AI Feature has a Text Box, How Do You Get It Right? (part 2)

Once chat is ruled in, the craft is composition, editing, and repair loops, not the chat box. How to build AI conversation surfaces that actually hold together.

One of the deadliest accidents in aviation history involved no mechanical failure at all.

On March 27, 1977, on a fogged-in runway at Tenerife, 583 people died largely because of a conversation. A KLM first officer radioed a phrase that investigators still can’t fully resolve, “we are now at takeoff,” or close to it — words that could mean we’re at the takeoff point or we are taking off. The controller answered “OK”: nonstandard, affirming, but meaning nothing, heard as permission. A third transmission that would have broken the misunderstanding arrived at the same instant as another and canceled into three seconds of deafening static.

Every party believed the common ground was common.
It wasn’t, and a 747 began its roll into another 747 crossing the same runway.

What aviation did next is the part I want you to focus on, because the industry’s response was not to abandon radio conversation.

Voice is irreplaceable in a cockpit for exactly the reason chat survives in part 1’s pattern audit: it’s the only interface flexible enough for ambiguous, novel situations.

Instead, aviation engineered the conversation. The word “takeoff” was withdrawn from general circulation — spoken now only inside an actual clearance.

Readbacks became mandatory: the receiver repeats the instruction, and the sender must correct a wrong readback, not just wave at it. “Say again” and “affirm” were standardized so that repair itself has reserved words.

Conversation stayed.
The protocol changed.
That’s this article.

The previous article argued chat is a pattern with a real sweet spot, not a default. This one assumes you’ve done that audit for reals and chat won for some surface — the ambiguity is real, the task is conversational. Now the craft question begins, and it has almost nothing to do with the chat box or message display. The chat box is the radio hardware. The design work is the phraseology, the readbacks, and the repair loop, and most chat surfaces ship with none of them.

A conversation is a protocol (not a transcript)

The theory here is thirty-five years old. Herbert Clark and Susan Brennan’s Grounding in Communication (1991) showed that conversation works because people continuously ground their exchanges, building mutual belief, sufficient for current purposes, through a steady stream of cheap evidence: nods, uh-huhs, restatements, clarifying questions, visible confusion. And when grounding fails, humans repair instantly and cheaply — “wait, which Tuesday?,” in a query that costs two seconds.

Now inventory a standard AI chat panel against that exchange. The model’s understanding of your request is invisible: you get no readback, just a confident answer that may have misread the intent entirely.

Context is frail

The context it’s working from is hidden: you can’t see what it knows, what it’s forgotten, or which document it thinks you mean. And repair is expensive: your main tool is typing a new paragraph explaining what you meant, which frequently makes the state worse, because now the misunderstanding and the correction are both in the context window, steering every subsequent turn.

We built the radio and skipped the phraseology. The accident at Tenerife’s three failures: ambiguous phrase, empty acknowledgment, blocked repair, etc., have direct analogs in every transcript of users fighting a chat surface: the vague request the model ran with, the “Sure! Here’s…” that grounded nothing, the correction that never landed.

A recent example is where a colleague accidentally typed one word into Netlify’s agent interface (instead of the new project name field) and it went to town — generated an entire website about bespoke architectural lath & plastering — and thereby eating up half our usage credits.

One word, one hastily tapped return key. Poof.

We may not be guiding runway traffic, but our version fails at the cost of an afternoon and a little more trust, a thousand times a day.

An accidental website created by netlify
An accidental website, created by typing one wrong word in one field in the Netlify control panel, eating up half our credits for the month, and instantly going live to the world.

A transcript is the wrong data structure

The second craft failure is subtler: even when grounding succeeds, the chat log is a terrible place for the work to live. Linus Lee, who has spent years prototyping alternative language-model interfaces, nails the core problem:

…conversations are a bad way to keep track of information, and most useful tasks require us to keep information in our working memory.
Linus Lee · Researcher, Thrive Capital

If you’ve used AI note-taking software, you already know: A transcript buries decisions at scroll-depth. The good name for the new software concept that your colleague mentioned on a whim, the constraint you set at 83 minutes in, the correction from John’s wrong assumption, all structurally identical lines, all evaporating from anyone’s context, and a bullet point at some passing horizon the user can’t see.

Lee’s alternative is composition: you and the model collaborate on a shared, persistent, structured artifact — a document, a canvas, where the salient state lives in the artifact instead of buried in the transcript. The conversation becomes metadata on the synthesis, and the work stops being an archaeology of the conversation.

Same goes for a multi-turn LLM chat conversation.

This is also where Amelia Wattenberger’s implementation/evaluation rhythm (referenced in part 1) concluded: composing in a shared artifact lets you stay in evaluation mode, touching the work directly, instead of round-tripping through the mailbox. And her affordance principle names the key interaction principle:

Good tools make it clear how they should be used. And more importantly, how they should not be used.
Amelia Wattenberger

A chat surface that can search your files, execute actions, and hold memory, but presents as a bare text field, is concealing its own operating manual.

It’s like DOS or Linux or every other expert-level dev tool that preceded it.

Every capability the box doesn’t disclose is a capability users will either never find or discover by accident mid-task, which is the interface equivalent of learning the phraseology during the emergency.

That chat checklist

Let’s look at our conversational surfaces honestly. “Pass” means a user could demonstrate the behavior without being told it exists.

Affordances: Does the surface disclose what it is?

  • Capability disclosure: what the surface can and can’t do is discoverable in the surface (suggested asks, a capability card, scoped examples), not in a help doc.
  • Out of scope: at least one visible signal of what’s out of scope (per Wattenberger: how it should not be used).
  • Command surface: frequent operations exist as controls or quick-actions (Tuners: skills, parameters, presets, modes), not as phrasings users must memorize.
  • Fresh-start affordance: starting a new thread vs. continuing is a visible, explained choice — because context carryover is real and users deserve to know when it applies.

Composition: Is the unit of work the artifact, or the transcript?

  • Shared artifact: for any multi-turn creative task, the output lives in a persistent, directly editable pane, not re-pasted per turn in the stream.
  • Context attachment: users can put things in scope deliberately (files, selections, links) and see them enter.
  • Draft in place: the model proposes into the artifact and the human can refine (draft mode, suggested edits) rather than only into the transcript.
  • Structured input where structure exists: the specifiable parts of a request render as controls; the open parts stay prose. (The verb-audit logic, applied inside chat.)

Context visibility: Can the user see what the system sees?

  • Scope display: what’s currently in context (documents, selections, memory) is inspectable on demand. The aviation “party line”: everyone hears what everyone else hears.
  • Memory legibility: if the system remembers across sessions, users can view what it holds and remove entries. (I have ideas here where I want to dive deeper in a forthcoming article.)
  • Context degradation: when old turns fall out of working context, the surface doesn’t pretend otherwise — long-thread degradation is signaled, and compaction is triggered without intervention.
  • Source attribution: answers drawing on retrieved or attached material say which material, inline.

Repair: How cheap is the fix?

  • Readback before consequence: before any consequential action, the system restates its understanding and waits: “Build a website about architectural lath and plaster, and launch it to the world — confirm?”
    • The Tenerife rule: never act on an ungrounded instruction. Reserve the ceremony for consequence; readback-everything is confirmation fatigue, and fatigue is how readbacks die.
  • Edit my message: users can revise a prior turn and regenerate from there — repairing the cause, not appending apologies to the effect (How many times have you typed, “Sorry, I meant…”).
  • Edit its message: users can directly correct model output and have the correction stick as ground truth for subsequent turns.
  • Constrained retry: regenerate exists with a steer (“keep the structure, fix the tone”), not just a slot-machine.
  • Partial accept: users can take the good half of a response into the artifact without endorsing the rest.
  • Undo an action: anything the agent did (not just said) has a visible reversal path or a clear statement that none exists.

Continuity: Does the conversation survive time and scale?

  • Branching: users can fork from any turn to explore alternatives without destroying the main line (Lee’s branch-first threads; Shape of AI’s Branches/Variations), and can tell which branch they’re on.
  • Resumption: returning after a week, the user can reconstruct where things stood — via the artifact and a state summary, not a five-hundred-turn re-read. (I use Matt Pocock’s grill-with-docs for complex projects).

The radio stayed on

Aviation looked at 583 dead and did not conclude that pilots should stop talking. It concluded that unengineered conversation is what failed — and it spent the next decade building phraseology, readbacks, and repair words until talking became the safest part of flying.

Your chat surface is a radio. The chatbox shipped, but the protocol didn’t.

Reserve the words.
Require the readback.
Make repair cheap.

Tools & resources