Updated July 20, 2026 — The world is wild, and moving fast!
A head of research told me about a customer interview that died in its first ninety seconds. The participant joined the call, smiled, made small talk — and then a third attendee slid into the participant list: the recording bot her team’s AI notetaker sends into every meeting. The participant’s eyes flicked to it. “Is that recording? Where does that go?” The researcher explained. The consent form had been signed days earlier; the disclosure was made; every box was checked. The participant said, “Sure, fine…”
…And then spent the next forty minutes giving the careful, lawyered answers of someone talking near a microphone instead of to a person.
The transcript, of course, came back flawless. It captured everything except anything true.
Nothing in that session violated a checklist. The failure happened somewhere our discipline mostly doesn’t look: the session wasn’t running on the discussion guide; it was running on trust — and the balance went negative before the first real question was asked.
Josh Clark has been arguing since 2017 that machine intelligence is a design material — something to shape the way we shape layout or prose — and his new book with Veronika Kindred, Sentient Design, builds a whole practice on that idea. I buy the argument. But I want to push it one layer down. AI is the medium. The material you’re actually shaping — the thing that holds its form or cracks under every single decision you ship is the user’s willingness to rely on you. Google’s PAIR team gives the cleanest definition of it I’ve found:
Trust is the willingness to take a risk based on the expectation of a benefit.
Read that as a designer and something clicks: “willingness to take a risk” is a balance. It accumulates. It erodes. Every AI surface you ship either deposits into it or withdraws from it.
Here’s why you came here: a review lens I’ve been calling the Trust Ledger — a way to walk any AI surface and mark, decision by decision, what it deposits and what it withdraws. It’s a critique habit, not a process. Your team can run it this week.
You’re launching into a deficit
First, because the conditions changed underneath us, and a lot of roadmaps haven’t caught up.
Nielsen Norman Group’s State of UX 2026 — Jan 2026, names the problem:
In 2026, trust will be a major design problem for AI experiences.
Their diagnosis is that we’ve moved from AI hype to AI fatigue — people who’ve been burned by half-ready AI features are more hesitant to adopt the next one, and rebuilding that confidence requires fundamentals: transparency, control, consistency, and support when the system fails. Notice what those four things have in common. None of them is a model capability.
All decisions are design decisions
The population-level numbers say the same thing. The KPMG and University of Melbourne global study of trust in AI: 48,000 people across 47 countries, found only 46% are willing to trust AI systems, and willingness to rely on AI actually fell between 2022 and 2024, from 52% to 43%, even as usage exploded. Pew’s February 2026 survey found 71% of US adults believe increased AI use will make their personal information less secure. Seventy-one percent. That’s not a skeptical fringe; that’s your typical user.
Your AI feature does not launch into neutral territory. It launches into a preexisting deficit, established by every “delightful” sparkle-button and every silently harvested dataset your users have already lived through.
You don’t start at zero. You start below it.
Huge withdrawals happen off-screen
Here’s where it gets awkward, and the reason this article exists:
The AI products we’re shipping don’t just process what people type into them. They listen to meetings through transcripts. They read behavior in feeds and social apps. They infer patterns from signals nobody consciously handed over. In my own org, the sharpest version of this is research — we sit on hours of recorded customer conversations, and every AI vendor in our stack would love to synthesize them. The people in those recordings consented to a human taking notes. Whether they consented to becoming training data is not clear, or overtly represented.
And the recent record shows exactly what happens when companies answer that question with merely a terms-of-service update:
- Zoom, 2023. New ToS language claimed rights to use customer content for AI training. Days of public backlash later, Zoom rewrote the terms entirely and called the original language a process failure.
- Slack, May 2024. Users discovered customer data fed Slack’s models by default, with opt-out only via an admin email. The policy language got rewritten within a week.
- LinkedIn, September 2024. A default-ON setting sharing user data for generative-AI training appeared before the privacy policy was updated; the UK regulator’s pressure suspended training on UK data.
- Otter.ai, 2025. A class action — filed by a man who wasn’t even an Otter user — alleges the Notetaker records meeting participants without their affirmative consent and trains on the recordings. The suit is pending.
There’s a trend here. It’s the default nobody chose. It’s the clause nobody read. It’s the data flow nobody drew in the Figma file. Not one of these withdrawals was made by a designer. They were made in legal reviews and data-pipeline decisions. But every one of them was experienced in the product, as betrayal, by users.
Design didn’t make these withdrawals. Design just gets the overdraft notice.
And the overdraft is expensive. Thales’s 2025 Digital Trust Index found 82% of consumers had abandoned a brand in the previous year over concerns about how their data was used; Cisco’s 2024 privacy survey found more than 75% of consumers won’t buy from an organization they don’t trust with their data. The takeaway for us as leaders: a trust review that only covers the screens is auditing the teller window while the vault door stands open. Your ledger has to include entries design doesn’t traditionally own.
The goal isn’t trustmaxxing
Here’s where the ledger metaphor needs a correction, because the obvious reading: maximize deposits, minimize withdrawals, biggest balance wins is wrong, and the research is clear about why.
The foundational paper here is Lee and See’s 2004 Trust in Automation: Designing for Appropriate Reliance, and its core argument has aged perfectly into the AI era: when trust is miscalibrated relative to what the system can actually do, both directions fail. Too little trust and people abandon a system that would have helped them. Too much trust and they rely on it past the edge of its competence. I’d argue this is worse, because now the system’s failures become their failures. Google’s PAIR guidance says: because AI products are statistical, the user shouldn’t trust the system completely — the whole first principle of their trust chapter is “help users calibrate their trust.”
Overtrust isn’t hypothetical in the field, either. NN/g’s Caleb Sponheim calls the failure mode magic-8-ball thinking — users stop verifying answers and just trust them. NN/g’s diary studies keep finding inflated trust in AI tools, not the skepticism we’d design for by instinct.
Which brings me back to Clark, who saw this coming before most of us had shipped anything with a model in it:
Our answer machines have an over-confidence problem.
So the ledger needs a fourth entry type beyond deposits and withdrawals: the counterfeit deposit. A confidently presented wrong answer feels like a deposit, but it’s counterfeit currency, and it bounces the moment the user discovers the error. As Page Laubheimer puts it in NN/g’s hallucination guidance, “when a system lies to you some of the time, it becomes hard to trust it at any time.” Counterfeits are the most dangerous entries in the book, because they inflate the balance right up until they collapse it. A hedge rendered honestly, “here’s what I’m not sure about, here’s the source, check me,” is a smaller deposit, but it clears.
This is also why I think it’s dangerous letting teams treat “make it feel smarter” as a premium good. NN/g’s recent research suggests people trust AI more when it seems smart than when it performs sentience — emotional affect can actively reduce trust in task-oriented work. Personality is decoration. Calibration is structure.
But doesn’t transparency have a cost?
Here’s the objection I get from smart PMs when we push consent and disclosure: every disclosure is friction, and friction kills adoption. They’re not wrong about the mechanics. There’s even hard evidence that transparency itself can withdraw trust: Schilke and Reimann’s 2025 paper The Transparency Dilemma ran thirteen experiments and found that people who disclose their AI use are “trusted less than those who do not.”
But. Read the same paper one clause further: being exposed by a third party damages trust considerably more than disclosing yourself. That reframes the whole trade. Disclosure isn’t a choice between a withdrawal and no withdrawal — it’s a choice between a small, planned withdrawal now and a catastrophic, uncontrolled one later. Zoom, Slack, and LinkedIn each ran the experiment for us at production scale, and got the same result thirteen lab studies got.
So the design question isn’t whether to spend trust on transparency. It’s where the spend buys the most. PAIR’s guidance is that trust-building is slow and deliberate and starts before first use — which means the cheapest deposits happen early, in expectation-setting, where they cost nearly zero friction. Microsoft’s Guidelines for Human-AI Interaction puts its two trust guidelines first for a reason: make clear what the system can do, and make clear how well it can do it. Both are onboarding-and-framing moves. Neither adds a checkbox to a task flow.
What burns the friction budget without depositing anything is consent theater: the blanket disclaimer under every response, the wall-of-text modal before the user has seen any value, the “I agree” nobody reads. That’s friction spent on legal cover, not on trust. Real consent is asked at the moment the data becomes ambient, in language a human would use, with a no that actually works.
It costs more design effort and less user patience.
Here’s what that looks like shipped, because the contrast is instructive. My researcher friend didn’t fix her sessions with a better consent modal — she switched tools, to Granola, an AI notepad that never joins the call at all. It runs locally on her laptop, transcribes from system audio, and deletes the audio once transcription is done. No bot in the participant list. No third attendee to lawyer up at. The meeting looks like a conversation again.
But read that move through the ledger, because it’s subtler than invisible is better. Removing the bot removes the ambush, and it removes the interface’s only disclosure signal. The consent moment doesn’t disappear; it transfers from the UI to the human. Granola is unusually candid about this hand-off: its own docs put consent responsibility on the user and strongly recommend “always informing participants that you’re taking AI-enhanced notes,” and its participant-privacy playbook is blunt that “disclosure should happen before the recording begins, not after” — name it in the invite, ask a yes-or-no out loud, let the answer land in the notes. So the ledger entry is contingent: the tool deposits candor and honest data handling, and hands your team an IOU — the disclosure only happens if your process makes it happen. (And run the vendor itself through the ledger while you’re there: by default, anonymized data can improve Granola’s models, with an opt-out in settings that enterprise plans flip org-wide. That’s a priced, findable withdrawal — which is exactly the standard.)
My friend’s version of the IOU getting paid is one human sentence at the top of every session. That’s the friction budget spent right: a bot ambush traded for a sentence.
Friction isn’t the enemy of trust. Unspent friction in the wrong place is.
Taking it further than a research session, imagine if you were on a first date, and you found out that this happened:
Chang has also taken to recording most of her first dates, usually with Granola. She puts her phone out on the table or her chair and then once the date is through, she has the feed automatically sent to Anthropic’s Claude to report back on how she did.
The Trust Ledger
Everything above compresses into a review lens you can bolt onto your existing critique ritual. No new meeting, no new deck template. When an AI surface comes through review, you walk its moments and mark entries in four categories:
- Deposit: the surface honestly increases the user’s warranted willingness to rely on you: a cited source, an accurate confidence signal, a consent ask that respects the answer, a graceful and honest failure.
- Withdrawal: the surface spends trust: a default that serves you rather than the user, an ambient data grab, friction that exists for cover rather than clarity. Withdrawals aren’t forbidden — they’re priced. Some are worth it. All of them must be deliberate.
- Counterfeit: the surface fakes a deposit: unearned confidence, fake precision, fluency standing in for accuracy, sentience cosplay. These get flagged hardest, because they read as wins in the demo and detonate in the field.
- IOU: a promise about future behavior: “we never train on your content,” “you can delete this anytime,” “a human reviews every flag.” IOUs are fine — but every IOU is a future withdrawal the day it’s broken, so each one gets verified against what the pipeline actually does before it ships.
Here’s a compressed example: an AI feature that summarizes a customer-research repository, which is about as ambient-data-heavy as design tooling gets:
| Moment | What ships | Ledger entry |
|---|---|---|
| First encounter | “✨ Ask AI about your customers” button, no framing of limits | Counterfeit — sparkle promises magic it can’t price |
| The ask | Repository ingests every recorded call by default; participants consented to recording, not synthesis | Withdrawal — silent, and it’s other people’s trust being spent |
| The act | Summary cites and links the source clips | Deposit — verifiable beats fluent |
| The act | “Users prefer option B 3-to-1” derived from six interviews | Counterfeit — fake statistical precision |
| The miss | No path when the summary is wrong; researcher must simply know to distrust it | Withdrawal — failure with no floor |
| The promise | “Your data never leaves your workspace” in marketing copy | IOU — verified against the vendor contract, or cut |
Fifteen minutes of walking that table with a team produces a sharper roadmap conversation than most quarter-long “responsible AI” initiatives, because every row is a specific, arguable, fixable decision.
If you want to run a Trust Ledger review
Here’s the sequence I’d hand to anyone who wants this up and running.
-
Pick three surfaces, not an audit. Choose by exposure, not by flagship status: (a) the surface that touches the most ambient data — transcripts, behavioral signals, anything users didn’t consciously hand you; (b) the surface with the highest-stakes failure; (c) the surface making the biggest promise. Decision point: your PM will lobby for the flagship feature because it’s demo-ready. Resist. The flagship gets scrutiny for free; the notetaker-shaped things rot in the dark.
-
Recruit one non-designer who can see the pipeline. A data scientist, a platform engineer, or your privacy counsel — someone who knows where the data actually goes, because half your withdrawals live in systems no designer drew. Trade-off: it slows the session and imports legal vocabulary into critique. Accept the cost for the first pass; a ledger that only covers pixels certifies the teller window and ignores the vault.
-
Walk five moments per surface, and mark entries. First encounter, the ask, the act, the miss, and — the one teams always skip — the discovery: the moment a user learns something about how the system works that you never told them. If the discovery moment reads worse than the disclosure would have, that’s the transparency dilemma telling you you’ve chosen exposure over disclosure. Rule for arguments: if the room debates an entry for more than two minutes, mark it a withdrawal and move on — genuine deposits don’t need a defense counsel.
-
Hunt the silent withdrawals. Go through defaults, retention, training use, and what changes when the model updates — the Zoom/Slack/LinkedIn class of entries. The test question I use: what would a journalist find? If the answer makes anyone shift in their chair, it’s a withdrawal, currently scheduled to be discovered by someone else. Decision point: fixing a default is a product decision, not a design one — so your output here is a named, priced entry you escalate, not a redesign you sneak in.
-
Price every IOU. List each promise the surfaces make — in UI copy, empty states, marketing pages. For each: can someone in the room verify it’s true today, and who gets notified when the pipeline changes tomorrow? An IOU nobody can verify gets cut from the copy before it becomes a class-action exhibit. This is where you’ll find marketing has been writing checks your data pipeline can’t cash.
-
Fix calibration before deficits. Counterintuitive, but do the counterfeits first — the overclaiming button, the fake precision, the hedge-free wrong answer — before the underfunded areas. A deficit costs you adoption; a counterfeit costs you the ability to recover, because it converts every future failure into evidence of bad faith. Show confidence only where a user can act on it — PAIR’s guidance, and the cheapest calibration fix there is: if a confidence signal changes no decision, cut it.
-
Allocate the friction budget, then make the lens recurring. Pick the one or two moments where added friction is the deposit: A real consent ask where data goes ambient, a visible “don’t use this part” control, and strip disclaimer-theater everywhere else, because blanket friction depletes patience without depositing anything. Then add a ledger column to your crit template and re-run the walk whenever the model or the data pipeline changes. Trade-off called out: a standing lens means some reviews get longer. The alternative is re-learning this article’s second section in the press.
Will your first ledger be mostly withdrawals and IOUs? Most likely. So was your first accessibility audit. The point is that it turns a vague queasiness about “AI trust” into a list of decisions with owners.
Tools & Resources
Everything referenced above, and what each is actually useful for:
- State of UX 2026 — Moran, Budiu & Gibbons, NN/g — the trend report naming trust as 2026’s design problem; useful for making the “why now” case to your exec peers.
- People + AI Guidebook — Google PAIR — especially the Explainability + Trust and Data Collection + Evaluation chapters; the closest thing to a canonical playbook for calibration and consent patterns.
- Sentient Design — Josh Clark & Veronika Kindred (Rosenfeld Media, 2026) — the AI-as-design-material argument in full, including two chapters on defensive design for AI failure; also the free foundational article.
- Systems Smart Enough to Know When They’re Not Smart Enough — Josh Clark — the 2017 essay on machine over-confidence; still the best short read on designing the display of uncertainty.
- Guidelines for Human-AI Interaction — Microsoft HAX Toolkit — 18 validated guidelines (Amershi et al., CHI 2019); G1 and G2 are your zero-friction expectation-setting deposits.
- Trust in Automation: Designing for Appropriate Reliance — Lee & See, 2004 — the calibration-not-maximization argument, twenty years before it was about LLMs.
- Trust, Attitudes and Use of AI: A Global Study 2025 — KPMG & University of Melbourne — 48,000-person baseline numbers for your next leadership deck.
- Americans and AI 2026 — Pew Research Center — the 71%-less-secure stat; the single most quotable figure on the data-trust deficit.
- The Transparency Dilemma — Schilke & Reimann, 2025 — thirteen experiments on why disclosure costs trust but exposure costs more; ammunition for the “pay it anyway” argument.
- AI Magic-8-Ball Thinking — Caleb Sponheim, NN/g and AI Hallucinations: What Designers Need to Know — Page Laubheimer, NN/g — the overtrust failure mode and what to do about it in the interface.
- Prioritize Smarts over Sentience — Evan Sunwall, NN/g — evidence that performed personality can reduce trust in task-oriented AI.
- The consent-controversy case file: Zoom’s ToS reversal (TechCrunch), Slack’s default opt-in backlash (TechCrunch), LinkedIn’s backtrack (ITPro), and the Otter.ai class action (National Law Review) — four precedents for the “silent withdrawal” pattern, useful when someone claims the risk is theoretical.
- Granola — the bot-less AI notepad from the friction example; its security & privacy FAQ and participant-privacy playbook are worth reading as a model of a vendor documenting the consent hand-off honestly instead of burying it.
- Thales 2025 Digital Trust Index and Cisco 2024 Consumer Privacy Survey — the abandonment numbers that turn trust from a values conversation into a revenue one.
Every question you ship is a withdrawal from somewhere. Fund the account first.