Trust Calibration: The Best Users Are Half-Skeptical
Full trust pastes fabricated case law into court filings; no trust turns AI into a paperweight. Judge trust yourself across six scenarios; clickable citations, proxy confidence signals, and conditional warnings
THE QUESTION THIS PAGE ANSWERS
ANSWER FIRSTWhat is the key idea behind “Trust Calibration: The Best Users Are Half-Skeptical”?
Full trust pastes fabricated case law into court filings; no trust turns AI into a paperweight. Judge trust yourself across six scenarios; clickable citations, proxy confidence signals, and conditional warnings
Turn taste into a behavior the product can repeat. The useful outcome is not a nice opinion. It is a visible rule, a small example, and a way to tell when the experience falls below the bar.
Capture one before-and-after example that shows the quality bar without extra explanation.
Polish that improves the surface while leaving the user's uncertainty untouched.
⚠ Overtrust: treating hallucination as truth
2023, New York—Mata v. Avianca: a practicing lawyer used ChatGPT to find case law and pasted 6 fabricated cases that don’t exist straight into court filings. After the judge checked each one, the lawyer was sanctioned and made global news. Similar accidents kept coming: AI output is fluent, confident, and well formatted—every surface cue nudges the judgment “this looks solid”, and those cues have nothing to do with whether the content is true.
⚠ Undertrust: AI becomes an expensive paperweight
The crash at the other end is quieter: a company buys AI tools, an employee hits an error once, and from then on every line of output gets sentence-by-sentence review. Review costs more than writing it yourself, so people stop using it. Procurement keeps paying; efficiency never rises. Undertrust doesn’t make the news—it only shows up in the internal postmortem titled “AI tool active usage: 8%.”
Here’s the judgment mantra first: watch “risk × verifiability”, not “how powerful AI is.” The same full adoption is calibrated on a weekly report and overtrust in court. Six scenarios—you’re the calibrator.
Calibration tool one: source citations—the kind you can open. Users who want to verify jump to the original in one click, which reins in overtrust; users too lazy to verify still see “there’s a source” and get reasonable confidence, which patches undertrust. The win-win requires citations that are real and clickable: a decorative fake citation exposed once is ten times worse than no citation. The AI answer below hangs three superscripts—open each one and compare with the source text.
Tool two: confidence. Part One already covered it—the model doesn’t know what it doesn’t know, so self-reported certainty is unreliable. But the product layer has honest proxy confidence signals: how many docs retrieval hit, how relevant they are, whether sources agree, whether the knowledge cutoff covers the question. Below is the same medication answer; three switch groups map to three product decisions—flip them and watch how the answer on the right and the user’s trust calibration change.
Tool three does one job: that tiny line—“AI may err; please verify important information”—is anyone actually reading it? Psychology’s answer is banner blindness: Benway & Lane (1998) found with eye-tracking that elements constantly shown in a fixed spot get filtered out by the brain as background texture. Tap through five answers below and watch that line disappear with your own eyes.
Turn the feeling in “Two ways to crash first” into a judgment
“2023, New York—Mata v.” points out that AI has lowered the bar for making something usable. The skill readers need is noticing what is wrong and turning that feeling into an actionable requirement.
Watch the user's next action, not just the surface
Turn “The crash at the other end is quieter: a company buys AI tools, an employee hits an error once, and from then on every line of output gets sentence-by-sentence review .” into observable questions: does the user know what happened, what to do next, and how to recover from an empty or failed state? Does the hierarchy make the important information visible first?
- Set the trust target at calibration : full trust causes accidents, no trust wastes money—pull users toward the diagonal
- Make citations clickable : swap “trust me” for “you can verify,” and watch for decorative fake citations biting back
- Say confidence with proxy signals : retrieval hit count, relevance, source agreement—more honest than the model’s self-reported certainty
Pretty is not the same as usable
Apply “Tool three does one job: that tiny line—“AI may err;” to a second screen or flow. Record one moment of hesitation and the user action after the change; observable behavior is stronger evidence than polish alone.
From “Two ways to crash first” to “Six users, pulse-check one by one”
“Two ways to crash first” grounds the problem in “2023, New York—Mata v. Avianca: a practicing lawyer used ChatGPT to find case law and pasted 6 fabricated cases that don’t exist straight into court filings. After the judge checked each one, the lawyer was san…”. “Six users, pulse-check one by one” then moves it toward “Here’s the judgment mantra first: watch “risk × verifiability” , not “how powerful AI is.” The same full adoption is calibrated on a weekly report and overtrust in court. Six scenarios—you’re the calibrator”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.
Carry the judgment into the next situation
For experience work, turn abstract impressions into user actions: did the person understand the state, find the next step, recover from an error, and want to continue?
- “Two ways to crash first”: 2023, New York—Mata v. Avianca: a practicing lawyer used ChatGPT to find case law and pasted 6 fabricated cases that don’t exist straight into court filings. After the judge checked each one, the lawyer was san…
- “Six users, pulse-check one by one”: Here’s the judgment mantra first: watch “risk × verifiability” , not “how powerful AI is.” The same full adoption is calibrated on a weekly report and overtrust in court. Six scenarios—you’re the calibrator
- “The closing point”: Show warnings conditionally : a constant disclaimer goes invisible in three days; a low-confidence popup with a reason is what gets read
The final “The closing point” brings the discussion to “Show warnings conditionally : a constant disclaimer goes invisible in three days; a low-confidence popup with a reason is what gets read”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.
✅ What this lesson wants to share
- Set the trust target at calibration: full trust causes accidents, no trust wastes money—pull users toward the diagonal
- Make citations clickable: swap “trust me” for “you can verify,” and watch for decorative fake citations biting back
- Say confidence with proxy signals: retrieval hit count, relevance, source agreement—more honest than the model’s self-reported certainty
- Show warnings conditionally: a constant disclaimer goes invisible in three days; a low-confidence popup with a reason is what gets read
I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.
After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.
When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.
No discussion on this article yet.