Lesson 07 — Trust Calibration: when a Claude answer earns your trust as-is
Standard: — · Bloom's: — · Structure: —.
Notice the Say-See-Do cycles running on Priya S.'s actual work — not a canned exercise — and that every capability claim is cited to the frozen doc-set. The exit ticket climbs Bloom's to —, and the lesson closes by writing to the ledger.
Learning objective
(Bloom's: Evaluate) Before you act on a Claude-drafted passage or answer in The Last Mill or your interview prep, evaluate whether the current model, its knowledge cutoff, and its active settings justify trusting the output as-is — or whether the task calls for web search, extended thinking, or your own verification instead.
Standard named: AIHC.1.B3 (Trust calibration). Trust, per the standards backbone, is evidence-based, not a feeling (X1), and where the trustworthy line sits moves as models and their settings change (X2) — this lesson teaches you to locate that line for a specific claim, not to feel generally confident or generally suspicious.
Cycle 1 — Trust is evidence, not a feeling
SAY — A confident-sounding Claude answer and a well-evidenced one are not the same thing. Before you decide whether to trust a passage as-is, two facts settle it, not the tone of the prose: (1) which model answered, and (2) whether that model's knowledge reaches the claim's date. Per the doc-set, the model menu sits "next to the send button and governs three things together — which model, what effort level, and whether extended thinking is on" [S14], and every current model carries a stated cutoff — for example "Sonnet 5, Fable 5, Opus 4.8, Opus 4.7 — January 2026" while "Sonnet 4.6, Opus 4.6 — August 2025" and "Haiku 4.5 — July 2025" [S15]. Verbatim: "These models may not be aware of events or information that occurred after their respective cutoff dates" [S15].
SEE — A described two-panel graphic: left panel, the model-name control next to the send button (per S14's stated location) with an arrow labeled "1. which model answered you." Right panel, the S15 cutoff table, with an arrow labeled "2. does the claim's date fall before or after that line?" A third small callout notes that S14's effort-tier list names "Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6" specifically — if your active model isn't on that list, don't assume the same tier controls apply; check your own menu rather than extrapolate.
DO — Open your real "The Last Mill" chat. Click the model name (S14) and write down which model you're on. Look it up in the S15 table and write down its cutoff date. Now pick one factual claim currently in your draft and note whether that claim's date falls before or after the cutoff you just found.
⏸ Pause point. You now have: your model, its cutoff, and one claim's date relative to it. Stop here if you need to — nothing further depends on doing this in one sitting.
Cycle 2 — Sort the question before you trust the answer
SAY — Not every open question needs the same tool. The doc-set draws the line explicitly: "Web search — best for quick factual lookups needing one or two tool calls... Extended thinking — best for hard reasoning that doesn't need fresh web info... Research — best for heavier information-gathering: five-plus tool calls over roughly 1–3 minutes" [S10]. Web search is switched on via "the slider icon in the chat input, find 'Web search' in the dropdown, toggle on" [S9], and once it's on, an answer comes with "direct citations plus links back to the sources" [S9] — which is itself evidence you can check, not just a claim to accept.
SEE — A described three-lane decision flow: Lane 1 "quick fact/current event → toggle Web search [S9]"; Lane 2 "hard reasoning, no new facts needed → toggle Extended thinking [S14]"; Lane 3 "heavy multi-source synthesis → Research (requires web search on first) [S10, S11]."
DO — Take two real open questions from your interview-prep notes or draft right now. Sort each into one of the three lanes, and turn on the matching toggle for at least one of them in your real chat.
⏸ Pause point. Two questions sorted, one toggle live. That's a complete, resumable unit — pick this back up whenever your block allows.
Cycle 3 — The line moves; recheck it, don't assume it
SAY — Trust calibration isn't a setting you pick once. X2, the moving line: as a new model or a new cutoff ships, where "trust as-is" ends and "verify first" begins shifts with it. The doc-set says this about itself: its own cutoff table and model names are "exactly the kind of fact that goes stale" and should be rechecked "on a rotating cadence" (doc-set frontmatter). Referent (IPL cricket, flavor only): the pitch changes character between overs — a shot that worked two overs ago can fail on the next ball if the wicket's turned. You re-read the conditions each time; you don't trust last over's pitch report forever.
SEE — A described annotated excerpt of the doc-set's own frontmatter warning (the exact sentence about model names and cutoffs going stale), with a callout: "this warning is about the doc-set, but the same discipline applies to your own trust calibration."
DO — In your real Project (wherever you keep working notes for this piece), write one standing line for yourself: today's date, your current model, and its cutoff — a marker you'll compare against the next time you sit down to write.
⏸ Pause point. One line written. Nothing else required before stepping away.
Cycle 4 — Ask the model what it's unsure about
SAY — Trust calibration isn't only you checking the model — it's also asking the model to tell you where it's shaky. When web search is on, an answer citing sources shows its work [S9]; when it's off, that same claim rests on the model's own training, unlabeled. Directly asking "did you use web search for that, and how recent is your information?" turns a silent assumption into a checkable answer.
SEE — A described transcript snippet: Priya's real question, followed by a described Claude reply distinguishing a search-sourced claim (with a citation link, per S9) from a claim answered from training knowledge alone (no citation, no search indicator).
DO — In your real current thread, ask Claude directly whether it used web search for the claim you flagged in Cycle 1, and how current its information on that topic is. Record the answer next to the claim.
Independent at-bat
Unscaffolded: take the next three factual claims in your real draft or interview-prep list. For each, decide and write down a trust verdict — trust as-is, needs web search, or needs a different setting (extended thinking / research) — and justify it using what you now know about your model's cutoff and the three-lane sort from Cycle 2. No checklist provided this time.
Exit ticket (climbing to Evaluate; graded against the doc-set, not model agreement)
- (Remember) Besides which model, name the two other things the model menu controls together, per S14.
- (Understand) Why can't you assume Sonnet 5 (or any current model) knows about something that happened after its stated cutoff, even if it answers confidently? [S15]
- (Apply) Your draft states the mill's original 1987 closure date. Which lane from Cycle 2 does this claim belong in, and why?
- (Analyze) Compare that 1987 claim to a claim about last month's reunion event mentioned in the same chat. Why do they land in different trust postures even though a single model answered both?
- (Evaluate — objective level) Here is an unverified paragraph from your real draft (three claims). Render a trust verdict on each claim — trust as-is / needs search / needs a different setting — and justify every verdict against your model's cutoff and the S9/S10 tool-choice logic, not against how confident the sentence sounds.
Ledger write
ledger_write:
learner_id: L3-CONS-PRIYA
lesson_id: L3-07-B3
standards: [AIHC.1.B3]
tags: [X1, X2]
exit_ticket:
score:
bloom_reached:
auto_score:
self_score:
calibration_gap:
journal_prompt: >
Pick one claim you trusted as-is today without checking a cutoff or a citation. Now that
you've run this lesson, would you still trust it the same way — and what specifically
changed your mind (or didn't)?
structure_used: PBL
referents_used: [cricket-IPL]
next_lesson_seed: "carry the model-menu + cutoff habit forward into Lesson 12 (D2 Regeneration), where a new model or feature moving the line is the trigger, not the topic."
Rubric self-audit (R1–R13, against Exemplar_and_Rigor_Rubric_WebChat_x_SmallBusiness_2026-07-19_v01_I.md Part B)
| # | Indicator | Verdict | Evidence |
|---|---|---|---|
| R1 | One Bloom's-leveled objective, ≥1 named AIHC standard, learner-visible | PASS | Objective block names Evaluate + AIHC.1.B3 explicitly |
| R2 | Every product claim traces to the frozen doc-set; no invented UI | PASS | All feature claims cite S9/S10/S14/S15; the effort-tier model list is quoted exactly, including the honest note that some models (e.g. plain Sonnet 5) aren't on S14's stated list — not extrapolated |
| R3 | 3–6 SSD cycles, one point each, complete | PASS | 4 cycles, each one distinct point (model+cutoff, tool-lane sort, moving line, ask-the-model) |
| R4 | Each SEE anchors its SAY, ground-truth-verified | PASS | SEEs describe only UI/table elements stated in S9/S10/S14/S15 and the doc-set's own frontmatter, no fabricated screens |
| R5 | Every DO acts on the learner's real work | PASS | DOs operate on Priya's real chat, real draft claims, real interview-prep questions |
| R6 | Media doctrine: static vs. genuine motion | PASS | No motion involved (menus, tables, transcripts); all SEEs are static/annotated |
| R7 | Exit ticket 3–5 Qs, Bloom's-climbing, doing-focused, SSOT-graded | PASS | 5 Qs, Remember→Evaluate, graded against S9/S10/S14/S15 facts |
| R8 | Ledger write: standards, both scores, journal preserved | PASS | ledger_write block present with anchor codes, scores, journal_prompt |
| R9 | Band-appropriate scaffolding with fade; independent at-bat present | PASS | Cycles 1–2 fully scaffolded, Cycle 4 lighter, at-bat unscaffolded (3 claims, no checklist) |
| R10 | Referents elected-only, flavor-only, anti-stereotype clean | PASS | Cricket-IPL used once, as pacing analogy only, never substitutes for a doc-set claim |
| R11 | Timing-tolerance honored: pause points, no unbroken long block | PASS | Pause points after Cycles 1, 2, 3 with explicit "stop here if you need to" |
| R12 | Non-replication: profile swap changes DOs + exit ticket materially | PASS | Every DO/question keys to Priya's specific model, draft claims, and cricket referent — a different learner profile changes all of it |
| R13 | Consumer trust boundaries taught alongside features, never over-reassured | PASS | Cycle 1 explicitly flags the S14 effort-tier list's model-coverage gap rather than smoothing over it |
Escalation verdict: all load-bearing indicators PASS → clears; auto-ships per the versioned policy (Playbook §5).