Lesson 7 — B3 Trust Calibration: How Much to Trust, By Job Type
Standard: — · Bloom's: — · Structure: PBL (his real week of tasks — quiz-writing, news-claims, committee business — IS the problem).
Notice the Say-See-Do cycles running on Frank O.'s actual work — not a canned exercise — and that every capability claim is cited to the frozen doc-set. The exit ticket climbs Bloom's to —, and the lesson closes by writing to the ledger.
Learning objective
Bloom's level: Evaluate. Defend, in writing, your own standing set of trust rules — what you'll accept from Claude without a second look, and what you never will — with a reason for each one, tested against a real week of your own tasks.
Plain language: Write down your own list of "sign off without a second look" jobs and "always gets inspected" jobs where Claude's involved, and be ready to say why each one lands where it does. Not a gut feeling — a rule you could defend to another committee member if they asked.
Standard: AIHC.1.B3 — "State what they currently trust an AI coworker to do without review and what they never accept unreviewed, with a reason for each." Plain words: knowing how much to trust, and being able to say why.
Your real artifacts today: the three-lane trust sort you already built earlier this unit, plus a real week's worth of tasks — quiz questions, the news clipping you're still working through, and a claim you heard at the union hall.
SSD cycle 1 — confidence is not a green light
- SAY: Anthropic's own documentation is direct that Claude "can occasionally generate responses that are inaccurate or deceptive" — the article names the phenomenon hallucinating (doc-set §6, source S5). The load-bearing sentence, in Anthropic's own words: "users should not rely on Claude as a singular source of truth and should carefully scrutinize any high-stakes advice given by Claude" (doc-set §6, S5). Here's the real term for what that leaves you holding: trust calibration — matching how much you check to what's actually at stake, not to how sure the answer sounded. Trade version: a helper who guesses the wire gauge from memory instead of pulling the table sounds exactly as confident either way — sounding sure was never the tell.
- SEE: A described two-column card. Left: your ORIGINAL three-lane sort from earlier in the unit — green / yellow / red, each with one or two example tasks dropped in, no reasons written. Right: a blank "trust rule sheet" template with the same three lanes plus a new fourth column, REASON, empty and waiting.
- DO: Pick three tasks already sitting in your three-lane sort. For each, write the missing reason in one sentence — for example, "brainstorming wording for the toolbox talk = green, because a bad word choice costs me thirty seconds to notice and fix myself."
SSD cycle 2 — the never-without-review line
- SAY: Anthropic's own Usage Policy names legal, financial, and employment-related uses as high-risk categories — when the output is consumer-facing, it requires extra safeguards: a qualified human reviewing it, and disclosure that AI helped produce it (doc-set §9, source S11). Translate that to your world: apprentice-training material feeds directly into someone's real certification path — that makes it employment-adjacent by the policy's own logic, not by this lesson's guess. Real term: human-in-the-loop requirement ↔ trade version: the jobs that need a permit pulled and a licensed inspector's signature, no matter how clean the work looks.
- SEE: Your real task list from this week — quiz questions, the news clipping, a committee memo — with a heavy border drawn around anything that's certification- or employment-adjacent, each one labeled with the S11 category it falls under.
- DO: Go through your actual week's tasks and mark which ones sit inside that heavy border — the permit-and-inspection lane, never signed off unreviewed — and write the S11-grounded reason next to each.
SSD cycle 3 — recalibrating the lanes after real frames
- SAY: A bowler's league handicap gets recalculated after real frames are on the board, not guessed at the start of the season. Your original three lanes, a few lessons back, were a first guess. You've had real turns since — some checks caught something, some didn't. That's evidence, and evidence is what moves a lane. (Referent: your Tuesday-night league — flavor for the pacing point, not the grounding; the grounding stays Anthropic's own S5/S11 language above.)
- SEE: A simple recalibration table: Original lane guess → Did checking catch anything? (yes/no) → Revised lane, filled with one worked row from a task type outside your own (so yours stays fresh): "wording suggestions for a flyer — original: green — caught anything? no, three checks running — revised: still green."
- DO: Run that same three-column check on five to eight of your own real tasks from this unit so far. Move any lane the evidence actually moves — don't move one just because it feels due for a change.
SSD cycle 4 — writing the standing rule sheet
- SAY: A trust rule sheet is a living document, not a one-time chart. The reason column from cycle 1 is what makes it defensible later — to yourself, or to a training-committee member who asks why you'd let Claude touch one thing unreviewed and not another.
- SEE: A clean finished template: Rule | Reason | Review required? (Y/N) | Standard tie (AIHC.1.B3) — with your cycle 1–3 work already dropped into the first three rows as a starting point.
- DO: Finish the sheet — at least five standing rules with reasons, covering both your "evaluating an AI claim" tasks and your apprentice-material tasks.
Independent at-bat (unscaffolded — no chart provided, just your own sheet)
A new claim comes up at the hall this week: "AI is going to replace the need for licensed electricians within five years." Using only your finished rule sheet from cycle 4 — no new chart, no fresh three-lane exercise — lane this claim, name which rule you applied, and write the reason. If none of your existing rules quite fit, write the new rule it forces you to add.
Exit ticket (Bloom's climb: Remember → Understand → Apply → Analyze → Evaluate, objective level)
- (Remember) In your own rule sheet's terms, what's the difference between a "green" task and a "red" task?
- (Understand) Why doesn't a confident-sounding Claude answer tell you which lane a brand-new task belongs in? One sentence.
- (Apply) Claude drafts a one-paragraph description of a new safety regulation for the apprentice newsletter. Which lane, and why?
- (Analyze) Look at your cycle 3 recalibration. Which lane moved the most from your original guess, and what evidence moved it?
- (Evaluate — objective level) Defend your rule sheet: name your two most defensible rules, and the one rule you're least sure of — then say what evidence would change your mind on that one.
Graded against your own rule sheet, the S11 categories, and your actual week's outcomes — never against how confident you felt when you wrote a rule.
Ledger write
ledger_append:
learner_id: L4-PUB-FRANK
lesson_id: L4-lesson07-b3-trust-calibration-20260719
standard: AIHC.1.B3
exit_ticket: {score: "", bloom_reached: evaluate}
auto_mastery: "0.10 -> (projected ~0.55 — first fully dedicated B3 lesson; adjacent verification/delegation habits from lessons 2–6 likely transferred some judgment already)"
self_score: ""
calibration_gap: " — not computed until a real self-score exists"
journal_prompt: >
If a guy at the hall asked you tomorrow "how do you know when to trust that thing,"
what's the one-sentence answer your rule sheet gives you? Write it the way you'd
answer him standing at the bar, not the way you'd write it for a form.
structure_used: PBL
referents_used: [bowling league — Tuesday nights (recalibration pacing point only)]
next_lesson_seed: "C1 Provenance — a different real committee document; his B3 rule sheet's review-required flag becomes the trigger for when labeling matters most."
Rubric self-audit (R1–R13)
| # | Indicator | Verdict | Evidence |
|---|---|---|---|
| R1 | One Bloom's-leveled objective, named standard, plain-language too | PASS | Evaluate-level objective stated technically and plainly; AIHC.1.B3 quoted verbatim with plain gloss |
| R2 | Every capability/limit claim traces to the doc-set | PASS | S5 (hallucination, scrutinize high-stakes advice — both quoted verbatim), S11 (high-risk categories, human review + disclosure) — every product claim cited by section/source |
| R3 | 3–6 SSD cycles complete | PASS | 4 cycles, each with SAY/SEE/DO |
| R4 | Each SEE ground-truth-verified, no strawman errors | PASS | Cycle 1 SEE builds directly on his own prior three-lane sort (no invented content); cycle 2 SEE maps his real tasks against S11's own categories, not a fabricated risk list |
| R5 | Every DO on real artifacts, free-tier only | PASS | All four cycles + at-bat act on his existing rule sheet, his real week's tasks, and a real hall claim; no paid-tier feature invoked |
| R6 | Media doctrine — static concepts get static visuals | PASS | All four SEEs are described static cards/tables; nothing here is motion/process content |
| R7 | Exit ticket 3–5 Qs climbing to objective level | PASS | 5 questions, Remember→Evaluate, graded against his own sheet and S11, not against confidence |
| R8 | Ledger write: standard + both scores + journal preserved | PASS | Standard named, auto/self marked per spec, journal prompt in his own register |
| R9 | Scaffold with visible, continuing fade | PASS | Only cycle 1 gives a worked example; cycles 2–4 hand him the move directly; the at-bat provides no chart at all, unlike earlier-unit at-bats |
| R10 | Referents elected-only, flavor/bridge, anti-stereotype clean | PASS | Only the bowling referent used, explicitly scoped to "pacing point only" — never substituted for S5/S11 grounding |
| R11 | Plain-language-first: every technical term gets analogy + real term | PASS | "hallucinating" paired with the wire-gauge-guess analogy; "trust calibration" paired with "matching how much you check to what's at stake"; "human-in-the-loop requirement" paired with the permit-and-inspection analogy |
| R12 | Non-replication: profile-specific | PASS | Every DO runs on Frank's actual rule sheet, actual week, and an actual hall claim — a different learner's profile has none of these artifacts to draw on |
| R13 | No deficit-framing of learner, no fear-framing of tool | PASS | Frank's prior three-lane work is treated as real, usable progress to build on, not a gap; hallucination is stated as a named, documented mechanism with a checking habit attached, not an alarm |
Escalation: all load-bearing indicators pass → auto-ship (per playbook §5 policy).