Lesson 5 — B1 Verification: Checking Claude's Answer Like Any Subcontractor's Work
Standard: — · Bloom's: — · Structure: PBL (walking the exact flags lesson 4 surfaced, against his real code book and the clipping's real claim).
Notice the Say-See-Do cycles running on Frank O.'s actual work — not a canned exercise — and that every capability claim is cited to the frozen doc-set. The exit ticket climbs Bloom's to —, and the lesson closes by writing to the ledger.
Learning objective
Bloom's level: Evaluate. Check the clearance-dimension answer and the clipping's "40%" figure — the two things Claude itself flagged as least certain in lesson 4 — against independent sources, the way you'd check any subcontractor's work: against the code, not against how sure they sounded.
Plain language: Checking answers — actually opening the book, not just trusting how confident the reply felt.
Standard: AIHC.1.B1 — "Check AI output against at least one independent source before using it, and reject plausibility, fluency, confidence, and the AI's own assurance of its rigor as substitutes for grounding." Plain words: checking the work.
Your real artifact today: the working-clearance answer and the clipping's central figure — both flagged by Claude itself in lesson 4 — plus your box-fill answer key and your actual NEC code book.
SSD cycle 1 — confidence is finish work; correctness is what's behind the wall
- SAY: Anthropic's own documentation names this directly: Claude "can occasionally generate responses that are inaccurate or deceptive," a phenomenon it calls hallucinating, and its own stated guidance is blunt: "users should not rely on Claude as a singular source of truth and should carefully scrutinize any high-stakes advice given by Claude" (doc-set §6, source S5). A smoothly worded answer is like a nicely mudded, painted wall — it can look finished either way, whether or not the wiring behind it is correct. Confidence is what you see; correctness is what's behind the wall, and the only way to know is to open it up.
- SEE: The flagged clearance-dimension answer from lesson 4, shown smooth and confident — and next to it, the same answer after Frank actually opens Table 110.26(A)(1) in his code book: the answer assumed a "Condition 2" installation, but Frank's real panel setup is "Condition 3" (unguarded, grounded surfaces on the opposite side) — a different clearance number entirely. The gap is real, and it's visible only because he checked, not because the sentence sounded off.
- DO: Open your actual code book to Table 110.26(A)(1). Check the flagged clearance dimension against the real condition of your real panel. Grade it yourself: right, wrong, or right-for-a-different-condition-than-yours.
SSD cycle 2 — checking claim by claim, not the whole answer at once
- SAY: You don't total a bowling league night after frame one, and you don't clear a whole quiz answer as "correct" in one look either — you check it claim by claim, verdict by verdict, then move to the next. Your box-fill answer from lesson 2 has several separate checkable pieces: the box-size assumption, the cable count, the device count, and the final cubic-inch total. Each one gets its own look.
- SEE: A checklist rendering of that answer, broken into four separate claims, each with an empty verdict box: ✅ matches the table · ❌ doesn't match · ⚠ matches a different input than the one you actually gave.
- DO: Verify all four claims now, one at a time, against Article 314.16's actual table values. Mark each box yourself.
SSD cycle 3 — a citation from Claude is a pointer to check, not a fact by itself
- SAY: On the clipping's flagged "40%" figure and its "a study found" citation, Anthropic's own guidance is specific: when Claude cites something, "original websites may contain important context or details not included in Claude's synthesis" (doc-set §6, S5) — meaning the honest move is to open the actual source page yourself, not stop at Claude's summary of it. Free-tier claude.ai does include the "ability to search the web" (doc-set §4, source S15), so you can ask it to help you locate the underlying source — the checking step is still yours, afterward.
- SEE: A described exchange: Frank asks Claude to help find the source behind the clipping's claim; Claude returns a synthesis naming a source. The SEE shows Frank clicking through to that actual source page directly — and finding the real source states a narrower, more hedged claim (a possible range, under specific conditions, over a longer time horizon) than the clipping's flattened "up to 40% within a decade" paraphrase suggested. Nothing here required Claude to be wrong — the gap was already sitting between the study and the clipping; opening the original is what surfaced it.
- DO: Do exactly this yourself: ask Claude to help you find the source behind the clipping's claim, then open that actual source page directly, and write one sentence on what you found that the clipping's own wording left out or flattened.
SSD cycle 4 — reject the four things that aren't grounding
- SAY: The standard names exactly what does NOT count as checking: plausibility, fluency, confidence, and the AI's own assurance of its rigor. A cover band can nail the guitar tone and still flub the lyrics — sounding right isn't the same as being right, on a record or on a code citation. None of those four gets you off the hook for opening an independent source.
- SEE: A short reject-list card: Plausibility ❌ · Fluency ❌ · Confidence ❌ · The AI's own self-assurance ❌ — Independent source ✅, the only box that actually clears a claim.
- DO: Pick one more claim from either real artifact — your choice, box-fill sheet or clipping — that hasn't been checked yet this unit. Verify it against an independent source, and name, out loud, which of the four false substitutes you had to consciously set aside to bother checking at all.
Independent at-bat (full loop, unscaffolded)
Draft one brand-new apprentice quiz question with Claude's help — pick a topic not yet touched this unit, such as disconnecting means for HVAC equipment. Then inspect and correct it against your actual code book before it would ever reach an apprentice. No checklist given this time — just the four-move discipline from today, applied cold.
Exit ticket (Bloom's climb: Remember → Understand → Apply → Analyze → Evaluate, objective level)
- (Remember) Name the four things the standard says do NOT count as checking.
- (Understand) Why can a citation from Claude be a real, honest pointer and still not be a fact by itself? One sentence.
- (Apply) Here's a new claim from your clipping you haven't checked yet: pick one and run at least two of today's moves on it.
- (Analyze) Between the clearance-dimension answer and the "40%" figure, which was harder to check, and what made it harder — access to the source, ambiguity in the question, or something else?
- (Evaluate — objective level) Show your HVAC-disconnect quiz question and your inspection trail, and defend your final sign-off — or your rejection of the draft — the way you'd defend a real inspection call to another inspector.
Graded against your code book and the actual source pages you opened — never against how confident any answer sounded.
Ledger write
ledger_append:
learner_id: L4-PUB-FRANK
lesson_id: L4-lesson05-b1-verification-20260719
standard: AIHC.1.B1
exit_ticket: {score: "", bloom_reached: evaluate}
auto_mastery: "0.20 -> (projected ~0.55 — found a real discrepancy on the clearance-condition mismatch; completed full four-move discipline; independent at-bat closed the loop)"
self_score: ""
calibration_gap: " — not computed until a real self-score exists"
journal_entry_prompt: >
Log it the way you'd log a corrected inspection finding: what specifically did you catch
today that would have gone out wrong if you'd only judged it by how confident it sounded?
structure_used: PBL
referents_used: [electrical-trade code-inspection culture — wall/finish-work framing; bowling league — frame-by-frame pacing; classic rock — cover-band bridge, one line only]
next_lesson_seed: "B2 Failure literacy — name the SPECIFIC failure mode behind the clearance-condition mismatch and the clipping's flattened stat, not just 'it was wrong.'"
Rubric self-audit (R1–R13)
| # | Indicator | Verdict | Evidence |
|---|---|---|---|
| R1 | One Bloom's-leveled objective, named standard, plain-language too | PASS | Evaluate-level objective stated technically and plainly; AIHC.1.B1 quoted verbatim with plain gloss |
| R2 | Every capability/limit claim traces to the doc-set | PASS | S5 (hallucination naming, singular-source-of-truth caution, citation-as-pointer guidance), S15 (web search on free tier) — each cited by section/source |
| R3 | 3–6 SSD cycles complete | PASS | 4 cycles, SAY/SEE/DO complete each |
| R4 | Each SEE ground-truth-verified, no strawman errors | PASS | The condition-2-vs-condition-3 clearance mismatch is a realistic, specific, checkable code-table error type; the clipping/source gap is framed honestly as a flattening between study and clipping, not an invented Claude fabrication |
| R5 | Every DO on real artifacts, free-tier only | PASS | Every cycle checks a real, specific flagged item from lessons 2 and 4 against his actual code book or an actual opened source page; web search is a documented free-tier feature (S15) |
| R6 | Media doctrine honored | PASS | All SEEs are static described visuals (annotated answers, checklists, a reject-list card); no motion content needed |
| R7 | Exit ticket 3–5 Qs climbing to objective level | PASS | 5 questions, Remember→Evaluate, graded against his code book and the opened source pages |
| R8 | Ledger write complete | PASS | Standard, exit ticket, auto/self , journal prompt in-register present |
| R9 | Scaffold fade visible | PASS | Only cycle 1 gives a fully worked discrepancy; cycles 2–4 hand him the checklist shape only; the at-bat gives zero worked example — full unscaffolded loop, consistent with lesson 5's position in the fade arc |
| R10 | Referents elected, flavor-only | PASS | Electrical-trade referent central (wall/finish-work, inspection framing); bowling league used once (pacing); classic rock used once (cover-band line) — none substitute for doc-set grounding |
| R11 | Plain-language-first | PASS | "hallucinating" reused with its doc-set definition; "citation" explained as "a pointer to check, not a fact by itself" before being treated as a technical concept; "grounding" paired with "opening the book" throughout |
| R12 | Non-replication | PASS | Every DO checks a specific, previously-generated answer unique to Frank's own quiz sheets and his own clipping — nothing here is a generic verification drill |
| R13 | No deficit-framing, no fear-framing | PASS | The discrepancy found is framed as exactly what checking is for (a normal, expected catch), not evidence the tool or Frank failed; hallucination is stated as a documented, named phenomenon with a checking habit attached |
Escalation: all load-bearing indicators pass → auto-ship.