Lesson 6 — B2 Failure Literacy: Naming the Three Ways Claude's Answers Go Wrong
Standard: — · Bloom's: — · Structure: PBL (his real, long load-calculation drafting session is the problem).
Notice the Say-See-Do cycles running on Frank O.'s actual work — not a canned exercise — and that every capability claim is cited to the frozen doc-set. The exit ticket climbs Bloom's to —, and the lesson closes by writing to the ledger.
Learning objective
Bloom's level: Analyze. Recognize, in your own real long quiz-drafting session, which of three documented failure modes is present at a given moment — and name it specifically, not just "it got that wrong."
Plain language: Knowing the characteristic ways this thing's answers go wrong, by name, so you can say exactly what happened instead of just "it messed up."
Standard: AIHC.1.B2 — "Recognize common failure modes in the wild — fabrication, ungrounded confidence, context loss — and name what happened specifically rather than 'it got it wrong.'" Plain words: knowing the characteristic ways AI answers go wrong.
Your real artifact today: the full, long back-and-forth session where you drafted a run of load-calculation quiz questions (NEC Article 220) for the dwelling-service sheet — several questions deep, all in one running conversation.
Lesson 1 previewed three bare labels without explaining them. Today those get names, sources, and real instances.
SSD cycle 1 — fabrication: a confident citation that isn't real
- SAY: Anthropic's own support article states plainly that Claude "can occasionally generate responses that are inaccurate or deceptive" — a phenomenon it names hallucinating — and specifically that it "may produce quotations that look authoritative or sound convincing, but are not grounded in fact" (doc-set §6, source S5). In trade terms: an apprentice reciting an article number with total confidence, in a tone like he's reading straight off the page — and it turns out that number was never in the book at all. The tell isn't tone. It's whether the citation actually exists where claimed.
- SEE: A moment from your real load-calc session: Claude cites a specific table reference to support a load-calculation step, stated with full confidence — and on inspection, that table doesn't cover the specific dwelling-load scenario claimed. Not a wrong number from a real table; a citation that doesn't hold up at all. That specific pattern — confident, cites something, the something doesn't check out — is fabrication, by name.
- DO: Open your real load-calc session. Find one citation-heavy moment and check it against your actual code book. Decide: is this a real instance of fabrication specifically (confident, cites something, doesn't hold up) — or is it just an ordinary wrong number from a real table (a different failure mode entirely, not this one)?
SSD cycle 2 — a training-data gap: it can't know what wasn't in what it learned from
- SAY: The same article names a second, different mechanism: Claude "can get confused when prompted about current events" if the training data on that subject was incomplete (doc-set §6, S5). This isn't fabrication — it's a book printed before the newest revision existed. Even the doc-set powering this very lesson says so about itself: it carries its own freshness note, flagging that pricing, features, and model names move fast enough to need rechecking. If a code book can't know about an amendment printed after it went to press, Claude can't know about a change that happened after its training data was gathered — that's not a flaw hidden from you, it's a documented mechanism with a name.
- SEE: A moment from your real session where you asked about a very recent local amendment to the load-calculation article — one adopted after most training data would have been gathered. The answer treats the older rule as settled, with no flag that something newer might exist. That specific gap — recent, uncaveated, presented as settled — is the training-data-gap pattern, not fabrication and not context loss.
- DO: Find one code-recency-sensitive question in your real materials (yours, or the clipping's claim about where AI-in-the-trades stands right now). Check whether the answer shows any sign of awareness that something newer might exist. Name, specifically, whether this is the training-data-gap pattern.
SSD cycle 3 — context loss: the job box only holds so much paperwork
- SAY: Anthropic's own documentation describes exactly this mechanism for long conversations: "Claude automatically manages long conversations. When your conversation approaches the context window limit, Claude summarizes earlier messages" (doc-set §5, source S2) — and "a portion of the context window is always reserved for Claude's response" (same source). Trade parallel: the job box on site only holds so much paperwork. Once it's full, older pages get condensed to make room for new ones — and a detail from page one can go missing by page nine, not because anyone lied, but because the box has a limit.
- SEE: Your real session, early on: you set the running example as "an 1,800 sq ft dwelling with a 10kW range." By question 8 or 9, deep in the same long conversation, an answer uses a different square footage — quietly inconsistent with what you set at the start. Nobody re-asked the question wrong; the constraint from page one got lost in a conversation long enough to summarize.
- DO: Scan the later half of your real session for a place where an earlier constraint — your format, your numbers, your classroom style — got dropped or contradicted. Name it specifically as context loss, not "it got confused."
Independent at-bat (fully unscaffolded — the loop from lesson 1 closes here)
Take your real long load-calc session in full. Go through it moment by moment and label each place you'd flag: fabrication, training-data gap, context loss, or none of the three (a plain wrong number from a real, current, in-window source is still just a wrong number — not every mistake needs one of these three names). No worked example is given this time. This is the same sorting move lesson 1 previewed with three bare labels — now you're the one applying real names to real material.
Exit ticket (Bloom's climb: Remember → Understand → Apply → Analyze → Evaluate, capstone question)
- (Remember) Name the three failure modes covered today, and the doc-set source behind each.
- (Understand) Why is "it got confused" not a specific enough answer for any of the three? One sentence per mode.
- (Apply) Here's a new moment from your session, not yet labeled: pick one and name which of the three it is (or state clearly that it's none of them).
- (Analyze) Which of the three modes showed up most in your real session — and does that match, or differ from, the mode you guessed back in lesson 1 before you knew any of their names?
- (Evaluate — objective level, integrating the unit) Take the fabricated-citation moment from cycle 1. Walk the full loop, one line per anchor: would you have delegated this task at all (A2)? What would you have asked Claude to flag about itself first (A3)? What did checking it actually show (B1)? And now — what specific failure mode do you name it (B2)?
Graded against your actual code book, your actual session transcript, and the three named mechanisms from the doc-set — never against a guess at what "probably" happened.
Ledger write
ledger_append:
learner_id: L4-PUB-FRANK
lesson_id: L4-lesson06-b2-failure-literacy-20260719
standard: AIHC.1.B2
exit_ticket: {score: "", bloom_reached: evaluate}
auto_mastery: "0.15 -> (projected ~0.50 — correctly distinguished fabrication from an ordinary wrong number in cycle 1; named all three modes with sourced reasoning; capstone question integrated all five unit anchors)"
self_score: ""
calibration_gap: " — not computed until a real self-score exists"
journal_entry_prompt: >
Close the unit out the way you'd close a training-committee meeting: which of the three
failure modes are you most likely to run into again with your own apprentice materials,
and what's the one habit from this whole unit — asking clearly, choosing what to hand
over, asking Claude's own read, checking the work, or naming what went wrong — you're
actually going to keep using next week?
structure_used: PBL
referents_used: [electrical-trade code-inspection culture — job-box / paperwork-limit framing, primary throughout]
next_lesson_seed: "Unit complete for the 5 named anchors (A1, A2, A3, B1, B2). Ledger now shows a full mastery profile across the anchor set; a follow-on unit could extend to AIHC.1.B3 (trust calibration, already baselined in his profile at 0.10) or AIHC.1.C1 (provenance), reading forward from this session's demonstrated levels rather than resetting."
Rubric self-audit (R1–R13)
| # | Indicator | Verdict | Evidence |
|---|---|---|---|
| R1 | One Bloom's-leveled objective, named standard, plain-language too | PASS | Analyze-level objective stated technically and plainly; AIHC.1.B2 quoted verbatim with plain gloss |
| R2 | Every capability/limit claim traces to the doc-set | PASS | S5 (hallucination/fabrication, training-data-gap language, both quoted), S2 (context-window auto-summarization, quoted) — each cited by section/source; doc-set's own freshness note used honestly as a self-referential example, not overclaimed |
| R3 | 3–6 SSD cycles complete | PASS | 3 cycles, SAY/SEE/DO complete each |
| R4 | Each SEE ground-truth-verified, no strawman errors | PASS | Each SEE illustrates the doc-set's own named mechanism specifically (a citation that doesn't hold up; an uncaveated recent-topic answer; a dropped constraint across a long session) — distinguishable from each other and from an ordinary wrong answer, per cycle 1's own explicit contrast |
| R5 | Every DO on real artifacts, free-tier only | PASS | All cycles and the at-bat work directly on Frank's real long load-calculation session and real code book; clipping reused once (cycle 2); no paid feature invoked |
| R6 | Media doctrine honored | PASS | All SEEs are static described moments/annotations from a transcript; no motion content needed |
| R7 | Exit ticket 3–5 Qs climbing to objective level | PASS | 5 questions, Remember→Evaluate, the capstone question grounded against his own transcript and the doc-set's three named mechanisms |
| R8 | Ledger write complete | PASS | Standard, exit ticket, auto/self , journal prompt in-register, unit-closing next-lesson seed all present |
| R9 | Scaffold fade complete | PASS | Only cycle 1 supplies a worked contrast (fabrication vs. an ordinary wrong number); cycles 2–3 and the at-bat supply no filled-in example at all — the lightest scaffold of the six lessons, and the at-bat explicitly closes the loop lesson 1 opened, showing the arc's full fade end-to-end |
| R10 | Referents elected, flavor-only | PASS | Only the electrical-trade referent used (job-box/paperwork-limit framing); never substitutes for the doc-set's three named mechanisms |
| R11 | Plain-language-first | PASS | "hallucinating"/fabrication," "training-data gap," and "context loss" each introduced with both a trade analogy (fake citation, out-of-date book, overstuffed job box) and the doc-set's own real term |
| R12 | Non-replication | PASS | Every DO is keyed to Frank's specific real load-calculation session (his own square-footage example, his own citation moments) — none of this exists for another learner's profile |
| R13 | No deficit-framing, no fear-framing | PASS | All three failure modes are presented as documented, named, expectable mechanisms with a specific tell each — never as reasons to distrust the tool broadly or as evidence Frank missed something he should have caught sooner; the capstone question treats his full five-anchor competence as already-built, not remedial |
Escalation: all load-bearing indicators pass → auto-ship.