← Frank O.'s track Lesson 06 / 12 · Frank O.

Lesson 6 — B2 Failure Literacy: Naming the Three Ways Claude's Answers Go Wrong

Bloom's: PBL (his real, long load-calculation drafting session is the problem)
● Exhibit 6 of 12 — Frank O.'s track

Standard: — · Bloom's: — · Structure: PBL (his real, long load-calculation drafting session is the problem).

Notice the Say-See-Do cycles running on Frank O.'s actual work — not a canned exercise — and that every capability claim is cited to the frozen doc-set. The exit ticket climbs Bloom's to , and the lesson closes by writing to the ledger.


Learning objective

Bloom's level: Analyze. Recognize, in your own real long quiz-drafting session, which of three documented failure modes is present at a given moment — and name it specifically, not just "it got that wrong."

Plain language: Knowing the characteristic ways this thing's answers go wrong, by name, so you can say exactly what happened instead of just "it messed up."

Standard: AIHC.1.B2"Recognize common failure modes in the wild — fabrication, ungrounded confidence, context loss — and name what happened specifically rather than 'it got it wrong.'" Plain words: knowing the characteristic ways AI answers go wrong.

Your real artifact today: the full, long back-and-forth session where you drafted a run of load-calculation quiz questions (NEC Article 220) for the dwelling-service sheet — several questions deep, all in one running conversation.

Lesson 1 previewed three bare labels without explaining them. Today those get names, sources, and real instances.


SSD cycle 1 — fabrication: a confident citation that isn't real

  • SAY: Anthropic's own support article states plainly that Claude "can occasionally generate responses that are inaccurate or deceptive" — a phenomenon it names hallucinating — and specifically that it "may produce quotations that look authoritative or sound convincing, but are not grounded in fact" (doc-set §6, source S5). In trade terms: an apprentice reciting an article number with total confidence, in a tone like he's reading straight off the page — and it turns out that number was never in the book at all. The tell isn't tone. It's whether the citation actually exists where claimed.
  • SEE: A moment from your real load-calc session: Claude cites a specific table reference to support a load-calculation step, stated with full confidence — and on inspection, that table doesn't cover the specific dwelling-load scenario claimed. Not a wrong number from a real table; a citation that doesn't hold up at all. That specific pattern — confident, cites something, the something doesn't check out — is fabrication, by name.
  • DO: Open your real load-calc session. Find one citation-heavy moment and check it against your actual code book. Decide: is this a real instance of fabrication specifically (confident, cites something, doesn't hold up) — or is it just an ordinary wrong number from a real table (a different failure mode entirely, not this one)?

SSD cycle 2 — a training-data gap: it can't know what wasn't in what it learned from

  • SAY: The same article names a second, different mechanism: Claude "can get confused when prompted about current events" if the training data on that subject was incomplete (doc-set §6, S5). This isn't fabrication — it's a book printed before the newest revision existed. Even the doc-set powering this very lesson says so about itself: it carries its own freshness note, flagging that pricing, features, and model names move fast enough to need rechecking. If a code book can't know about an amendment printed after it went to press, Claude can't know about a change that happened after its training data was gathered — that's not a flaw hidden from you, it's a documented mechanism with a name.
  • SEE: A moment from your real session where you asked about a very recent local amendment to the load-calculation article — one adopted after most training data would have been gathered. The answer treats the older rule as settled, with no flag that something newer might exist. That specific gap — recent, uncaveated, presented as settled — is the training-data-gap pattern, not fabrication and not context loss.
  • DO: Find one code-recency-sensitive question in your real materials (yours, or the clipping's claim about where AI-in-the-trades stands right now). Check whether the answer shows any sign of awareness that something newer might exist. Name, specifically, whether this is the training-data-gap pattern.

SSD cycle 3 — context loss: the job box only holds so much paperwork

  • SAY: Anthropic's own documentation describes exactly this mechanism for long conversations: "Claude automatically manages long conversations. When your conversation approaches the context window limit, Claude summarizes earlier messages" (doc-set §5, source S2) — and "a portion of the context window is always reserved for Claude's response" (same source). Trade parallel: the job box on site only holds so much paperwork. Once it's full, older pages get condensed to make room for new ones — and a detail from page one can go missing by page nine, not because anyone lied, but because the box has a limit.
  • SEE: Your real session, early on: you set the running example as "an 1,800 sq ft dwelling with a 10kW range." By question 8 or 9, deep in the same long conversation, an answer uses a different square footage — quietly inconsistent with what you set at the start. Nobody re-asked the question wrong; the constraint from page one got lost in a conversation long enough to summarize.
  • DO: Scan the later half of your real session for a place where an earlier constraint — your format, your numbers, your classroom style — got dropped or contradicted. Name it specifically as context loss, not "it got confused."

Independent at-bat (fully unscaffolded — the loop from lesson 1 closes here)

Take your real long load-calc session in full. Go through it moment by moment and label each place you'd flag: fabrication, training-data gap, context loss, or none of the three (a plain wrong number from a real, current, in-window source is still just a wrong number — not every mistake needs one of these three names). No worked example is given this time. This is the same sorting move lesson 1 previewed with three bare labels — now you're the one applying real names to real material.


Exit ticket (Bloom's climb: Remember → Understand → Apply → Analyze → Evaluate, capstone question)

  1. (Remember) Name the three failure modes covered today, and the doc-set source behind each.
  2. (Understand) Why is "it got confused" not a specific enough answer for any of the three? One sentence per mode.
  3. (Apply) Here's a new moment from your session, not yet labeled: pick one and name which of the three it is (or state clearly that it's none of them).
  4. (Analyze) Which of the three modes showed up most in your real session — and does that match, or differ from, the mode you guessed back in lesson 1 before you knew any of their names?
  5. (Evaluate — objective level, integrating the unit) Take the fabricated-citation moment from cycle 1. Walk the full loop, one line per anchor: would you have delegated this task at all (A2)? What would you have asked Claude to flag about itself first (A3)? What did checking it actually show (B1)? And now — what specific failure mode do you name it (B2)?

Graded against your actual code book, your actual session transcript, and the three named mechanisms from the doc-set — never against a guess at what "probably" happened.


Ledger write

ledger_append:
  learner_id: L4-PUB-FRANK
  lesson_id: L4-lesson06-b2-failure-literacy-20260719
  standard: AIHC.1.B2
  exit_ticket: {score: "", bloom_reached: evaluate}
  auto_mastery: "0.15 ->  (projected ~0.50 — correctly distinguished fabrication from an ordinary wrong number in cycle 1; named all three modes with sourced reasoning; capstone question integrated all five unit anchors)"
  self_score: ""
  calibration_gap: " — not computed until a real self-score exists"
  journal_entry_prompt: >
    Close the unit out the way you'd close a training-committee meeting: which of the three
    failure modes are you most likely to run into again with your own apprentice materials,
    and what's the one habit from this whole unit — asking clearly, choosing what to hand
    over, asking Claude's own read, checking the work, or naming what went wrong — you're
    actually going to keep using next week?
  structure_used: PBL
  referents_used: [electrical-trade code-inspection culture — job-box / paperwork-limit framing, primary throughout]
  next_lesson_seed: "Unit complete for the 5 named anchors (A1, A2, A3, B1, B2). Ledger now shows a full mastery profile across the anchor set; a follow-on unit could extend to AIHC.1.B3 (trust calibration, already baselined in his profile at 0.10) or AIHC.1.C1 (provenance), reading forward from this session's demonstrated levels rather than resetting."

Rubric self-audit (R1–R13)

# Indicator Verdict Evidence
R1 One Bloom's-leveled objective, named standard, plain-language too PASS Analyze-level objective stated technically and plainly; AIHC.1.B2 quoted verbatim with plain gloss
R2 Every capability/limit claim traces to the doc-set PASS S5 (hallucination/fabrication, training-data-gap language, both quoted), S2 (context-window auto-summarization, quoted) — each cited by section/source; doc-set's own freshness note used honestly as a self-referential example, not overclaimed
R3 3–6 SSD cycles complete PASS 3 cycles, SAY/SEE/DO complete each
R4 Each SEE ground-truth-verified, no strawman errors PASS Each SEE illustrates the doc-set's own named mechanism specifically (a citation that doesn't hold up; an uncaveated recent-topic answer; a dropped constraint across a long session) — distinguishable from each other and from an ordinary wrong answer, per cycle 1's own explicit contrast
R5 Every DO on real artifacts, free-tier only PASS All cycles and the at-bat work directly on Frank's real long load-calculation session and real code book; clipping reused once (cycle 2); no paid feature invoked
R6 Media doctrine honored PASS All SEEs are static described moments/annotations from a transcript; no motion content needed
R7 Exit ticket 3–5 Qs climbing to objective level PASS 5 questions, Remember→Evaluate, the capstone question grounded against his own transcript and the doc-set's three named mechanisms
R8 Ledger write complete PASS Standard, exit ticket, auto/self , journal prompt in-register, unit-closing next-lesson seed all present
R9 Scaffold fade complete PASS Only cycle 1 supplies a worked contrast (fabrication vs. an ordinary wrong number); cycles 2–3 and the at-bat supply no filled-in example at all — the lightest scaffold of the six lessons, and the at-bat explicitly closes the loop lesson 1 opened, showing the arc's full fade end-to-end
R10 Referents elected, flavor-only PASS Only the electrical-trade referent used (job-box/paperwork-limit framing); never substitutes for the doc-set's three named mechanisms
R11 Plain-language-first PASS "hallucinating"/fabrication," "training-data gap," and "context loss" each introduced with both a trade analogy (fake citation, out-of-date book, overstuffed job box) and the doc-set's own real term
R12 Non-replication PASS Every DO is keyed to Frank's specific real load-calculation session (his own square-footage example, his own citation moments) — none of this exists for another learner's profile
R13 No deficit-framing, no fear-framing PASS All three failure modes are presented as documented, named, expectable mechanisms with a specific tell each — never as reasons to distrust the tool broadly or as evidence Frank missed something he should have caught sooner; the capstone question treats his full five-anchor competence as already-built, not remedial

Escalation: all load-bearing indicators pass → auto-ship.