Lesson 10 — Measurement: tracking what claude.ai actually improves in your work
Standard: — · Bloom's: — · Structure: —.
Notice the Say-See-Do cycles running on Priya S.'s actual work — not a canned exercise — and that every capability claim is cited to the frozen doc-set. The exit ticket climbs Bloom's to —, and the lesson closes by writing to the ledger.
Learning objective
(Bloom's: Analyze) Analyze your own claude.ai usage on The Last Mill and your interview-prep workflow to determine which settings and habits are actually saving you time or improving your draft, versus which are just consuming usage allowance without a measurable payoff.
Standard named: AIHC.1.C3 (Measurement). Measurement means instrumenting your own practice and calibrating from what it shows (P6) — because your attention and your usage allowance are both the scarcest resources here (X7), not infinite ones to spend on a hunch.
Cycle 1 — Not every turn costs the same
SAY — Some actions draw more from your usage allowance than a plain chat turn. Verbatim: "File creation draws on the same usage allowance as chatting, and consumes more of it than a plain-text turn" [S08]. Separately, "Research uses the same usage-limit pool as normal chat, but sessions burn through it faster" [S11]. The doc-set gives no exact numbers or ratios — only this ordering — so the fair comparison is relative cost, not invented figures.
SEE — A described ordering diagram (no fabricated numbers): "plain chat" on the left, "file creation" in the middle labeled "costs more than plain chat [S08]," "research" on the right labeled "burns the pool faster than normal chat [S11]" — an order, not a scale.
DO — List your last five real Claude interactions on this project and tag each by type: plain chat / file creation / web search / research.
⏸ Pause point. Five interactions tagged. Stop here if needed.
Cycle 2 — Build the log Claude can build for you
SAY — Claude can produce working files, including spreadsheets "with working formulas" [S08], for workflows like "financial models with live formulas" [S08] — the same mechanism repurposed as a personal productivity tracker instead of a budget. Referent (true-crime podcasts, flavor only): a careful investigator keeps a running case log, not a memory of what helped last time.
SEE — A described spreadsheet artifact with columns: date / task / setting used / time spent / self-rated payoff.
DO — Ask Claude to create this tracker as a real file [S08], seeded with the five real interactions you tagged in Cycle 1.
⏸ Pause point. Tracker created and seeded. Stop here if needed.
Cycle 3 — Read your own log for a pattern
SAY — Measurement isn't logging for its own sake — it's reading the log and acting on what it shows. Per the two facts from Cycle 1, a workflow that costs more (file creation, research) only earns its keep if the payoff column backs it up.
SEE — A described filled-in tracker after a week of real-shaped entries, with one visible pattern: "research" rows clustering at low self-rated payoff for simple fact-checks; "extended thinking" rows clustering at high payoff for structure work.
DO — Review your real (even if still sparse) log entries and write one sentence: which setting the data says to use more of, which to cut.
Independent at-bat
Unscaffolded: run one full week logging every real session in your tracker without a template prompt, then write one verdict paragraph — what claude.ai measurably improves for your work, and what doesn't earn its usage cost — citing your own log, not a general impression.
Exit ticket (climbing to Analyze; graded against the doc-set)
- (Remember) Name two actions that S08 and S11 say draw down usage allowance faster than a plain chat turn.
- (Understand) Why is "it felt helpful" not the same claim as "it measurably improved my draft"? Tie your answer to what a tracker column actually records versus a feeling.
- (Apply) You used Research for a source background-check last week and extended thinking for restructuring an outline. Fill in both rows of your tracker's format.
- (Analyze — objective level) From your real tracker so far, name one workflow habit the data supports keeping and one it supports cutting — justified from the log, not a gut sense.
Ledger write
ledger_write:
learner_id: L3-CONS-PRIYA
lesson_id: L3-10-C3
standards: [AIHC.1.C3]
tags: [P6, X7]
exit_ticket:
score:
bloom_reached:
auto_score:
self_score:
calibration_gap:
journal_prompt: >
Before this lesson, how did you decide whether a claude.ai habit was "working" for your
writing — a feeling, or something written down? What's the first thing your new tracker
would have told you that you didn't already know?
structure_used: PBL
referents_used: [true-crime-podcasts]
next_lesson_seed: "the tracker itself becomes a case study in Lesson 12 (D2 Regeneration) — a workflow habit worth re-measuring whenever the model or its costs change."
Rubric self-audit (R1–R13)
| # | Indicator | Verdict | Evidence |
|---|---|---|---|
| R1 | One Bloom's-leveled objective, ≥1 named AIHC standard, learner-visible | PASS | Objective names Analyze + AIHC.1.C3 |
| R2 | Every product claim traces to the frozen doc-set; no invented UI | PASS | S08 (file-creation cost, formulas), S11 (research pool-burn rate) quoted exactly; explicitly refuses to invent numeric ratios the doc-set doesn't give (Cycle 1 SAY/SEE) |
| R3 | 3–6 SSD cycles, complete | PASS | 3 cycles (within the 3–6 band): relative cost, build the tracker, read the pattern |
| R4 | Each SEE anchors its SAY | PASS | SEEs describe only an ordering diagram (no fabricated numbers) and a spreadsheet mockup matching S08's stated capability |
| R5 | Every DO acts on the learner's real work | PASS | Real interactions on her real project, real tracker seeded with her real data |
| R6 | Media doctrine | PASS | No motion; static/annotated only |
| R7 | Exit ticket 3–5 Qs, Bloom's-climbing, SSOT-graded | PASS | 4 Qs, Remember→Analyze |
| R8 | Ledger write complete | PASS | anchor codes, scores, journal_prompt present |
| R9 | Scaffolding with fade; at-bat present | PASS | Cycle 1–2 scaffolded, Cycle 3 lighter, at-bat (full week, no template) unscaffolded |
| R10 | Referents elected-only, flavor-only | PASS | True-crime used once, case-log analogy only |
| R11 | Timing-tolerance honored | PASS | Pause points after Cycles 1–2 |
| R12 | Non-replication | PASS | Keyed to Priya's own five real interactions, her own tracker data — cannot be reused unchanged for another learner |
| R13 | Consumer trust boundaries named, never over-reassured | PASS | Cycle 1 explicitly states the doc-set gives no numeric ratios rather than inventing false precision |
Escalation verdict: all load-bearing indicators PASS → clears; auto-ships. One non-load-bearing gap flagged above (no consumer usage-analytics dashboard exists in the doc-set) — handled by teaching Priya to build her own instrument rather than presenting a feature that doesn't exist.