Lesson 10 — Measurement: Outcomes He Defines, Not Dashboard Completions
Standard: — · Bloom's: — · Structure: PBL.
Notice the Say-See-Do cycles running on Devon T.'s actual work — not a canned exercise — and that every capability claim is cited to the frozen doc-set. The exit ticket climbs Bloom's to —, and the lesson closes by writing to the ledger.
Generation-time decisions (logged, per Playbook §3–5)
| Decision | Value | Why |
|---|---|---|
| Learning structure | PBL | "is my rollout working?" is Devon's live question, not a case study |
| Objective (one) | Evaluate the rollout against measures of working that Devon defines, using the dashboard only for what it can actually answer | first C3 lesson |
| Standard | AIHC.2.C3 | Band 2: instrument the collaboration and act on what the data says — with measures designed before collection, not after |
| Bloom's | Evaluate | judging metric validity against a stated purpose is evaluation, not lookup |
| Cycles | 4 | within the 3–6 band |
| Accommodation | option-suppression | ONE recommended measurement set; alternatives on request |
Learning objective
(Bloom's: Evaluate) Evaluate whether Bridgeway's Claude rollout is working, using (a) what the usage-analytics dashboard actually reports, and (b) outcome measures you define yourself for the grant-report workflow — and defend why each chosen measure evidences "working" rather than merely "being used."
Standard named: AIHC.2.C3 (Measurement). Trust what the data shows, not what adoption feels like (X1); and watch the scale–quality tension (X3) — usage going up tells you scale, it does not tell you quality. Note the boundary with Lesson 09 (C2): governance built the rules; measurement asks whether the rules — and the rollout they govern — are producing anything worth having.
Cycle 1 — What the dashboard actually gives you (and to whom)
SAY — Per the doc-set, the analytics surface is real and specific: click your initials (lower
left) → "Analytics." Named overview metrics: "Weekly active members" · "PRs created in
Code" · "Sessions in Cowork," plus adoption, product-usage, results, and spend sections.
Spend export: Settings > Analytics → "How much is Claude costing?" → "Export spend report"
→ MTD / Last Month / Last 90 Days / Custom → "Download" — a CSV with user email, model,
token counts, and "Net spend (total_net_spend_usd)." Two honest boundaries at your tier,
named now so they don't surprise you later: Analytics chat — asking Claude questions about
your org's usage in plain language — is Enterprise only, so Bridgeway reads the dashboard
directly; and the finer-grained Enterprise viewing rules (like "Admins can view all analytics
except Spend") govern a plan you're not on.
SEE — (static annotated list) The dashboard's named sections, each tagged with the one question it can answer: Weekly active members → "are people showing up?" · Sessions in Cowork → "are they using the tool we actually rolled out?" · Spend export → "what is it costing, by person and model?" — and a struck-through row: Analytics chat → Enterprise only, not available to Bridgeway.
DO — Open your real Analytics view and record three numbers as your baseline: weekly active members, Sessions in Cowork, and month-to-date spend from the export. Recommended path: record them in a dated row in your rollout-plan.docx, since that document already carries your enforcement-gap notes from Lessons 07–09. (Alternatives available on request.)
Cycle 2 — The scorecard test: usage is entries, not scores (BBQ referent)
SAY — A BBQ competition doesn't rank teams by how many entries they turned in — the judges score appearance, tenderness, taste, per box, against a standard. "Forty entries submitted" is participation; it says nothing about whether any brisket was good. Your dashboard numbers are entries: 30 of 40 staff active weekly proves adoption, not that a single grant report got better, faster, or safer. (Referent: BBQ judging scorecard — flavor only; the measurement-validity point stands on its own.) This is the vanity-metric trap the whole EdTech field ships by default: completion and satisfaction standing in for outcomes.
SEE — (static two-column table)
| Dashboard can answer (usage) | Dashboard cannot answer (outcomes) |
|---|---|
| Who used Cowork this week | Whether the Q3 Outcomes section needed fewer correction rounds |
| Sessions per team | Whether a donor letter shipped with an unverified claim |
| Spend per person/model | Whether your point-of-use rules are being followed |
DO — For each of your three baseline numbers from Cycle 1, write one sentence: the question it does answer, and the rollout question it cannot answer. No fixes yet — just the honest sort.
Cycle 3 — Define "working" before you collect (the measures are yours)
SAY — Design before collection: decide what "working" means, then instrument for exactly that — never collect first and rationalize after. The platform will not do this step for you; no doc-set source defines outcome measures for a grant-report workflow, and this lesson won't pretend one does. Recommended path — three outcome measures for the grant-report production workflow, one per thing you actually care about: (1) verification coverage — of the claims in each shipped Outcomes section, how many carry the Lesson 08 attribution footer (drafter + verifier + source file); (2) correction burden — how many funder-bound figures got corrected after Devon's review stage rather than before it; (3) rule adherence — in a monthly spot-check of one Cowork session per team, was the Lesson 07 permission lane actually the one in use. (Alternatives — cycle-time measures, staff-confidence surveys — available on request; not expanded here.)
SEE — (static worked example) One row of a measurement log, filled in for the real Q3
report: Q3 Family Support Outcomes · claims: 3 · footered: 3/3 · post-review corrections: 0 ·
spot-check: Grants team, Manual lane confirmed — each cell traceable to a thing Devon can
actually count, no estimate anywhere.
DO — Adopt or adapt the three measures for your real workflow and write the empty log template into your rollout-plan.docx, with a named owner and a monthly cadence for each measure. The platform's contribution stays in its lane: the dashboard feeds usage context; these three feed the "working" verdict.
Cycle 4 — Act on what you find, and check the cost side honestly
SAY — A measure no decision hangs on is decoration. Attach one pre-committed action to each measure: what changes if verification coverage drops below 100%? If post-review corrections climb, which lesson's habit gets re-run with which team? And pair the outcome view with the real cost view you already have: the spend export's per-user, per-model CSV. One caution from the doc-set's own scope notes, so you don't import someone else's numbers as your evidence: the deployment guide's landing page advertises a five-level maturity model and roadmap, but the numbered stats live in a separate PDF eBook the doc-set explicitly did not freeze — do not cite those figures as measurement evidence for your rollout. Your log, not a vendor eBook, is your ground truth.
SEE — (static three-line policy) IF footer coverage < 100% on any shipped report → that
team re-runs the Lesson 08 at-bat before its next report ships. IF post-review corrections > 0
two months running → the affected team's Lesson 05 verification habit is re-drilled. IF
spot-check finds a lane mismatch → Lesson 07's policy-only finding escalates to a standing
agenda item.
DO — Write your own three IF→THEN lines, keyed to your three measures, into the rollout-plan.docx next to the log template. One line each, pre-committed now — not decided in the moment the number goes bad.
Independent at-bat
The Development team's donor-acknowledgment workflow has no measures yet. Design its measurement row unscaffolded: three outcome measures Devon defines (not usage stats), a named owner and cadence for each, and one IF→THEN action per measure — using the dashboard only where it genuinely answers the question.
Exit ticket (climbing to Evaluate; graded against the doc-set + his own artifacts)
- (Remember) Name the three overview metrics the Analytics page reports, per the doc-set, and the one analytics feature Bridgeway's tier does not have.
- (Understand) Why can "Sessions in Cowork" go up while the rollout is failing? One sentence, in usage-versus-outcome terms.
- (Apply) The board asks: "Is the Claude thing working?" Give the two-part answer structure this lesson taught — which number comes from the dashboard, and which comes from your log.
- (Analyze) Your spend export shows one staffer at triple the median spend. Name two different explanations — one that's a problem, one that isn't — and the outcome measure that distinguishes them.
- (Evaluate — objective level) A colleague proposes "percent of staff who completed Claude onboarding" as the rollout's headline success metric. Render a verdict on that metric for Bridgeway's actual goal (grant reports shipping verified, on time, with clean provenance), and defend your verdict using the design-before-collection rule and at least one measure you defined today.
Ledger write
ledger_write:
learner_id: L2-ADMIN-DEVON
lesson_id: L2-admin-devon-10-c3-measurement
anchors: [AIHC.2.C3]
tags: [X1, X3]
exit_ticket:
score: ""
bloom_reached: ""
auto_score: ""
self_score: ""
calibration_gap: ""
journal_prompt: >
Before this lesson, if the board had asked "is it working?", what number would you have
reached for — and does that number survive the entries-versus-scores test you ran today?
structure_used: PBL
referents_used: [bbq-judging-scorecard]
next_lesson_seed: "Lesson 11 (D1 Orchestration) turns the grant-report workflow the measures now watch into an explicit multi-person production line — stages, handoffs, and who verifies what."
RUBRIC SELF-AUDIT (against Exemplar_and_Rigor_Rubric_Cowork_x_Nonprofits_2026-07-19_v01_I.md)
| # | Indicator | Verdict | Evidence |
|---|---|---|---|
| R1 | ONE Bloom's-leveled objective, ≥1 named AIHC standard | PASS | One Evaluate objective; names AIHC.2.C3 |
| R2 | Every platform claim traces to the doc-set; gaps named, not invented | PASS | All analytics/spend-export claims quoted from doc-set §6; analytics-chat named Enterprise-only; eBook stats explicitly excluded per doc-set §8/§10 |
| R3 | 3–6 SSD cycles, complete | PASS | 4 cycles, all SAY/SEE/DO complete |
| R4 | Each SEE ground-truth-verified | PASS | SEEs render only doc-set-named metrics/paths and Devon's own artifacts; the worked log row uses his real Q3 claims count |
| R5 | Every DO on real artifacts | PASS | DOs write into the real rollout-plan.docx and measure the real grant-report workflow; at-bat on the real donor-letter workflow |
| R6 | Media doctrine (static concepts → static visuals) | PASS | All SEEs static lists/tables; no video |
| R7 | Exit ticket 3–5 Qs climbing, SSOT-graded | PASS | 5 Qs, Remember→Evaluate, graded against doc-set §6 facts + his own defined measures |
| R8 | Ledger write: standards, both scores, journal | PASS | anchors/tags present, scores, journal_prompt is a question |
| R9 | Band-appropriate scaffolding with fade; at-bat independent | PASS | Cycle 3 gives a fully worked measure set, Cycle 4 gives worked IF→THENs then demands his own; at-bat designs a full measurement row unscaffolded |
| R10 | Referents elected-only, flavor-only | PASS | BBQ judging scorecard used once, framing the vanity-metric point; the measurement claims stand on doc-set + design-before-collection grounds |
| R11 | Option-suppression honored | PASS | Cycles 1 and 3 each state ONE recommended path with alternatives noted parenthetically, never enumerated |
| R12 | Non-replication | PASS | Measures are keyed to Bridgeway's specific workflows, prior-lesson footers, and permission lanes; a different learner's ledger would produce different measures and IF→THENs |
| R13 | No unenforceable/unavailable practice advised without naming the gap | PASS | Analytics chat named Enterprise-only up front; the dashboard's outcome-blindness named as a platform limit Devon's own log covers; eBook figures barred as evidence |
Escalation check: no load-bearing indicator fails → auto-ships per the versioned escalation policy (Playbook §5).