← Home 05 / The derivation protocol

How the standards were derived — in the open.

Rigor you can audit. From ten months of raw transcripts to a ratified taxonomy, every step is documented, and every quote is verifiable. This is the rigor exhibit — the JD's “rigorous about measurement” applied to the standards themselves.

● Docent note

Read this as a chain of custody. The claim is not “an expert says these are the standards” — it is “here is the corpus, here is how it was mined, verified, adjudicated, and versioned, and here is where you can check any link.” The one thing it is not is a consensus standard: it is a single-practitioner method demonstration, stated plainly in the scope note below.

The chain, step by step

Raw transcripts → ratified taxonomy

01

Frame first

NGSS/CCSS hybrid method

The construct — AI-human collaborative work capability — was defined before the evidence was mined, then the corpus populated it rather than inventing it. Three dimensions (practices, core ideas, crosscutting concepts) map to anchors, back-written across four proficiency bands with published criteria. This is the same evidence-centered-design chain used to build real academic standards.

02

Mine recall-first

Sonnet finder fleets under a published spec

Candidate practices were mined from ten months of the practitioner's real working record (web-chat and CLI eras) by fleets of Sonnet finders operating under a published finder specification — a minimum-articulation threshold and a verbatim-or-dead quoting rule. Recall-first: cast wide, verify hard.

03

Verify every quote

Byte-verification against raw transcripts

Every candidate quote was byte-verified against its raw transcript window. Zero fabrications across 1,409 pilot and fleet candidates. A quote that could not be located verbatim was killed, not kept.

04

Adjudicate the structure

Human ruling, with licensed dissent

The family structure was adjudicated by Krystal through a decision walkthrough — seven structural rulings (D1–D7) plus three cross-cutting rules. Cross-tier dissent was licensed: one divergence was adopted over the author's recommendation, which is itself now evidence for the elicitation anchor.

05

Adjudicate every member

Full-Opus pass after a blind audit

A blind 15% audit measured 18.0% single-pass disagreement — so Krystal ruled a full Opus adjudication of all 1,702 members. 48 records were double-adjudicated at 88% independent agreement; six residual disagreements were tie-broken with per-record documented reasoning.

06

Version for regeneration

v01 → v02 → v03

v01 derived practice evidence from a paraphrase of a handwritten index; v02 re-derived it from the adjudicated corpus with exact per-family receipts; v03 applied Krystal's seven walkthrough rulings. Superseded versions deprecate, never delete — the standard is built to be re-derived, not perfected against obsolescence.

The recursion

The standards were produced under the standards.

The corpus was mined under a published spec (specification), by tiered fleets (delegation), byte-verified (verification), with failures turned into rules mid-run (failure literacy), full provenance and corrected attribution (provenance), instrumented estimation (measurement), licensed cross-tier dissent that changed a ruling (elicitation), state surviving multiple session deaths (orchestration), and every structural change adjudicated by the human (governance). The evidence for the standards produced itself under the standards.

Scope, stated plainly

A single-practitioner method demonstration

These standards demonstrate that the derivation method works on a real corpus. Their evidence base is one expert practitioner's adjudicated working record plus one hiring organization's readiness specification. They are not validated for adoption as a team, organizational, or public standard — standards for public consumption would require re-derivation from a diverse set of practitioner corpuses. The method scales; this instance is deliberately, disclosed-ly, one person.

The corpus records what an effective practitioner does; it is observational, not experimental. It does not by itself establish that these practices cause superior outcomes — the causal warrant is the employer's (they pay for this capability) and the corpus's descriptive one, with an outcomes study named as future work.