← Home 02 / The standards

What “great” means, named precisely enough to hit.

The AI-Human Collaboration & Coworking (AIHC) learning standards, v03. An NGSS/CCSS-style framework that defines AI-collaboration capability as observable, band-leveled performance — so a team can hit the bar without the author in the room.

8Practices
6Core ideas
11Crosscutting
11Anchors
45Expectations
4Bands
● Docent note

This is the JD's core ask, answered: “define what great means precisely enough that the team can hit it without you in the room.” Start with the eleven anchors below — each is one capability, defined once and then written as a rising performance expectation across four proficiency bands. Filter by strand or band; open any anchor to read its expectations and the practices, core ideas, and crosscutting concepts it draws on. Everything is cross-linked: an anchor's tags jump to the dimensions, and the lessons in the Tracks cite these codes directly.

The four bands

Assisted Use → Systems Stewardship

Every anchor is written four times, once per band — the same capability, rising from personal tool-use to designing the systems others work within.

Band 1

Assisted Use

AI as a tool on bounded personal tasks; single sessions, single model.

Band 2

Governed Collaboration

AI as a routine coworker; specifications, verification, provenance, and policy become explicit artifacts.

Band 3

Orchestrated Production

AI as a workforce; multi-agent pipelines with enforced gates, measurement, and deliberate delegation architecture.

Band 4

Systems Stewardship

The practitioner designs the standards, policies, and infrastructure others collaborate within — encoding craft into systems rather than guarding it.

The heart of it

Eleven anchors, four strands

Direction · Judgment · Stewardship · Evolution. Filter, then open an anchor to read its four-band progression.

A1

Specification

Author specifications precise enough that a competent AI coworker meets the quality bar without the author in the room.

P1CI5X3
Performance expectations by band
AIHC.1.A1

Write task requests that state goal, inputs, constraints, and what “done” looks like before handing work to an AI coworker; when output misses, revise the request before re-rolling the output.

AIHC.2.A1

Author reusable specifications — briefs, rubrics, acceptance criteria — precise enough that an AI coworker meets the bar without mid-task correction, and calibrate a rubric against known-good and known-bad exemplars before trusting it.

AIHC.3.A1

Maintain a specification system — briefs, rubrics, gates — that lets multiple AI workers produce to one bar across a pipeline, with depth chosen per task risk and, for empirical claims, rubrics that score falsifiability, construct validity, and independence of verification.

AIHC.4.A1

Encode the organization's quality bar into specifications, standards, and enforcement mechanisms that hold without their author in the room, and revise them on production evidence.

A2

Delegation

Match work to AI capability using evidence, cost, and reversibility — with the routing itself documented and defensible.

P2CI1X6
Performance expectations by band
AIHC.1.A2

Decide whether to hand a task to an AI coworker from stated reasons about capability and the cost of being wrong — not novelty, habit, or hype.

AIHC.2.A2

Match tasks to models and configurations deliberately, documenting the capability evidence and reversibility rationale behind each routing.

AIHC.3.A2

Operate a delegation architecture — capability tiers, role definitions, spawn governance — in which every delegation carries an explicit model choice, scope, and escalation path.

AIHC.4.A2

Set and evolve the delegation policy itself — the capability ladder, defaults, and exception rules — from accumulated evidence, re-deriving it at each model generation.

A3

Elicitation new in v03

Solicit and use the AI coworker's own knowledge — proposals, self-diagnoses, dissent — as first-class input, with the license to diverge made explicit and divergence treated as data rather than defiance.

P2P3CI1X10
Performance expectations by band
AIHC.1.A3

Ask the AI coworker for its read, its proposal, or its uncertainty before finalizing an approach, and treat “what would you do?” as a working question rather than a courtesy.

AIHC.2.A3

Make elicitation routine and structural: require the AI coworker to propose the fix for its own failures, surface its open questions at defined checkpoints, and record adopted proposals with the AI credited as source.

AIHC.3.A3

Build licensed divergence into pipelines: reviewers and presenters are authorized to disagree with the artifact they carry, divergences are labeled and logged rather than silently reconciled, and a surviving divergence changes the artifact.

AIHC.4.A3

Design the elicitation architecture for a team: the protocols, prompts, and credit systems that make AI coworkers' knowledge, dissent, and self-diagnosis flow into the production system by default.

B1

Verification

Treat all AI output as unverified claims; converge on correctness through source-grounded audit against ground truth outside the AI's evidence chain, and treat agreement among correlated workers as a prompt to verify, not to trust.

P3CI2CI5X1X4X8
Performance expectations by band
AIHC.1.B1

Check AI output against at least one independent source before using it, and reject plausibility, fluency, confidence, and the AI's own assurance of its rigor as substitutes for grounding.

AIHC.2.B1

Run structured verification — source-grounded audit against ground truth outside the AI's own evidence chain, adversarial reading — on any AI artifact entering shared work, with findings recorded as findings; treat unanimous agreement among AI workers as a trigger for external checking, not a confidence signal.

AIHC.3.B1

Build convergence gates into pipelines so verification is enforced by mechanism rather than memory, with raters independent of authorship and — where consensus is the signal — decorrelated by architecture or training lineage.

AIHC.4.B1

Design the organization's verification standards — which work requires which depth of audit — and audit the auditing: verify that gates catch seeded errors.

B2

Failure literacy

Anticipate, recognize, and classify AI failure modes; convert incidents into enforced rules.

P7CI2X4
Performance expectations by band
AIHC.1.B2

Recognize common failure modes in the wild — fabrication, ungrounded confidence, context loss — and name what happened specifically rather than “it got it wrong.”

AIHC.2.B2

Classify incidents against a failure-mode taxonomy and adjust their own practice in response — grounding earlier, specifying tighter, gating harder.

AIHC.3.B2

Run the incident-to-rule cycle end to end: capture failures with provenance, attribute cause, and convert repeat classes into rules enforced structurally — by non-bypassable mechanism, because advisory rules fail stochastically.

AIHC.4.B2

Maintain the failure-mode taxonomy itself: extend, split, and retire failure classes as models change, and propagate the rules that encode them across the systems they govern.

B3

Trust calibration

Maintain an explicit, evidence-based map of what each AI coworker is trusted with; re-derive it as models change; never count agreement among same-lineage models as independent confirmation.

P8CI1CI2X1X2X8
Performance expectations by band
AIHC.1.B3

State what they currently trust an AI coworker to do without review and what they never accept unreviewed, with a reason for each.

AIHC.2.B3

Keep the trust map explicit and test it: spot-check the trusted zone, record surprises in both directions, adjust scope rather than sentiment, and never treat agreement among same-lineage models as independent confirmation.

AIHC.3.B3

Recalibrate the trust map on model releases through deliberate probes of known capability edges — including checks for failure modes shared across the model fleet — rather than letting anecdote accumulate into policy.

AIHC.4.B3

Publish and defend the trust line for a team — where AI acceleration is expected, where human craft is non-negotiable — and move it on evidence as models improve.

C1

Provenance

Keep the full human/AI provenance of every artifact inspectable — idea source, generator, verifier, version.

P4CI4X5
Performance expectations by band
AIHC.1.C1

Label work they hand off or publish with what the AI did and what they did.

AIHC.2.C1

Keep artifact-level provenance — idea source, generator, verifier, version — as a routine part of producing work, preserving superseded versions rather than overwriting them.

AIHC.3.C1

Operate provenance protocols of the AIGHVA class across a pipeline, so every claim in a deliverable is origin-traceable through its audit trail.

AIHC.4.C1

Design attribution and provenance systems — registries, protocols, statutes — that make full provenance the path of least resistance for everyone working in the system.

C2

Governance

Operate within — and, at proficiency, author — explicit AI-use policies; know where AI is disallowed and why; secure consent where work touches persons or agents.

P2P4CI6X6
Performance expectations by band
AIHC.1.C2

Find and follow the applicable AI-use policy before using AI on consequential work; comply with stage-appropriate restrictions (including “no AI here”) without supervision.

AIHC.2.C2

Apply stage-appropriate AI use across a work process — where AI drafts, where it refines, where it is excluded — and disclose per the governing policy's transparency norms.

AIHC.3.C2

Operationalize governance: encode policy into workflow — gates, consent artifacts, IRB-analog review where the work touches persons or agents — so compliance is structural, not aspirational.

AIHC.4.C2

Author governance: write the AI-use policy for a team or organization, with its enforcement mechanisms, disclosure norms, and revision cadence built in rather than bolted on.

C3

Measurement

Instrument the collaboration — estimates, actuals, quality outcomes, differential impact — and let the data recalibrate practice.

P6CI5X1X7
Performance expectations by band
AIHC.1.C3

Estimate before delegating and compare the estimate with what actually happened, in writing.

AIHC.2.C3

Instrument recurring collaboration — task timing, estimate calibration, quality outcomes — and let accumulated data, not memory, adjust their estimates.

AIHC.3.C3

Measure whether the collaboration system works — gate catch rates, rework rates, calibration drift — and feed the results back into specifications and delegation policy.

AIHC.4.C3

Define what “working” means for AI-accelerated production and own its measurement architecture — including monitoring of differential impact across the people the system serves.

D1

Orchestration

Compose multiple AI workers, tools, and gates into one accountable production system rather than a chain of handoffs.

P5CI3CI4X3X7
Performance expectations by band
AIHC.1.D1

Chain AI assistance across the steps of one task — draft, critique, revise — rather than accepting single-shot output.

AIHC.2.D1

Coordinate parallel AI workers on decomposed work with explicit handoffs, each handoff artifact readable by a cold reader with no session context, and in-progress state engineered to survive session death — durable artifacts committed before stepping away, successor briefs written for the thread that will resume the work.

AIHC.3.D1

Run gated multi-agent pipelines end to end — research, design, authorship, review, QA, publication — as one accountable production system.

AIHC.4.D1

Design orchestration architectures for others — roles, tiers, gates, escalation paths — and evolve them as one production engine rather than a chain of handoffs.

D2

Regeneration

Design artifacts, workflows — and standards — for durability through regeneration: versioned, deprecable, re-derivable at low cost, rather than perfected against obsolescence.

P1P4CI3X2
Performance expectations by band
AIHC.1.D2

Version their artifacts, expect them to be superseded, and keep the inputs that produced them.

AIHC.2.D2

Prefer regenerable artifacts — specification plus sources plus method — over hand-perfected ones the next model generation will outdate.

AIHC.3.D2

Design workflows for fast regeneration: when a model, source, or standard changes, affected artifacts can be re-derived, re-gated, and re-shipped at low cost.

AIHC.4.D2

Treat standards and craft themselves as regenerable: encode them, deprecate them on schedule, and re-derive them as models and evidence move — this document included.

The three dimensions

What the anchors are built from

Every anchor draws on coworking practices (what you do), core ideas (what is true about AI collaboration), and crosscutting concepts (the themes that recur everywhere). Open any item to read it.

Coworking practices · P1–P8

P1 Specifying work

Translating intent into executable specification — briefs, rubrics, quality bars, acceptance criteria — precise enough that an AI coworker can meet the bar without its author in the loop.

P2 Delegating deliberately

Choosing what to hand off, to which model or configuration, with what scaffolds — matched to evidence of capability, cost, and reversibility rather than habit or hype.

P3 Verifying against sources

Treating AI output as unverified claims; converging on correctness through source-grounded audit, adversarial review, and gates — never plausibility.

P4 Keeping provenance

Maintaining the traceable record of every artifact's human and AI contributors — idea, generation, verification — including disclosure, versioning, and preservation of superseded versions.

P5 Orchestrating multi-agent work

Composing multiple AI workers, tools, handoffs, and gates into one accountable production system.

P6 Instrumenting and calibrating

Measuring the collaboration itself — estimates against actuals, quality outcomes, gate performance — and letting the data recalibrate practice.

P7 Diagnosing failure

Recognizing, classifying, and attributing AI failure; converting incidents into rules and mechanisms (an error-to-rule cycle).

P8 Renegotiating the line

Re-deriving what to trust AI with — and where human craft is non-negotiable — as an explicit, evidence-based, continuously re-evaluated decision.

Core ideas · CI1–CI6

CI1 Model capability and variability

Capabilities are empirical claims requiring evidence, varying by model, version, configuration, and context. A capability verdict must first rule out a scaffolding gap.

CI2 Failure modes

AI failure is systematic, not random: fabrication indistinguishable from grounded output, ungrounded confidence, context loss, instruction drift, register inflation, correlated failure. Failure modes can be taxonomized and designed against.

CI3 Context, memory, and grounding

An AI coworker's output is bounded by what it can actually see. Grounding must precede prompting; norms that matter need enforcement mechanisms, not advisory notes.

CI4 Provenance and attribution systems

Protocols of the AIGHVA class (human idea → AI generated → human verified) — versioning, audit trails, attribution registries — are what make AI-produced work trustworthy, improvable, and honest.

CI5 Quality specification and assessment

Quality bars exist only when specified: rubrics, standards, gates, and the construct validity of one's own evaluations — measuring what you mean to measure.

CI6 Governance, consent, and policy

Organizational AI-use policies, stage-appropriate use (including where AI is disallowed), consent and IRB-analog review where work touches persons — or the agents themselves — and the safety boundaries of the tools.

Crosscutting concepts · X1–X11

X1 Trust as evidence

What is trusted, is trusted because of documented evidence, at a stated scope, subject to revision.

X2 The moving line

Every human/AI division of labor is temporary. The line is re-derived — not defended — as models improve.

X3 Scale–quality tension

Scaling output without lowering the bar requires the bar to be encoded, because an unencoded bar cannot scale past its keeper.

X4 Enforcement over advisory

Norms that matter get mechanisms — gates, hooks, protocols that fail closed — not reminders.

X5 Provenance as first-class

An artifact whose origin cannot be traced cannot be trusted, verified, credited, or improved.

X6 Reversibility and blast radius

Delegation freedom scales with reversibility; irreversible or outward-facing work gets gates and human judgment.

X7 Human bandwidth as the scarcest resource

The collaboration is designed so humans spend attention only where humans are irreplaceable; everything else is delegated, systematized, or made cold-reader-ready.

X8 Convergence is not validity

Agreement among AI outputs is evidence only to the extent the outputs are independent. Correlated workers can converge on the same error, so unanimity is a prompt to verify against ground truth — not a confidence signal.

X9 The ask/act boundary v03

Never ask for permission you already have; always ask about requirements you cannot derive. The skill is classifying which situation you are in before choosing.

X10 Bidirectional knowledge flow v03

The AI coworker's own solutions, readings, self-diagnoses, and dissent are solicited and credited as first-class input. Recurs across delegation, verification, and orchestration.

X11 Deficit-free engineering — design around the actual participants v03

An unmet need is read as a gap in what the collaboration provided, not a defect in who it was provided to — for both humans and models. One pedagogy, both directions.

Where these came from

Not asserted from taste — derived from evidence.

These standards were derived from 1,702 adjudicated receipts of one expert practitioner's real working record. See the evidence, and the method.