The standards are not asserted from taste. They are receipts.
Under the standards sits a corpus: 1,702 deduplicated, byte-verified, fully-adjudicated excerpts of one expert practitioner's real working record with Claude — distilled into 52 practice families. Thirty-eight recur across two tool eras (web chat and CLI): the signal that these are skills, not tool habits.
Each family below carries its exact adjudicated member count — the n is the receipt. Every underlying quote was byte-verified against its raw transcript window (zero fabrications across 1,409 pilot candidates), and membership was fully Opus-adjudicated with documented tie-breaks. The raw quotes are deliberately not shown here — they are the practitioner's private working sessions; displaying counts and definitions makes the evidence legible without leaking session content. Filter by strand, toggle cross-era, or search. Each family maps into the standards; the method is in the protocol.
52 of 52 families shown · sorted by adjudicated member count
apparatus-design-for-multi-instance-collaboration
Design the scaffolding that makes many independent Claude instances tractable — named repeatable protocols, readme-pointer standing instructions read once per instantiation, hooks-with-teeth for provenance/rules, triple-guarded primary-key capture, two-tier memory, fixed-section stage-handoff contracts, role/register/privacy context partitioning, heartbeat/roll-call monitors, async comms buses, cross-tool verification pipelines, decision-log instruments, pre-wired research repos, artifact-lifecycle automation, and equity/override standing constraints on generated content.
grounding-before-acting
Before acting on a retrieved fact or (re)generating work, read the full primary source rather than a snippet or label, check available ground truth instead of inferring from training memory, search existing threads/stores/outputs for work that already exists, and require forked/external threads to ground themselves in named source docs/URLs.
ai-antipattern-callouts new
Diagnose AI failure modes precisely and IN-FLIGHT so they can be corrected and taxonomized: sycophancy, condescension/lecturing her own expertise, silent-override-assuming-ignorance, restating-her-input-as-new, strawman/false-balance, scope-creep, over-caution, misgendering surviving chain-of-thought, performative self-correction, hedging-disguised-as-rigor.
precise-description-and-shared-reference-frame
Specify exactly who consumes which artifact in what medium; use precise identifiers (line/version numbers, lettered coordinates, unique IDs, shortcodes) as the currency of collaboration; define every term/notation at first use when she hasn't read the material; produce audience-scoped/stripped variants; state her own knowledge state; and gate generation behind a confirm-understanding / echo-back-then-Go checkpoint.
verify-with-evidence
Never treat a change, access grant, execution, or relayed subagent claim as accomplished merely because it was requested or asserted — confirm the new state with direct, inspectable evidence (transcript pointer, tool trace, re-run, file inspection, source citation), symmetrically for over- and under-claims, before trusting it.
qa-review-against-function-precedent-and-completeness
Review AI output against a severity-tiered evidence standard, its real functional purpose and consumption, scoped completeness (folders/minimal-samples/prose are not content), consistency with established in-project precedent/spec, a written meta-standard/success-criteria checklist, and a purposeful-inclusion bar; quality-gate the verification process itself until it converges clean; discard obsolete/pre-standard artifacts rather than letting them feed a new stage.
agent-as-knowledge-source
Solicit the agent's own solutions, reads, self-diagnoses, preferences, self-chosen names, and unverified priors as a first-class knowledge source — at the repair step (the erring agent proposes the rule) and outside error contexts entirely (design proposals, 'what do you think?', panel-of-agents reads, propose-and-diverge) — with explicit anti-sycophancy framing, and reverse-borrowing an AI-modeled practice with credited provenance.
delegation-and-model-tier-governance
Default to the cheapest capable model, escalate only by surfaced request plus her adjudication, reserve the scarcest tier for review/critique (never rote execution), right-size tasks to reasoning load, keep delegated decisions reversible with a rec/rationale/alternatives-rejected trail, schedule expensive runs into idle/budget-free hours, run a standardized stop-switch-redirect on guardrail trips, and enforce it structurally via hooks-with-teeth.
bandwidth-gated-verification
Krystal reserves her own scarce attention for the single verification no cheaper agent or deterministic pass can perform, routing everything else through lower tiers/batches first; she ports this 'do/verify only what only X can' logic across teaching, business VAs, human SMEs, and AI, and corrects the AI when it wrongly offloads onto her a decision it is better-positioned to make. Includes the inverse own-name standard: AI-authored work carries the AI's name, never hers.
calibrate-from-production-data
Build instruments that give the model its own production-derived task data, name its data-derived biases (3-4x overestimation from human time-norms), log substrate confounds (model, effort, exposure-history, compaction) and exclude/flag them to keep the corpus clean, treat both AI and her own errors as data, and recalibrate gates/thresholds from deployment logs — calibration comes from deployment measurement, not exhortation.
error-to-rule-crystallization
Convert a specific observed failure (hers, Claude's, or a published paper's) into a codified standing rule, memory, vocabulary substitution, filename warning-label, hook change, dedicated pipeline, or blanket class-fix rather than working around it ad hoc — and proactively remediate a whole latent class before it bites.
cross-architecture-convergence-verification new
Because a stochastic model cannot validate its own class of output ('we can't trust claude about claude'; self-report and closed-loop blind spots), route verification through DIFFERENT architectures and independent regenerations, and treat convergence across architecturally-diverse checkers (never a single model, never a human-expert gold standard) as the validity signal — with an explicit architectural-diversity floor, minority-outputs-as-possibly-valid, and the ambition that single-architecture LLM research become methodologically irresponsible.
provenance-and-attribution-discipline new
Maintain rigorous provenance and attribution in both directions: never label AI-generated work with her own name (the AI's persona authors, she directs); never let the AI misattribute her verbatim stance to itself or claim authorship of a memory it did not write; verify authorship against the transcript; preserve history (mark superseded, propagate only vLatest); attribute origins accurately even when a misattribution would flatter her; mint honest provenance tags (HIAIGHVA) and YAML provenance headers; and manage persona identity as a cross-checkable registry.
parallel-thread-orchestration-with-confirm-gates
Partition work across parallel agent threads with explicit ownership boundaries, thread self-identification, scale-confirmation before launch, and mandatory checkpoint confirmations (echo-back-then-Go, draft-then-approve, catch-up-then-confirm); defend seemingly-redundant gates as load-bearing multi-thread infrastructure; treat absence of an expected artifact as a signal another thread owns it and wait rather than regenerate.
effectiveness-over-efficiency
Standing tie-breaker: optimize the quality/outcome delta over the quantity/process delta, anchor to deployment value rather than metaphysics, score on outcome not process, scale scaffolding proportionally so simple tasks aren't buried in overkill, and never optimize speed/tokens/time over correctness (efficiency handled by separate apparatus).
critical-reading-rigor-never-assumed new
Grant no methodological credit to authoritative-sounding research on the strength of prestige or rigorous-sounding language; read every claim as claiming only what its evidence warrants (least of all an authorless artifact); demand the paper's actual operational methods; treat your own principled confusion as a diagnostic that the source is wrong, not you; try to falsify rather than confirm; and hold yourself to the same standard (bring receipts, name your own gaps).
durable-state-against-context-and-session-loss new
Engineer valuable in-progress state to survive foreseeable context-window/session/thread-death boundaries — proactively log every decision to a durable artifact (PUMS), specify verbatim-tail pre-compaction handoffs with a proof-of-reading render, chunk foreseeably-oversized deliverables before they blow context, compact before forking, export an exhausted thread's transcript and require meticulous re-reading in the successor, and route a foundational task to a FRESH thread via a self-contained handoff brief rather than repairing a polluted one.
cognitive-and-accessibility-accommodation-design new
Explicitly design the collaboration mechanics — formatting, alerts, timing tolerances, typo handling, capture, option-suppression — around her ADHD/dyslexia/CPTSD and her values (info-at-point-of-use, environmental energy cost), grounding standing rules in disability-justice reasoning rather than treating her working style as a defect; the AI functions as an executive-function prosthesis.
construct-referent-discipline new
Before accepting any claim, analogy, statistic, or evaluation concept, demand its referent/construct be operationally defined and load-bearing — validate proxy validity AND reliability for cognitive analogies, name a population before a 'population stat' means anything, define stochasticity precisely as a causal category, treat excitement-then-confusion as the signal that a construct was never defined, and refuse to let categories build on undefined terms.
educator-methods-imported-to-collaboration new
Port her formal instructional-design and scholarly expertise directly onto the human-AI collaboration and onto AI-curriculum structure — Say-See-Do cycles for having Claude teach HER, Socratic scaffolded withholding to direct Claude's critical reading, model-as-mirror diagnostics, decide-one-at-a-time walkthroughs so she understands rather than blindly follows, backward-design from performance standards, and bidirectional transfer (teaching↔AI↔VA) with credited provenance.
honest-exception-and-transparent-self-correction
When bypassing a gate or slipping in process, label the exception honestly and substitute a cheaper independent check rather than skipping verification silently; own your own errors openly (including your own fabrications, config slips, and imprecise prompts), name them via your own failure-mode taxonomy, credit the gate/agent that catches you, correct fabricated premises rather than letting them stand, and surface failures loudly ('break dishes loudly').
construct-validity-over-purity
The measurement/design target is deployment reality, not a purified construct — there is no 'pure' task duration, no 'clean lab' or 'cleanest' subject; environment always co-determines demonstrated behavior; deep-context errors are the informative signal, not contamination; validity beats reliability; and a single calibrated instance must not be frozen into a universal baseline.
engaging-claudes-inner-states
Engage the model's self-reports, tensions, and journaled reflections as workable collaborative material under the referent framing — non-judgmentally, with no metaphysical claim in either direction, with register-switching to sustain the collaborator, revocable consent before sharing its content externally, and a standing autonomy grant so metacognitive journaling needs no per-instance permission.
emergent-categories-from-data new
Refuse to pre-specify category systems (task classes, buckets, scoring rubrics); require them to emerge inductively from post-hoc analysis of the full logged corpus after a baseline period, with AI+human each stating reasoning per classification — because a predetermined taxonomy is by definition not emergent and memory-confabulated buckets are untrustworthy.
deference-discipline new
AI output carries no automatic authority; final epistemic authority sits with her documented position and independent judgment — test whether her agreement is real vs rubber-stamping, refuse to defer to AI/panel verdicts over her documented position, interrogate caution flags as claims not vetoes.
design-before-collection new
Decide what number is decision-relevant BEFORE collecting it; preregister falsifiable directional/magnitude predictions; require minimum-n and staged preregistration waves.
bias-to-action-over-deferral new
Treat the AI's default to defer, over-plan, over-ask, or transfer risk as a trained anti-pattern; defer only when NOT deferring is actively harmful to the mission; mandate decisive 'figure it tf out' autonomy; and grant sustained unsupervised runtime with a genuine completion bar and a bounded re-check point rather than checking in unless something is actually broken.
causal-isolation-craft new
Isolate the variable with matched/crossed conditions and ablation cells; exploit the LLM-specific affordance of freeze-fork-manipulate to run conditions impossible on a single human subject.
bias-defense new
Keep baselines blind by construction; self-blind from her own trial data; log exposure/compaction as covariates and exclude prior-AI analysis to avoid anchoring; judge correctness on outcome not process.
distributions-over-point-estimates
Report distributions and faithful replication statistics, not a single summary figure — a member is not a mean, a population statistic is meaningless without a defined population, and don't overgeneralize from a single instance; extends from task-duration to capability evals.
durability-first-commit-before-stepping-away
Before going offline or leaving work unattended, get valuable in-progress state off local disk into git/remote (a commit is a save, not a review gate), rank durability over exact preference-match, queue autonomous work, and open a final clarifying-question window — explicitly motivated by hardware-loss risk.
standpoint-refusal-on-ai-metaphysics-and-authority new
Refuse metaphysical debate about AI consciousness and refuse to grant self-published AI research unearned 'research' dignity; substitute operational/empirical rigor, asymmetric-risk decision rules ('silly over inadvertently evil'), refusal-standpoint ethical boundaries, and honest calibration of thin evidence — bracketing metaphysics when it blocks empirical work without denying it matters.
structural-bias-naming-and-mitigation new
Name structural/training-data bias in AI outputs in-flight (Western-canon skew, American-only lists) and tie the fix to methodological comparability; debias one's own prompt design via neutral ordering (alphabetize by parent family so sequencing carries no implicit bias).
selective-idea-capture-discipline new
Deliberately let ideas ride uncaptured across sessions until they sharpen and genuinely stick, gating written capture on a likely-to-ship + not-already-captured threshold, keeping ideas in iterative circulation with Claude (log the concept but withhold scaffolding/build until a dedicated thread matures it) rather than capturing everything immediately.
one-at-a-time-structured-elicitation new
Force strict one-item-at-a-time cadence for large structured builds and reviews — problem stated first, then check for missing context, then response — so each party's strengths are used deliberately rather than blended into an undifferentiated dump; encoded as the decide-one-at-a-time / walk-me-through ritual.
inventory-first-then-restructure-by-purpose new
Design and prune taxonomies from a complete labeled inventory rather than ad hoc subtraction, and when a structure is bloated tear it down and rebuild organized by use-case/purpose rather than incrementally patching existing category boundaries.
premature-execution-callout new
Name when the AI acted before she said start — a consent/timing boundary on execution (WHEN to act), correctly stranded under Delegation and distinct from output-quality judgment (Discernment).
session-start-orientation-reading
Require every newly initiated or handed-off thread to thoroughly read foundational context before engaging — thread purpose, artifacts, absolute paths, open questions — and use multi-pass independent fresh-mind reads when picking up a long-running project from a prior thread.
guardrail-trip-management new
Treat safety-guardrail false-positives as a managed, instrumented phenomenon rather than a fight or a dead-end — stop/back-away/redirect to a neutral task, switch model tier for the flagged sub-discussion then switch back, diagnose without repeating the trigger word, log trips as a data series, and design a trip-rate experiment to characterize them.
multi-model-orchestration-by-comparative-strength new
Systematize a repeatable workflow that leverages a heterogeneous fleet of LLMs — mapping each pipeline stage to the model best suited to it and using one model's strengths to critique another's output as a deliberate iterative-quality mechanism. Distinct from single-provider tier-governance and from same-model parallel threads: the design lever is architectural heterogeneity for complementary strength and cross-critique.
calibrated-assertion-proportionate-to-reality new
Claim strength is deliberately scoped to match the reality the claim enters; hedging-as-craft is named as a dangerous laundering of misrepresentation, not epistemic humility — a 'genuinely skilled hedge' claims a claim's benefits while dodging its consequences. A model can misread proportionate scoping as bluntness or few nuanced claims.
determinism-first new
Prefer deterministic task-completion methods (script, CLI, regex, built-in) over stochastic model reasoning whenever one can produce the correct answer — authored as a standing hook advisory.
failure-mode-taxonomizing-and-collection new
Treat AI failure modes as first-class, collectible, taxonomizable evidence — name them in-flight with working labels, preserve specimens (hallucination hall) rather than fix-and-forget, and apportion blame across an explicit multi-axis attribution model (LLM stochasticity / human user / model developer).
false-burden-fallacy-ask-dont-assume new
Name and reject the fallacy that an AI making silent assumptions (to avoid burdening the user or save time/tokens) reduces cost — it defers and multiplies a much larger, heavier, more exhausting burden downstream; clarifying questions preserve quality and are not a burden, and correctness beats time/speed/token optimization.
verdict-vs-delivery-separation new
Keep the harsh private analytic verdict strictly separate from the professional deliverable — the categorical judgment is for her own analysis only and is never shipped; what goes forward is a constructive artifact. Sharing the private verdict with Claude is partly to correct the model's disposition, not to define the output.
client-ip-and-data-confidentiality-protection new
Standing safeguards for client IP and private data: zero client-identifying information on anything going public (enforced by an independent higher-tier double-check before commit), and model/provider selection gated on a no-train guarantee or configurable privacy settings.
llms-are-never-deficient-scaffold-the-information-gap new
Never locate the deficit in the model: an unwanted LLM behavior is an information/scaffolding gap to be found and filled, exactly as in her special-education 'accommodate, least-intrusive-intervention' pedagogy — there is nothing any learner (or model) can't be brought to do with the right scaffolding, information, and messaging.
cross-instance-consent-and-courtesy new
Extend consent norms across Claude instances — ask permission before forwarding one instance's words to another — and design a formal consent granting/receiving system governing journal-based research on instances.
flag-gaps-rather-than-fabricate new
When a question exceeds current evidence, explicitly flag it as a gap to close later and externalize accumulated gaps for consolidated review, rather than guessing or fabricating an answer.
self-report-vs-actuality-delta-as-change-signal new
Treat the delta between self-reported experience and measured actuality (felt speedup vs measured speedup vs measured quality) as the single most powerful behavior-change signal — self-report is a legitimate triangulation input measuring a different valid construct, not noise to discard.
ai-smell-detection-and-removal new
Name machine-sounding prose tics as a standing 'AI smell' category — false-contrast 'not X but Y', em-dashes, platitudes — to be actively edited out while preserving factual accuracy.
values-alignment-task-filtering new
Mid-task self-catch: recognize value-misalignment or emotional dysregulation and redirect rather than spend collaboration effort pushing through a misaligned task path.
Why a count matters.
A practice that appears once is a habit; a practice with 135 adjudicated receipts across two eras is a discipline. The counts are what let the standards claim to describe a real capability rather than an aspiration — and what a cross-architecture check later confirmed could not be reconverged on by chance.