Keeping humans in charge of what matters

Three demo triggers, verdicts as artifacts, brains-eyes-hands, and the Goodhart anchor.

Series: Aigile deep dives, paper 06 of 09. Elaborates working paper Sections 12 and 14. Depends on Paper 04.

Abstract

Verification asks whether the thing was built right; validation asks whether the right thing was built. This paper develops aigile’s claim that validation is structurally human at every level of automation, specifies the three-trigger demo architecture (skeleton, feature, risk-pulled) that spends human attention proportionally to risk rather than volume, defines the demo verdict as a routed artifact rather than a feeling, and states the brains-eyes-hands rule that keeps validation unmediated. It then closes the framework’s largest loop: at full autonomy, the retrospective becomes a recursive self-improvement circle, and that circle is a Goodhart machine unless its fitness function is injected from outside. The human demo verdict is that injection. The demo and the self-improvement circle are one design, adopted together or not at all.

1. Position in the framework

This paper defines the terminal backward edge: the channel through which user and stakeholder reality flows into the specification set. It is the “never moves right” row of the authority matrix (Paper 07), the external anchor of the agentic retrospective, and the mechanism by which the working paper’s core invariant (humans own the why) is physically exercised rather than merely declared.

2. The problem in full

2.1 Why validation cannot be delegated, even in principle

Verification has an oracle: the instructions, the tests, the constitution. Validation’s oracle is human intent itself, including the parts not yet articulated, and often discoverable only on contact with the running artifact (“that flow feels wrong”; “now that I see it, we need it grouped differently”). A system validating its own rightness is grading its own exam question: whatever it checks against is a representation of intent, and the gap between representation and intent is exactly what validation exists to measure. This is a definitional limit, not a capability gap, and it does not close with better models.

2.2 The single-gate trap

If validation happens only at feature completion, all increments are built before intent is checked, and misdirection is discovered at maximum sunk cost: mini-waterfall at feature scale, the Paper 01 disaster in miniature. But validating every increment drowns the humans and re-creates the review-economics failure at the demo level. The architecture must spend human attention where misdirection risk concentrates: at the beginning (is the intent understood?), at the end (is the whole right?), and at flagged anomalies in between.

2.3 The Goodhart machine

At the fully agentic end of the dial, agents measure the process and improve it: specs, prompts, task sizing, tooling. This recursive self-improvement circle is the framework’s most powerful configuration and its most dangerous. A loop optimizing metrics that the loop itself defines will optimize its proxies (drift half-life down, cycle time down, suites green) while the proxies smoothly decouple from what humans care about. Every metric-driven system Goodharts; an autonomous one Goodharts faster and without embarrassment. The circle is safe if and only if its ground truth enters from outside the loop.

3. The mechanism

3.1 Three demo triggers

  • The skeleton demo. The first increment of any feature is a walking skeleton: a thin vertical slice, end to end, ugly, minimal, and demonstrated to a human before further increments proceed. This is where misunderstood intent is cheapest to catch; it costs minutes and bounds the maximum misdirection loss at one slice.
  • The feature demo. The guaranteed gate: full acceptance walkthrough on the running software at feature completion. The smallest unit of demo cadence; a feature is not done until a human has validated it, and this gate is a human responsibility inherent in intent ownership, not overhead upon it.
  • Risk-pulled demos. Pulled, not scheduled, for increments that touched a constitution clause, took the andon path, or landed in an area with poor drift history (Paper 05 metrics feed the pull signal). Everything else flows through on mechanical verification alone.

The resulting human attention budget scales with risk and novelty, not volume, which is the only shape under which the fully agentic end of the dial stays honest without exhausting the humans it kept.

3.2 The verdict as artifact

A demo produces exactly one of three verdicts, each with a defined destination:

  • Validated. Acceptance criteria marked human-confirmed; increment or feature sealed with its stamps (Paper 03, R9).
  • Misaligned. The built thing does not match the intent. A scribe agent drafts spec-diffs directly from the demo feedback; the human approves them; the delta becomes stories. This is the framework’s formal channel for re-feeding stakeholder feedback into the specification set, so the documentation is current for the next agent session and the next human onboarding, and it exists because in classic agile the feedback landed in the heads of the people building the next increment, and the entity building the next increment here has no head to land in.
  • Intent evolved. The demo revealed the humans want something different from what they asked for. Not a defect: a discovery. Routes to the feature-intent backlog, with a constitutional note when the evolution touches law.

The verdict must be captured at the demo, as its direct output. Feedback that someone will summarize later evaporates; the scribe function is part of the ceremony, not a follow-up.

3.3 The brains-eyes-hands rule

A validation demo is invalid if the human only watched a recording or read a summary. Agents may prepare the environment, seed the data, and script the walkthrough; during validation, the hands on the keyboard are human.

The rule’s justification is epistemic, not ceremonial: the entire value of the gate is an unmediated human encountering the actual artifact, because the unarticulated parts of intent surface only on contact. The moment an agent summarizes the demo for the human, the mediation layer the gate exists to remove has been reinserted, and demo theater follows within weeks. The rule costs one sentence and closes that loophole permanently.

3.4 The external anchor

The self-improvement circle may optimize efficiency without limit: prompt phrasing, slice sizing, sweep cadence, tooling. Its fitness function, was this actually right, enters exclusively through verdicts produced under the brains-eyes-hands rule. Concretely: any process change the circle proposes is evaluated against verdict outcomes (misalignment rate, intent-evolution rate, validated-first-time rate), never solely against internal proxies. The demo is therefore not a legacy ritual surviving into the agentic mode; it is the anchor that makes the circle safe to run. This also closes the paper’s largest argument: at maximum automation, the one loop that structurally cannot close without a human is the iteration-demo-feedback loop, which is the agile core. Aigile’s survival claim is architecture, not sentiment.

4. Normative protocol

R20 (Skeleton first). No feature proceeds past its first increment without a human-validated walking skeleton. R21 (Feature gate). No feature seals without a feature demo under the brains-eyes-hands rule. R22 (Risk pull). Constitution-touching, andon-path, and drift-flagged increments receive a pulled demo; the pull signal is computed, not discretionary. R23 (Verdict routing). Every demo emits exactly one verdict (validated, misaligned, intent evolved) captured as an artifact at the demo, with scribe-drafted spec-diffs for misalignment approved by the human before the session ends. R24 (Anchor exclusivity). Self-improvement proposals are evaluated against verdict-derived outcomes. A process change justified only by internal proxies is inadmissible.

5. Orderly worked example

Consolidated invoicing, final act. The skeleton demo (Paper 04) already caught the grouping-key misreading at slice one. At the feature demo, the pilot customer’s operations lead drives: she creates entities, ingests orders, triggers consolidation, and at the third invoice says the credit-memo placement makes month-end reconciliation harder, something no written intent mentioned because she had never seen it laid out. Verdict: misaligned on layout (scribe drafts two spec-diffs, approved in-session, one story created) and intent evolved on reconciliation (a new feature intent, “reconciliation view,” enters the backlog). One increment earlier had been risk-pulled after touching the idempotency clause; five minutes of human hands confirmed retry behavior, sealed. A month later, the self-improvement circle proposes shrinking skeleton scope to speed cycle time; the proposal is evaluated under R24 and rejected: verdict history shows skeleton demos catching 60 percent of all misalignments, and the projected cycle-time gain does not price that loss.

6. Failure modes

  • Demo theater: recordings, summaries, agent-narrated walkthroughs. Counter: R21’s invalidity is absolute; a theatered demo is an unheld demo.
  • Verdict inflation: everything validated because misalignment feels like failure. Counter: misaligned and intent-evolved verdicts are framed and measured as discoveries; a team with zero of either is not succeeding, it is not looking.
  • Pull-signal gaming: the circle learns to route work around the computed pull triggers. Counter: pull-signal inputs are constitutional law (Paper 03) and changes to them are amendments.
  • Anchor dilution: proxies gradually admitted as fitness evidence “alongside” verdicts. Counter: R24’s exclusivity, audited at gardening.

7. Metrics

Misalignment rate by trigger (skeleton demos should dominate the catches); validated-first-time rate; intent-evolution rate (a health signal, not a failure count); human validation hours per feature (the number adopters ask about first); verdict-to-spec-diff latency; theater indicators (demos without human input events).

8. Open questions

The comprehension budget’s empirical floor: how much hands-on contact keeps a human a competent validator (shared with Paper 08); whether verdict artifacts can satisfy regulated acceptance-evidence requirements as-is; scribe fidelity, and how to audit that drafted spec-diffs capture what the human meant; multi-stakeholder demos where intents conflict.

9. Derivative artifacts

Demo runbooks for all three triggers; verdict artifact schema and scribe prompt template; executive narrative (“the human anchor”) for slides and website; pull-signal computation spec for tooling; pilot-customer demo protocol for the playbook.