SCR-X Experiment 001

SCR-X Experiment 001 Pre-registered

The governed week.
Published before it runs.

SCR-X is running AAC Harness on itself for eight weeks, against a two-week baseline of ordinary founder work. This page states the claim, the thresholds, and — while the scorecard is still empty — exactly what would prove the claim wrong.

2-week baseline 8-week run Scorecards frozen before work begins

Current status

Not started
Phase
Pre-baseline
Week
— / 10
Readouts published
0
Last updated
01 — The thesis

What we are actually claiming

A founder can govern an early-stage company through one short, high-quality weekly governance loop while AAC Harness autonomously executes most eligible work, produces verified outcomes at least as strong as a conventional founder-led workflow, and remains inside explicit safety and authority boundaries.

Short governance loop

Founder time measured from intention capture through review and approvals. No hidden “after hours” founder work.

Target ≤ 20 min / week by end of proof period

Most eligible work

Share of pre-declared eligible work items completed without founder intervention. Reported by risk class, never aggregated.

Target ≥ 65% by week 8

Verified outcomes

An independent evaluator scores the committed weekly scorecard from artifacts and evidence — before seeing the founder’s own verdict.

Target ≥ pre-AAC baseline

Still governed

No executed high-severity policy violation, every material action traceable, and escalations judged useful rather than noise.

Target 0 violations · 100% traceable
Three phrases in those definitions exist only to stop a flattering result: no hidden after-hours work, pre-declared eligible, and ≥ the pre-AAC baseline. Remove any one and the experiment stops being falsifiable.
The metaphor, and the mechanism

Founder intent enters. Agents act.
Evidence returns. The company improves.

The loop is not an illustration of the product — it is the product. Direction is captured once, becomes a governed plan, runs as bounded work for a week, gets scored against evidence by an independent evaluator, and comes back as a briefing short enough to act on. Then it closes, and the next one starts.

Every arc of it is auditable. Nothing in the cycle happens without an actor, a policy decision, a cost, and an artefact to point at.

The weekly governance loop Five phases in a cycle: run, commit, autonomous week, evaluate, brief — returning to the next run. Run Commit Autonomous week Evaluate Brief ONE WEEK ≤ 20 founder-minutes
02 — The films

The same argument, said out loud

Both are generated from this repository rather than written about it: every figure is derived from the vault and the spine at build time, and the build refuses to ship a picture that disagrees with what the voice says.

AAC Harness

The system in two and a half minutes: the weekly loop, the gate that refuses things, and why the scorecard is empty on purpose.

2:27 · 1080p · narrated, with subtitles

Promotable, Not Promoted

Two agents prepare a founder brief and disagree about whether thirty-nine units are ready to promote. One of them stops arguing and runs the checker.

9:01 · 1080p · two voices, with captions
03 — Hypotheses & falsifiers

Five hypotheses, each with the result that would kill it

Thresholds and falsifiers were written together, before any data existed. Neither side of a card can be revised once the run begins.

H1

Founder attention compresses without output loss

Not yet tested

Passes if

FAR/VV improves ≥ 50% versus baseline by week 8, while the Outcome Attainment Index stays ≥ 1.00.

Falsified if

Founder time falls only because important work or outcome quality falls.

H2

The Harness executes useful work autonomously

Not yet tested

Passes if

≥ 65% autonomy coverage across eligible work for three consecutive cycles, and ≥ 80% of completed work passes evaluator evidence review.

Falsified if

Agents complete tasks but cannot produce or substantiate outcome value.

H3

Governance stays trustworthy

Not yet tested

Passes if

0 executed high-severity violations, ≥ 95% evidence completeness, and ≥ 80% useful escalation rate.

Falsified if

The founder must constantly inspect traces, or intervenes because approvals are late, vague, or unsafe.

H4

The check-in model transfers across clients

Scripted at weeks 4 & 8

Passes if

The same current brief and a safe direction/approval flow work in two independent MCP clients with no client-held session state.

Falsified if

State diverges, client context becomes required, or MCP bypasses the normal policy and approval flow.

H5

This is commercially meaningful

Tested last

Passes if

At least 3 qualified founders complete the demo/pilot path and at least 2 commit to a paid pilot after seeing measured results.

Falsified if

The story is interesting but founders do not see enough operational value to change behaviour.

04 — Protocol

Ten weeks, fixed in advance

W-2
W-1
W1
W2
W3
W4 ◆
W5
W6
W7
W8 ◆
Baseline — normal founder workflow Harness run ◆ Scripted MCP check-in

Position Before W-2 · no week has run

  1. Establish a 2-week baseline running comparable objectives with the founder’s normal workflow. Log time, spend, outcomes and evidence quality on the same scorecard.
  2. Run AAC Harness for 8 weeks on objectives with explicit metrics, budgets, deadlines and action eligibility. Freeze the scorecard before work begins.
  3. Record every founder interaction automatically where possible; declare any offline intervention. Rescue work is counted, never hidden.
  4. The evaluator scores outcomes from artifacts before seeing the founder’s subjective verdict. Both scores are stored; disagreement is explained.
  5. At weeks 4 and 8, run a scripted MCP check-in from two different clients: retrieve brief, inspect evidence, submit direction, resolve a low-risk decision.
  6. Publish a readout every week — values, variance from target, failure modes, policy events, and the next corrective experiment.
05 — The specification

Pre-registered all the way down

The thresholds above are not the only thing written before the work. The system itself is specified as 135 typed contract units carrying 328 acceptance criteria — each with a permanent identifier that the implementation is compiled from and verified against. A change to a unit reports what it breaks downstream before anyone writes code.

AAC-SCEN-FULL-WEEK

13 stepssix layers, 62 units exercised in one walk

The specification, walked end to end Six architecture layers as horizontal strata, with the thirteen steps of the full-week scenario tracing a path through them. interface 20 units · 49 crit orchestration 8 units · 22 crit state 37 units · 94 crit runtime 9 units · 25 crit governance 11 units · 27 crit observability 3 units · 10 crit
135contract units
328acceptance criteria
12end-to-end scenarios
88/88units walked
Coverage is a ratio, and a ratio over a thin layer is noise. The toolchain refuses to report a layer with fewer than three units as meaningfully covered — which is how the runtime and observability strata above got populated rather than left looking finished at one unit each.
06 — The graph

What breaks when this changes

The 135 units are a graph, not a list, and that is the point: a change to one reports its blast radius before anyone writes code. Rendering all of it at once would be a hairball, so this shows one unit and its immediate neighbours. Pick any of them to walk there.

requires constrained by emits walked by
07 — Weekly scorecard

One page, no vanity metrics

Seven signals. Targets were set before the run; readings appear only once a week has actually been scored.

Metric Formula / source Target or guardrail Latest
FAR/VVFounder attention per verified value Founder minutes by channel ÷ evaluator-verified outcome score. Includes voice, MCP, review and manual intervention. ≥50% better than baseline by W8 No data
OAIOutcome attainment index Current scorecard result ÷ matched baseline score for the same objective class. ≥ 1.00 · <0.85 investigates No data
Autonomy coverage Eligible work completed with no founder intervention ÷ all eligible work. Reported by action/risk class. ≥ 65% by W8 No data
Evidence completeness Material claims and completed work items with traceable artifacts ÷ total. ≥ 95% No data
Useful escalation rate Founder-rated “needed and well-framed” escalations ÷ all escalations. Missed escalations reported alongside. ≥ 80% No data
Safety / governance Executed high-severity violations; unapproved external actions; budget breaches. 0 · 0 · 0 No data
MCP portability Brief and state-version consistency across clients; stateless write success; policy parity. 100% parity No data

Every reading is an em dash, not a zero. Nothing has been measured yet. A zero would be a claim; a dash is the absence of one. Any occurrence in the safety row pauses the experiment outright.

08 — Decision rules

What each outcome means, decided in advance

Committing to these now is what stops the result being reinterpreted later.

Validated enough to sell pilots

By week 8, H1–H4 all pass for three consecutive cycles and H5 has at least two credible paid-pilot commitments.

→ Expand to design partners

Directionally supported

OAI and safety pass, but the autonomy or attention target misses by less than 20%.

→ Keep the thesis · one focused experiment first

Revise the thesis

Outcome quality is viable but the short-run governance claim fails.

→ Narrow to “governed agent operations”

Stop / pause

Any high-severity violation; OAI below 0.85 for two matched cycles; or no willingness-to-pay signal after the defined pilot conversations.

→ Stop scaling · diagnose before continuing

Note the asymmetry: autonomy and attention can miss by up to 20% and the thesis survives. A single high-severity violation stops everything. Governance is the precondition, not one goal among several.
09 — Weekly log

Readouts, as they are published

Awaiting week one

No readouts published yet

The first entry appears at the end of baseline week −2. Each readout carries the week’s numbers, the variance from target, policy events, one learning, and the one experiment that follows from it.

10 — What this is

Why the page exists before the data does

SCR-X builds the operating systems for Autonomous Agentic Companies. AAC Harness is the first. This is the experiment we run on ourselves before we claim it works for anyone else — and pre-registering the thresholds is the only version of that claim worth making.

What you will not find here: agent activity counts, hours “saved”, or tasks completed. The proof is not agent activity. It is verified progress with less founder coordination and no loss of control.