← TERMS

Experiment Architecture

How organisations generate evidence that updates intelligence — not just metrics

Experiment architecture is how an organisation structures trials so that intelligence can update honestly — from hypotheses from data through committed design, separate observation streams, and outcome comparison to tier promotion or demotion. It is not “we ran an A/B test” as a slide title. It is the socio-technical layout that makes evidence eligible to change models, policies, and knowledge.

Without it, “experimentation” produces dashboards and anecdotes. Intelligence stagnates or overfits — the same failure as in-sample patterns at organisational scale.


What experiment architecture requires

Pre-commitment. Hypothesis, primary metric, arms, sample rules, and decision thresholds recorded before outcomes are inspected — as L0 or strongly referenced L1 facts. Peeking and post-hoc metric shopping are tier laundering, not learning.

Separate streams. Each arm produces its own measure and commit captures — routing is an L0 configuration event (Immutable Infrastructure). Comparison is a query over committed history, not a mutable aggregate (The Moment of Commitment — when done right, “experiment” is architecture, not a parallel shadow system).

Outcome linkage. Predictions and interventions register against subject, time, and version so prospective outcomes can confirm or refute. Failed predictions demote tier; replicated success promotes it — the validation path from Synthesizing Knowledge.

Closed loop. Results feed model version commits, policy changes, or explicit HypothesisRefuted / HypothesisSupported events — not only a wiki page. Event Sourced Science applies to product and policy experiments, not only ML papers — countering publication over provenance and provenance deferred when misaligned incentives would otherwise reward headlines alone. Experiments are one layer of the epistemic harness — structure that lets intelligence update without tier laundering.


What it is not

  • Analytics without commitment — traffic splits nobody can audit
  • Peeking until significance — p-hacking dressed as agility
  • Unregistered holdouts — training on the test set organisationally
  • Rubber-stamp review of model output — human in the loop without calibration signal

Role in Enabling Intelligence

Synthesis proposes; experiment architecture is how the organisation learns whether proposals survive contact with the Real. It sits between knowledge (what we think we know) and calibrated intelligence (what we may rely on when acting) — the subject of the Enabling Intelligence series.

Corpus stance

A1 — adopted on this site: Planned core of Enabling Intelligence; full treatment in that series.