LLM-Enabled MOE Tracking
Strong language models make automated effectiveness measurement plausible — if decisions are recorded on the log
AI can track measure-of-effectiveness signals to a useful degree — qualitative as well as quantitative — when organisational decisions live in decision records, not slides alone.
A2 — working stance T7 — hypothesis: advances in large language models make it plausible to track measure of effectiveness (MOE) signals automatically, at least to a useful degree — not only the measure of performance (MOP) proxies organisations already instrument.
Before strong LLMs, automation skewed MOP-heavy: throughput, latency, checklist completion, benchmark deltas — crisp numbers easy to wire into dashboards and comparators. Effectiveness — did the aim move? customer harm, strategic outcome, qualitative harm, narrative coherence with bind intent — was left to slow human review, sampling, or post-hoc story. Closing the loop on performance while effectiveness drifted was structurally easy (Goodhart’s Law, metric substitution, Cobra effect).
LLMs add qualitative analysis at scale: reading incident narratives, support threads, contest records, policy diffs, and outcome descriptions alongside quantitative series — asking whether the MOE class of question is moving, not only whether the MOP ticked green. That does not eliminate judgment or meta-loop humility; it may lower the cost of watching effectiveness signals that were previously too expensive to monitor consistently.
The dependency: decision records
Automated MOE trackingdoes not work on air. Models need grounded context: what was decided, under which assumptions, with what falsification plan, by whom — the bind the organisation claimed to be testing.
Consistent qualitative MOE analysis requires that consequential decisions affecting the organisation are recorded in decision records — e.g. according to the (A)DR Decision Log framework: named ownership, assumption tables, observables, triggers, tier honesty. Without that lineage, an LLM can summarise activity (MOP theatre) but cannot reliably ask whether the decision hypothesis is failing in the world (MOE).
This hypothesis is adriver behind research toward the epistemic harness and thinking exoskeleton: structure that lets human and machine capability act at the boundary honestly, then return effectiveness signal through closed-loop control — quantitative and qualitative — without treating model summary as committed fact.
What would strengthen or falsify
Would strengthen:
- MOE-class monitors that fire before MOP-only dashboards on real organisational cases, with logged binds as input
- Qualitative MOE reviews that agree often enough with expert judgment on held-out decision records to justify cadence
- Lower incident of green MOP / rotting MOE when LLM-assisted effectiveness sampling is wired to mandatory review
Would weaken or falsify:
- MOE automation that tracks proxy language in records but misses ground outcomes — Goodhart at the model layer
- Decision records too thin or absent for the model to anchor — harness without binds produces fluent MOP theatre
- Qualitative MOE noise so high that organisations rationally revert to MOP-only loops
T12: how much MOE automation is enough before trust in the model replaces contestability — open design question, not settled here.
Corpus stance
Working hypothesis A2 — motivates the public specification; not proven doctrine. Promotion requires evidence from built harness, recorded ADRs, and operational loops — likely developed in Enabling Intelligence and Operations threads.