When the Commitment Boundary Needs Reinforcing
Severity, task nature, and who is affected
Part IV measured trustworthiness — six facets in The Facets Composed — composed into how much may I trust this for this use? When Trustworthiness Is Not Enough separated genuine from false uncertainty at the bind: when many are affected under real uncertainty, involvement and reinforced boundaries may be warranted; when the answer is knowable, votes are the wrong tool.
This article classifies when the commitment boundary must be reinforced — and how strongly — before you commit. Not every bind needs an investment committee. Not every bind may be a three-second click. The right calibration depends on the nature of the task, the outcome severity of error, how often it runs, and how many bear the consequences.
Using a certified collective with AI stress testing to choose a sort order is absurd. Using a naked LLM to approve a large commercial loan is dangerous — wrong-boundary automation at catastrophic severity. Using consensus where speed × time (or an invariant proof) settles the question is false uncertainty — democratic theatre, not judgment.
Navigate task nature against outcome severity in the interactive Decision Framework territory map; the Reference Card tab lists decision types, reinforcement mechanisms, and the axiomatic rules below.
Open Decision Framework map →Open Enterprise Decision map →
The three determining dimensions
Dimension 1 — Task nature
The most important question is epistemological, not technical: does a deterministic correct answer exist, and is it computationally feasible?
Deterministic and feasible (D-F): a correct answer can be computed exactly. Tax calculation, sorting, rules-based eligibility checks, string matching. The same inputs always produce the same correct output.
Deterministic in principle, infeasible in practice (D-I): a theoretically correct answer exists but computing it exactly is impractical at scale. Complex optimisation, large search spaces, computationally hard planning problems. Approximation or learned functions are necessary.
Genuinely probabilistic (P): no single deterministic correct answer exists. Medical diagnosis, fraud scoring, demand forecasting, content relevance, legal reasoning. The task involves judgment under uncertainty.
This distinction is not about complexity. A task can be extremely complex and still deterministic. Sorting a million records is complex; it is also deterministic. “Is this email spam?” sounds simple; it requires judgment over ambiguous evidence.
False uncertainty (When Trustworthiness Is Not Enough) is when stakeholders experience uncertainty because they lack access to a knowable answer — not because the task is P. Classifying D-F work as “we need a workshop and a vote” is process inflation: it installs collective machinery where transparent derivation suffices. Rule 1 and the D-F matrix rows exist partly to block that error.
Dimension 2 — Outcome severity
Catastrophic: errors cause death, serious injury, irreversible financial ruin, or major legal consequences. No recourse.
Significant: errors cause meaningful, recoverable harm. Real cost, but correctable.
Minor: errors are quickly caught and cheaply corrected. No lasting harm.
Negligible: errors have no meaningful consequence.
Dimension 3 — Frequency and variation
High-volume uniform: same task, similar inputs, high volume. Transaction fraud scoring, document classification.
High-volume variable: high volume but diverse inputs and contexts. Customer support, content moderation.
Occasional predictable: runs rarely but follows a known pattern. Annual regulatory filings.
Ad-hoc variable: one-off, unpredictable. Novel legal situations, creative briefs, research questions.
Who bears the consequences. The three dimensions above are necessary but not sufficient. When many parties are severely affected by a bind under genuine uncertainty, voice before commit and contestability paths are more likely to be proportionate — not because truth is voted on, but because cost allocation and standing are contested even after evidence is shared (When Trustworthiness Is Not Enough). When the answer is knowable (D-F), affected scale increases the obligation to publish evidence and derivation, not to average opinions.
The axiomatic rules — override all calibration
Four rules that override every cell in the matrix below.
Rule 1: Once a D-F algorithm exists, execution needs derivation — not judgment, human or AI.
Rule 1 governs execution, not design. Discovering or improving the algorithm — compressing a complex problem into a short, reproducible procedure — is genuinely intelligent work. In the information-theoretic frame, intelligence compresses: a compact program that solves a large problem is an act of understanding, not a substitute for it.
Once the procedure is found, published, and bound as the authorised method, each instance follows the chain by logical inference. That is the first class of work to automate — and the class that should not be performed by humans at scale, except for education, audit sampling, or improving the algorithm itself.
At execution, substituting judgment — model inference, workshop, vote — for derivation is a degradation. ML yields approximation where correctness exists; democratic aggregation yields majority ignorance under false uncertainty. If someone proposes an LLM, a workshop, or a vote for routine D-F execution, ask what specific failure of the deterministic algorithm that addresses. The answer is almost always performance, access, or politics — not that the instance required intelligence to solve.
Rule 2: Catastrophic severity always requires a human at the commitment boundary. No confidence level, no volume argument, no efficiency gain removes this requirement. A fraud model with 99.99% accuracy at catastrophic stakes is still wrong 1 in 10,000 times — and those errors are catastrophic by definition. The human is not just a quality check. It is the locus of accountability for decisions whose errors are irreversible.
Rule 3: Rubber-stamp review is not review. If humans approve model outputs without modification more than 99% of the time, the human is adding latency and the appearance of accountability without its substance. Track override rates. Act on them. A zero override rate is a red flag, not a sign of quality. See rubber-stamp approval.
Rule 4: Severity is set by the highest downstream use, not by the point of generation. Output from a task classified as negligible that feeds a catastrophic decision must be evaluated at catastrophic severity.
The calibration matrix
The table below is the static reference. Explore the same territory interactively — task nature against outcome severity, region chooser, and Reference Card.
Open Decision Framework map →Open Enterprise Decision map →
Abbreviations: H = mandatory human at commitment boundary · H? = human escalation path · VM = validated domain-specific model (not a general LLM) · Cal = calibrated with published calibration curve · Drift = continuous drift monitoring · RAG = retrieval with verified citations · Log = logging for audit · OR = override rate tracking
Three observations on reading the matrix:
The escalation path is not a safety net — it is a control. For every cell with H?, ask: who staffs this path, at what capacity, and is that capacity sufficient for the actual escalation volume? An unstaffed escalation path does not protect against errors. It creates a queue that eventually reaches a fatigued human making rapid decisions without adequate context.
The full-harness LLM row (P · Significant · High-vol variable) is where most LLM enterprise deployments actually live — and where most deployments fail to meet the requirements. Variable high-volume significant decisions are exactly the zone that requires RAG, citations, uncertainty disclosure, and a staffed escalation path. These are not optional. They are what make the LLM appropriate for that cell.
The volume pressure on catastrophic cells is a structural risk. As volume grows, the pressure to remove the mandatory human from catastrophic decisions grows with it. The framework’s response: volume does not change severity. A catastrophic error at high volume is catastrophic and high-frequency. The investment required is not to remove the human but to scale the human review capacity.
Where LLMs sit in the framework
A naked LLM — no retrieval grounding, no citations — is an L2-U-high artifact: multiple transformation steps from ground truth, with high and uncharacterised uncertainty. For automated decisions at the commitment boundary, it is prohibited above negligible severity — per the matrix above and the Artifact Classification grid. At negligible stakes it may support pre-bind human use (drafting, exploration); that is not licence for uncommitted write-through to L0. In practice, naked or under-harnessed models are deployed where they do not belong: P · significant · high-volume variable — the row that requires full harness, citations, and staffed escalation (LLMs at the Commitment Boundary).
RAG with verified citations over a curated, versioned knowledge base: approaches L1-U-med. Suitable for Class C operational support and the full-harness row in the matrix.
LLM output stored as L0 without a commitment boundary crossing contaminates the permanent record — the inference-as-fact failure mode. Automated bind requires recorded delegation (L0 policy + calibration evidence) or human/collective commit; never write-through.
The harness determines the classification, not the model. The same underlying model sits at L3-U-high (naked) or approaches L1-U-med (RAG, verified citations, human review). The investment in epistemic harness is the investment in trustworthiness — developed in LLMs at the Commitment Boundary. The next article covers how to raise quality at the bind — blind review, stress testing, calibration-weighted challenge. Legitimacy — certification, authorization, democracy — follows in When Social Stability Matters More Than Quality.
Who evaluates the matrix — and how
The matrix tells you what approach to use for a task once you know where it sits in the three dimensions. But who determines where a specific task sits? And by what process?
This is itself a decision — and the commitment boundary principles from Part III — The Moment of Commitment, But Commitment Is Not Enough — apply to it directly.
Task classification is not neutral. Classifying a hiring decision as “significant-recoverable” rather than “catastrophic-irreversible” carries consequences for who reviews it, by what mechanism, and with what quality gates. Classification under pressure to reduce friction will drift toward lower severity than is warranted. Someone must own the classification and be accountable for it.
AI can assist but not classify. An LLM can map a task description to candidate matrix cells, surface examples of similar tasks and their consequences, and flag cases where the classification seems inconsistent with precedent. This is useful L2 evidence for the human classifying the task. The classification decision itself is an L0 committed act — a named authority weight commits to treating this task as having this severity and this nature.
The classification deserves a quality mechanism proportionate to its consequences. Classifying a single internal routing task requires no ceremony. Classifying a new automated decision system that will process a million consequential decisions per day deserves the reinforcement mechanisms of the next article: certified authority, independent review, and explicit documentation of the assumptions made. The classification is not an afterthought. It is the decision that governs all subsequent decisions.
Classification should be revisited. A task initially classified as minor may accumulate downstream uses that make its actual severity higher. A domain initially classified as well-understood may become genuinely probabilistic as edge cases accumulate. The classification is an L0 event in the system’s governance record. When evidence suggests it is wrong, the correction is a new classification event that propagates re-classification downstream via supersession — not a silent update.
Calibration-weighted challenge mechanisms apply here too. For high-stakes classification — where misclassification would be severe — blind peer review and adversarial challenge apply. Meritocratic collective classification is a separate design choice (Meritocracy — The Sweet Spot).
Where territories overlap — collective zones, stress-test overlays, underlying decision types — use the Decision Framework map region chooser to select the layer you mean.