← GENESIS
Part IV — The Facets of Trustworthiness · Article 18

What Information Theory Says

The formal underpinning of the facets

The motivators and the synthesis built the facets from intuition and examples. This article shows that the architecture is not merely intuitive — it is formally derived. Information theory, developed by Claude Shannon in the late 1940s (A Mathematical Theory of Communication), provides the mathematical foundation for why atomicity and uncertainty behave as they do and why their composition rules are what they claim to be.

This article is for readers who want to understand why the framework is right, not just that it is useful. It can be read independently or skipped without loss of practical understanding.


Shannon entropy — the formal measure of uncertainty

Shannon entropy is the foundational measure of information. For a random variable X with possible outcomes, entropy H(X) measures how uncertain the outcome is — equivalently, how much information is gained when the outcome is observed.

A deterministic event has zero entropy — the outcome is certain before it is observed, so no information is gained by observing it. A fair coin flip has maximum entropy for a binary variable. A fraud model’s output on a given transaction has entropy somewhere in between — the outcome is uncertain, but not uniformly so.

This connects directly to the facets — but two kinds of “lossless” must not be conflated.

An L0 decision event has zero entropy relative to the act of deciding — the decision is what the decision is; regress terminates there. An L2 probabilistic model output has positive entropy relative to its inputs — the outcome is genuinely uncertain even when inputs are complete.

Logical inference is different again. It is deterministic (η = 1 for probabilistic loss — the same premises always yield the same conclusion). It is not L0. Starting from concrete L0 facts, inference produces a generalized statement — a sum, a rule conclusion, a tier-2 restatement — that forgets which specific atoms supported it unless the artifact contract carries them explicitly. That is strict mutual-information loss toward the original L0 set via the data processing inequality, even when no stochastic step intervenes. Such conclusions sit at L1 (one characterized step from L0) or L2 (multi-step or multi-stream), not at L0.


The data processing inequality — the formal proof of the hierarchy

The most important result for the framework is the data processing inequality. For any Markov chain of processing steps X → Y → Z, the mutual information satisfies:

I(X ; Z) ≤ I(X ; Y)

In plain language: processing cannot increase the mutual information between an output and the original source. It can only preserve or decrease it.

This is the formal proof that the L0–L3 hierarchy is not arbitrary. Every transformation step in a processing chain can only lose information about ground truth — never gain it — which is why epistemic uncertainty must be characterised at each level. A L2 artifact has, at most, the same mutual information with ground truth as the L1 artifacts it was derived from. In practice, it has less.

The inequality is strict whenever information is genuinely lost — every non-invertible step (including logical inference that generalizes above concrete L0 facts), every η < 1 probabilistic step, every aggregation that discards detail, every inconsistent cross-stream join.

The practical consequence: the further an artifact is from ground truth in processing steps, the less it can tell you about ground truth — regardless of how sophisticated the processing is. A sophisticated L3 dashboard tells you less about the underlying ground truth than a simple L0 event, because the processing chain between them has accumulated information loss. This is the formal content of Why Atomicity Matters.


Patterns from data — three senses of “new information”

Mining raw captures for regularities — a correlation, a fitted law, a “speed from position” formula — feels like discovering something the world did not tell you before. Three distinctions keep that honest.

1. Relative to the sample already held. An in-sample pattern is P = f(dataset, method). It is logical inference at scale: L1/L2, typically tier 6–8 as a hypothesis from data. The data processing inequality still applies to past captures — processing the same fixed rows cannot increase mutual information with the concrete L0 events that produced them. The pattern compresses the sample; it does not add a new observation of the Real.

2. Relative to a reader who lacked the summary. Writing the formula on paper and handing it to someone else is new information for them — on the communication channel from author to reader. Shannon’s theory covers that transfer explicitly. That is not the same claim as “re-analysing past data created new ground truth.” One link is transmission to an uninformed agent; the other is derivation from captures already in the log.

3. Relative to future outcomes. A promoted hypothesis — kinematics, demand forecast, fraud score — can predict locations, sales, or flags you have not yet measured. That predictive power is why hypotheses matter. Architecturally, each forecast is still an L2 projection from model plus inputs; future ground truth enters only when new measure or commit captures arrive and are compared to what was predicted. Past mining did not substitute for those captures. Validation raises tier; systematic misprediction demotes it. Until then, treat the rule as useful compression with explicit uncertainty, not as tier-zero fact about every future instant.

The framework therefore separates derivation (cannot invent past Real from past data), communication (can inform agents who lacked the derivative), and prediction (can guide action until the world confirms or refutes on fresh evidence).

Information vs knowledge

The better frame for pattern discovery is often not “we gained Shannon information from the sample” but “we gained knowledge from data.” Shannon does not theorize knowledge — he theorizes signals and uncertainty. Meaning, justification, and action were outside his scope; mutual information is defined over random variables, not over “understanding” or “what an organisation knows.”

Knowledge synthesis is the ascent from data (events in the log) through information (structured, provenance-carrying statements) to knowledge (rules, models, ontologies — synthesized and usable). A derived formula has lower mutual information with the raw L0 rows than those rows had with themselves; it can still raise knowledge — compress experience into something teachable, predictable, and actionable. Tier tracks how far that knowledge is justified; atomicity tracks how far it sits from atoms; Shannon tracks how much uncertainty each step preserves or destroys. Three facets, not one word overloaded.

The conclusion for pattern discovery: identifying structure compresses Shannon information relative to raw captures — detail is discarded on purpose — and creates knowledge in the organisational sense. That is how we gain knowledge from data, not by pretending re-analysis invented new past observations. Statistical fitting, analytic derivation, and knowledge-graph organisation are all synthesis methods; they differ in technique, not in this basic trade.

A later Synthesizing Knowledge series will take up those methods in full — discovering patterns with statistical tools, building models analytically, and organizing explicit knowledge into graphs and ontologies — always with tier, provenance, and the causal chain from events intact. Until then, the knowledge thread marks where that work lives alongside data and information.

A3 — open mapping Genesis Part III may eventually align atomicity with the capture stack more tightly: L0 with data (committed events and designated observations); L1 with information (structured, single-stream aggregates with explicit references); L2 — often — with knowledge (synthesized models and graphs ready to act on); L3 with presentation (knowledge adjusted and rendered for different audiences). The facets and the stack are not identical — commitment status and tier cut across both — but the rhyme is intentional and will be developed in that series rather than settled here.


Channel capacity — the formal version of epistemic uncertainty

The data processing inequality tells you that information can only be lost. Channel capacity tells you how much can be preserved at each step.

Shannon’s channel capacity C is the maximum mutual information that can be transmitted through a communication channel — a processing step. A deterministic transformation has infinite capacity (or more precisely, its capacity is bounded only by the entropy of the input — it loses nothing). A noisy channel has limited capacity determined by its signal-to-noise ratio.

The epistemic uncertainty facet — Low, Medium, High — is an informal categorisation of channel capacity. A deterministic, low-latency, single-stream transformation has high channel capacity. A probabilistic model with uncalibrated confidence, applied to data with unknown quality, over multiple streams with no consistency guarantees, has low channel capacity.

The channel capacity framework also explains why the uncertainty threshold is contractual. Two channels can have the same capacity in different units — what counts as “low uncertainty” for a payment system and for a monthly audit are different thresholds of the same underlying property.


KL divergence — the formal measure of observation quality

Kullback–Leibler divergence (Kullback & Leibler, 1951) measures how much one probability distribution diverges from another. Applied to observation quality:

Quality(sensor) = 1 - D_KL(P(observation) || P(ground_truth))

A perfect sensor — whose output distribution exactly matches the distribution of the underlying phenomenon — has KL divergence of zero. Its observations are perfectly faithful. An uncalibrated or biased sensor has non-zero KL divergence — its output distribution differs systematically from the underlying phenomenon.

This formalises the distinction between pragmatic observations of different qualities. A primary standard measurement (calibrated against a physical constant, traceable) has near-zero KL divergence. An uncalibrated instrument has unknown KL divergence. The requirement that observation quality be “characterised and published” is, formally, the requirement that KL divergence be estimated and disclosed.

The recursive problem — that estimates of quality are themselves subject to quality — is formalised here as the problem of estimating KL divergence from finite samples. The estimate is itself a random variable with uncertainty. Bayesian hierarchical models handle this by placing a prior over possible KL divergences, updating with calibration data. The regress terminates at some level where the prior is treated as fixed by convention — which is itself a pragmatic tier-one designation.


The Cramér-Rao bound — the fundamental limit of measurement

The Cramér–Rao bound establishes that no unbiased estimator can have variance lower than the inverse of the Fisher information:

Var(estimator) ≥ 1 / I_Fisher(parameter)

In plain language: there is a fundamental lower bound on how precisely a physical quantity can be estimated from a measurement, determined by the Fisher information content of that measurement. No amount of clever processing of the measurement can beat this bound. The bound is tight — achievable in the limit by maximum likelihood estimation.

Applied to the framework: a sensor with low Fisher information about the quantity it measures cannot be improved by sophisticated downstream processing. The processing can approach the Cramér-Rao bound but cannot exceed it. The limit on trustworthiness is set at the point of observation, not downstream.

This is why pragmatic observation quality determines an upper bound on the trustworthiness of every artifact derived from it. No amount of sophisticated modelling or aggregation can recover information that was never captured. The quality ceiling is set by the observation.


The multiplicative distance formula

Putting these results together: the mutual information between an artifact and ground truth compounds across processing steps in a multiplicative way.

For a chain of processing steps with information efficiencies η₁, η₂, …, ηₙ (where η=1 means lossless and η=0 means total information loss):

I(Output ; Ground_Truth) ≤ I(Input ; Ground_Truth) × η₁ × η₂ × … × ηₙ

The distance from ground truth is inversely proportional to this product. Each step multiplies the remaining mutual information by its efficiency η. This is why long processing chains with many probabilistic steps produce outputs that are very weakly informative about the original ground truth — not as a design failure, but as a mathematical inevitability. This is the formal content of the “uncertainty compounds multiplicatively” composition rule in The Facets Composed.

A deterministic step has η = 1 for probabilistic loss — it introduces no stochastic uncertainty. That does not mean it adds no atomicity distance or preserves full mutual information with concrete L0 ground truth: logical inference and other generalizing transforms are typically strictly lossy under the data processing inequality even at η = 1. A probabilistic model step has η < 1 — it adds irreversible stochastic loss on top. A cross-stream join without consistency guarantees has η < 1 — the temporal inconsistency between streams reduces the mutual information of the combined artifact with any single ground truth state.

The practical consequence: every probabilistic or lossy step in a processing chain is permanently reducing the trustworthiness of the artifact relative to ground truth. Sophisticated processing can be valuable — a well-trained model may extract highly relevant signals from complex inputs — but it cannot violate the data processing inequality. It can only approach the ceiling set by the Fisher information of the original observations.


The recursion and its management

The framework’s recursive property — uncertainty about uncertainty, quality of quality measures — has a formal treatment in Bayesian statistics.

The standard Bayesian framework handles second-order uncertainty (uncertainty about the probability distribution itself) through hierarchical models: a prior over probability distributions, updated by calibration data. Third-order uncertainty (uncertainty about the prior) is handled by a hyper-prior. The recursion is real but manageable: in practice, the third and higher orders contribute negligibly to uncertainty, and the hierarchy terminates at a level where the prior is treated as fixed by convention.

This is exactly the pragmatic tier-one designation: the system designates a level of granularity as its input boundary and treats uncertainty at that level as the starting point rather than asking what produced it. The choice of that boundary — and the prior placed on uncertainty at that level — is itself a design decision, and a consequential one.


What the framework is, formally

The atomicity hierarchy (L0–L3) is an informal but structurally correct discretisation of the mutual information chain established by the data processing inequality.

The epistemic uncertainty facet (Low/Medium/High) is an informal but structurally correct categorisation of channel capacity per step.

The consumer class framework (A–D) is an informal but structurally correct specification of the minimum mutual information with ground truth required for each category of decision.

The dependency matrix is the matching of required mutual information (consumer class) against provided mutual information (source level and uncertainty). The ✗ cells are cells where the provided mutual information is below the required minimum. This is not a preference — it is a formal mismatch.

Information theory does not invent the framework. It confirms it. The intuitions that produced the facets were correct because they were tracking real properties of information processing. The formalism makes those properties precise and makes the conclusions derivable rather than asserted.


The practical takeaway

For practitioners who do not need the formalism, the practical takeaways are:

Processing past captures can only preserve or lose mutual information with those captures, never gain it. Every transformation step is a potential point of loss. Telling an uninformed agent a derived summary is still genuine information on the communication channel. Predicting the future from a validated hypothesis is valuable — and is tested only when new captures arrive.

The quality ceiling is set at the point of observation. A poor-quality sensor or uncalibrated instrument sets an upper bound on the trustworthiness of everything derived from it. No downstream processing can breach this ceiling.

Probabilistic processing introduces irreversible stochastic loss. Unlike deterministic derivation (which can be exact given premises but still generalizes away concrete L0 detail), probabilistic processing adds η < 1 loss that cannot be recovered from the output alone.

In-sample patterns are hypotheses, not discoveries of new ground truth. Fit on the training set is tier 6–8 until holdout, prospective test, or replication promotes the claim.

Pattern discovery can increase knowledge while Shannon information relative to raw captures decreases. Knowledge synthesis and Shannon limits answer different questions — do not collapse them.

Long chains with many probabilistic steps compound information loss multiplicatively. The trust in a highly derived artifact may be very low relative to ground truth — not because any single step failed, but because many small information losses compounded.

Uncertainty must be characterised to be managed. Unknown uncertainty is worse than known uncertainty. The requirement to characterise and publish uncertainty properties is not bureaucratic — it is the minimum necessary for consumers to make calibrated decisions.

Part IV closes here. Part V — Uncertainty at the Commitment Boundary asks what happens when epistemic trustworthiness is necessary but not sufficient — beginning with When Trustworthiness Is Not Enough.