← TERMS

In-Sample Pattern

Structure found in a fixed dataset — derivable, not a new observation

An in-sample pattern is regularity found inside a fixed dataset by a stated procedure — correlation, cluster, fitted curve, “rule” that holds on the rows you already have. Given dataset D and algorithm A, the pattern P is in principle P = f(D, A): deterministic relative to what was searched and how. It was latent in D, not transmitted from outside the channel.

Architecturally: L1 or L2 derivation from tier-one captures; epistemically tier 6–8 until independently tested — a hypothesis from data, not tier-zero fact about the Real. The data processing inequality applies: P cannot increase mutual information with the concrete L0 events that produced D beyond what D already encoded. Strict inequality is normal — the pattern generalizes and forgets individual witnesses, like any logical inference.

Spurious in-sample patterns arise when search is wide (many features, many thresholds): something will fit noise. The existence of a line on a chart is real; existence of the structure in the Real is not — multiple-comparisons failure, frozen collective subjectivity in ML, The Confident Deck in analytics.

Honest handling: record D version, A, P, and holdout / prospective plan; never promote in-sample fit to tier 2 or L0 without new captures.

Corpus stance

A1 — adopted on this site: Names the gap between “we discovered a pattern” (human surprise) and “processing created new ground truth” (it did not, relative to the sample already held).