AI production assurance

Connect signals, thresholds and decision authority

Build controls in which every alert corresponds to a risk, a population, an action and a person who is genuinely empowered.

On this page
Control

An indicator becomes useful when it changes conduct

A dashboard can describe the past without protecting the service. A control names the risk sought, signal, population, window, baseline, threshold, authority, action and evidence of closure. If authority or action is missing, it is an observation.

Signal

Test the quality of the observation before its value

The signal must arrive in time, be linked to the right version and be sensitive enough to the harm sought. We examine missing data, delays, selection bias, small volumes, noise and opportunities for circumvention.

  • Provenance — source, timestamp, version and transformation chain
  • Coverage — populations, exclusions and unobservable events
  • Delay — time between effect, detection and capacity to act
  • Robustness — noise, seasonality, manipulation and supplier dependence
  • Validation — reconciliation with human review or a second source
Threshold

Base the limit on risk, not ease of measurement

An explicit threshold includes the metric or condition, population, window, baseline, uncertainty and associated decision. Pre-alert, restriction and stop levels need not all be numerical: an unknown version or inability to interrupt the system can be a critical condition.

Cadence

Set frequency from the tolerable delay before harm

Cadence depends on how quickly risk emerges and how long action takes. Availability, version integrity or security may require continuous controls; quality analysis may run in batches; a change or incident triggers an event-driven review; contextual effects may require human judgement.

Authority

Write who can limit the system before the threshold is crossed

The decision matrix distinguishes who receives the alert, qualifies it, decides, executes and verifies recovery. Authority must have the competence, support and access required. An on-call role without the power to stop is not an authority.

Decision

Keep, restrict or stop within a precise scope

Keeping may include dated enhanced monitoring. Restriction identifies the affected function, population, channel, region, version or volume. Stopping leads to the defined safe state. Every decision records its rationale, duration, residual risk and exit test.

Minimum decision record
DecisionScope to writeExit condition
KeepVersion and accepted reservationsDefined reassessment or signal
RestrictCapability or population actually limitedReturn test and authority
StopSafe state and suspended operationsCorrection, regression test and authorised recovery
Validation

Exercise the route before the incident

We inject a synthetic signal or replay a known event to test detection, escalation, execution of restriction and preservation of evidence. The exercise also measures actual delays and human or technical single points of failure.

Primary sources

Sources this page relies on

Last documentary review: 6 September 2026.

Before you write to us

Frequently asked questions

Must a threshold always be numerical?

No. An unknown version, missing log or inability to stop the system can be a critical condition.

Is there a standard monitoring cadence?

No. It depends on the speed of risk, tolerable delay before harm, response time and applicable duties.

How is a control tested?

By injecting a signal or replaying a known case, then observing detection, escalation, decision, action and closure evidence.

Decision chain

Which alert currently triggers no action?

Present the signal, risk and current route. We identify the link preventing an executable decision.

Exercise the control