AI portfolio arbitration

Build criteria and anchors that withstand scrutiny

A grid becomes defensible when every value denotes an observable state, the options are comparable and two readers can reconstruct the same reasoning.

On this page
Frame

Start with a decision, not a list of criteria

The same option may be useful for learning and unsuitable for industrialisation. Before measurement, the mandate fixes the decision, objective, horizon, scarce resources, affected population and baseline option.

Question
A decision phrased so it can actually be resolved
Baseline
What happens if no new option is commissioned
Options
Complete candidates that are comparable and distinct
Authority
The person or body empowered to accept trade-offs
Criteria

Retain only what can genuinely change the choice

A criterion enters the model only if it describes a useful difference between options. We search for missing effects, overlaps and causal dependencies before writing a scale.

  • Relevance — the measured effect matters to the objective or an affected party
  • Discrimination — the options can genuinely differ on this criterion
  • Measurability — data, observation or a structured judgement is available
  • Non-redundancy — the same effect is not valued under two headings
  • Compensability — a better result elsewhere may legitimately offset a loss here
Scales

Move from raw measure to interpretable value

A value function translates performance into the language of the decision. It may be linear, contain a threshold or reflect diminishing returns; its form must be justified, not selected to improve a ranking.

Value-function control
ElementControl questionRecord
Raw measureWhat is observed before transformation?Unit, source, population and date
Low anchorWhat is the least desirable state relevant to this choice?Value 0 and rationale
High anchorWhat is the most desirable state relevant to this choice?Value 100 and rationale
ShapeDoes one raw unit create the same value everywhere?Function, breakpoints and assumptions
Admissibility

Check that a total is meaningful before displaying it

The additive model assumes that preferences on one criterion remain interpretable regardless of the level taken by the others. If the effect of one criterion structurally depends on another, aggregation can tell a false story.

Same question
Identical objective, horizon and baseline for every option
Valid scales
Anchors and intermediate values reviewed before consolidation
Independence
No criterion replicates or mechanically conditions another
Accepted trade-offs
Residual compensation accepted by the authority
Calibration

Surface divergence before the live portfolio

Two assessors independently appraise calibration options. The comparison improves the scale, clarifies the evidence and separates disagreement about facts from disagreement about preferences.

  • Score separately — retain the initial value, evidence and rationale
  • Compare — locate the anchor, datum or scope producing the difference
  • Correct — rewrite a scale that permits incompatible interpretations
  • Arbitrate — retain the initial values and the selected value with its authority
  • Replay — verify the corrected scale on a second calibration option
Weighting

Compare changes of state, never words

Swing weighting compares the value of moving from the low to the high anchor on each criterion, over the range actually observed between options. An important-sounding criterion may carry little weight if all candidates are equivalent on it.

StageDecision requiredControl
ReferenceWhich swing creates the most value?Assign it 100 comparison points
ComparisonWhat is every other swing worth against the first?Document pairwise judgements
NormalisationWhat share does each swing represent?Weights sum to 1
ScenariosWhich disagreements must remain visible?Test alternative weight sets separately
Output

Deliver a grid a third party can read without an interpreter

The calibration record assembles the frame, criteria tree, raw measures, value functions, calibration options, initial appraisals, calibration decisions and weight scenarios. Every object carries a version and owner.

Primary sources

Sources this page relies on

Last documentary review: 7 September 2026.

Before you write to us

Frequently asked questions

Why is a 0-to-100 scale not enough?

Because the numbers remain arbitrary until 0, 100 and intermediate states are tied to observable performance and a justified value function.

How is the same benefit prevented from being counted twice?

We connect each criterion to a distinct effect, examine causal dependencies and remove or restructure criteria that describe the same change.

What does a weight actually measure?

The relative value of moving between the low and high anchors over the range considered. It does not measure the abstract importance of the criterion name.

Criteria and anchors

Do your criteria describe facts or impressions?

Present the decision, options and current grid. We identify overlaps, ambiguous anchors and missing evidence.

Review the grid