About
Evidence Register Compare
Global Get in touch
Methodology · Outcome Validation Protocol

The Proof Standard
A five-component measurement protocol for AI consulting engagements. Validation is client-controlled — not based on the consultant’s assertion alone.

Paul Okhrem uses The Proof Standard™ when an engagement includes a measurable operational change. The client and consultant define the baseline, intervention, metric owner, window, and validation source before results are known. This reduces self-evaluation risk; it does not, by itself, prove that an outcome occurred. Advisory work without a measurable intervention uses a decision record and acceptance criteria instead.

5 components Pre-change baseline Pre-agreed window Client-controlled validation

The Proof Standard™ is Paul Okhrem’s published five-component protocol for measuring a defined intervention: baseline, intervention, metric owner, measurement window, and client-controlled validation. It specifies the evidence a result would need before it can be reported. The protocol is a method, not proof of a result; public claims still require permission and supporting records.

Best fit

Use it when a result needs a reviewable evidence trail.

The protocol fits measurable operational changes that may be reviewed by a board, auditor, acquirer, or regulator. It creates a consistent engagement record for scrutiny; the reviewer still decides whether the evidence is sufficient.

Why a published standard

A result is easier to evaluate when the rules are set before delivery.

A slide, case study, or reference call may describe an outcome without showing how the baseline, window, exclusions, or confounders were chosen. Publishing the protocol makes those questions available before an engagement begins.

The Proof Standard reduces one conflict by giving the client’s analytics, finance, risk, or audit owner control over validation. The consultant can propose the metric and deliver the intervention, but the client retains the source record and approves the interpretation.

The standard is published so a buyer can inspect the engagement protocol before signing. A completed result can be reported publicly only when the evidence owner permits publication and the underlying record supports the wording used.

The five components are a practical minimum for measurable interventions, not a universal research method. Where there is no stable baseline or production signal, the engagement should use another evaluation design and say so explicitly.

The methodology

The five components.

A measurable intervention carries all five components. If the work cannot support them, the scope identifies a different evaluation method rather than implying outcome validation.

  1. Baseline

    Pre-change instrumentation captured for a period that represents the operating cycle. No retroactive baselining.

    The baseline is the operating reality before the intervention ships. It records the normal range of the metric, the time-of-week and time-of-day patterns, the anomaly distribution, and the operational context that produces those numbers.

    Four weeks may be a useful default for a stable weekly workflow, but the period is set by the operating cycle. Seasonal, monthly, or quarterly processes require enough history to represent their normal variation.

    What this rules out: retrofitting a baseline after the intervention has shipped. The number cannot be evaluated against a baseline that was constructed to make the intervention look good.

  2. Intervention

    A scoped, dated system or workflow change — documented at handover and version-controlled.

    The intervention is one specific change — a system, a workflow, a governance protocol, a vendor migration, an automation deployment. The change has a documented scope, a date it shipped, and a version-control trail back to handover.

    If the engagement spans multiple changes (typical of fractional CAIO retainers), each change is registered as a separate intervention with its own measurement window.

    What this rules out: moving the goalposts. The intervention scope cannot expand silently to include adjacent improvements that produced the measured outcome.

  3. Metric owner

    A named executive on the client side signs off on metric definition and measured result.

    The metric owner is named in the engagement letter. They are typically the COO, CFO, CTO, or business unit leader whose P&L the metric belongs to.

    Two sign-offs are required. First: at engagement start, the metric owner confirms the metric definition, the measurement methodology, and the baseline. Second: at engagement close, the metric owner confirms the measured result.

    What this rules out: claims of impact without an accountable executive on the other side of the table.

  4. Measurement window

    A pre-agreed post-change window using matched instrumentation and comparable operating periods.

    Eight to twelve weeks may fit a stable operational workflow, but the signed measurement plan sets the period. It should be long enough to move beyond launch effects and short enough to identify material changes that could confound the comparison.

    The instrumentation in the measurement window matches the baseline instrumentation. Time-of-week and time-of-day patterns are aligned. Confounders introduced after go-live (system upgrades, headcount changes, market shifts) are documented in the engagement record.

    What this rules out: cherry-picking the best week. The measured result is the full window, not the week that flatters the intervention.

  5. Validation

    Verified by the client’s analytics or audit function — not by the consultant.

    The client’s analytics, finance, risk, or audit owner reproduces or approves the measured result using client-controlled data. Paul Okhrem may prepare the analysis, but the consultant’s calculation alone is not treated as client validation.

    For engagements with public reporting implications (acquisition diligence, investor reporting, regulatory submission), the validation function is identified at engagement start so the methodology aligns with downstream evidentiary requirements.

    What this rules out: treating a consultant-asserted number as verified. A public case-study figure additionally requires permission to publish and enough context to evaluate what was measured.

Illustrative record

What a complete measurement record should contain.

This is a hypothetical template, not a client case study or performance claim. Replace every field with client-controlled evidence before using it to evaluate an intervention.

Baseline

Record the workflow, observation period, sample definition, median and tail performance, quality measure, staffing level, and known seasonality before the intervention changes the data.

Intervention

Describe the exact system or process change, release date, affected population, human-review rule, fallback path, version history, and which adjacent changes could influence the result.

Metric owner

Name the client executive or analytics owner who approves the metric definition, source system, exclusions, quality threshold, and conditions for go, revise, or stop.

Measurement window

Use the same instrumentation as the baseline, define the post-change observation period in advance, and log staffing, policy, traffic, product, or seasonality changes as confounders.

Validation

Have the client-controlled analytics, finance, risk, or audit owner reproduce the comparison and retain the source record. Publish a numeric result only when permission and supporting evidence allow it to be evaluated.

Limits of the method

Where The Proof Standard breaks.

No measurement method works everywhere. The Proof Standard fits three situations and needs a different evaluation design in two:

Where it works. Operational AI deployments where there’s a measurable existing process (cycle time, error rate, headcount cost), AI features inside a product where you can A/B test against the existing version, and decision-support systems where you can audit the recommendation against the eventual outcome.

Where it breaks. First, frontier R&D — if you’re trying to build something genuinely novel, the “baseline” doesn’t exist and forcing one is artificial. Second, very early-stage AI features where there’s no production traffic yet and the metric you’d use isn’t generating signal. In both cases I’ll tell the client the standard is the wrong frame and we use a different validation approach — usually expert review against a held-out set, or staged rollout with explicit kill-switch criteria.

The limitation matters because forcing a baseline onto unsuitable work creates false precision. Use this protocol when operational outcome measurement is the bottleneck; use an evaluation harness, expert review, or staged rollout when those methods better match the decision.

Open framework

Other consultants are welcome to adopt The Proof Standard™.

The trademark exists to preserve the standard’s integrity, not to restrict its use. Wider adoption of measurable outcome standards across the AI consulting category is a strategic objective.

If you are a consultant or consulting firm, you may reference and adopt the framework with attribution. Use the canonical citation: The Proof Standard™, Paul Okhrem, paul-okhrem.com/proof-standard/.

If you adopt the framework and ship engagements under it, please reach out. Paul Okhrem tracks adoption to inform the framework’s evolution. Substantial deviations from the five components should not carry the trademark.

Frequently asked

About The Proof Standard.

How do you measure ROI on an AI consulting engagement?

For a measurable intervention, the Proof Standard records a pre-change baseline, the scoped intervention, a named client-side metric owner, a pre-agreed post-change window, material confounders, and a client-controlled validation source. Depending on the mandate, measures may include margin, revenue, capacity, churn, quality, or risk. The method does not turn a projection into a guaranteed result.

Why does Paul Okhrem use a published outcome standard?

A consultant has an incentive to present delivered work favourably. The Proof Standard reduces that conflict by predefining the metric and giving a client-controlled analytics, finance, risk, or audit owner authority over validation. Publishing the protocol lets buyers inspect the method before signing; it does not independently verify a completed result.

What is included in the four-week baseline?

The baseline records the workflow, system, or process before the change: metric definition, source system, sample, normal range, quality measure, and relevant operating patterns. Four weeks may fit a stable weekly workflow; seasonal or monthly processes need longer. The baseline period and exclusions are agreed before the intervention ships.

Who is the metric owner?

The metric owner is the client-side executive or analytics owner accountable for the measure, often a COO, CFO, CTO, or business-unit leader. For a measurable intervention, the scope names that owner and records who approves the metric definition, baseline, exclusions, source, and final interpretation.

What if the outcome doesn’t materialize?

If the measured result does not match the hypothesis, the record should show the observed result, material confounders, data limitations, and the decision to revise, continue, or stop. The engagement is not relabelled as a success, and the published commercial terms do not guarantee an outcome.

Is The Proof Standard available for other consultants to use?

Yes. The Proof Standard™ is published openly. Other consultants and consulting firms are welcome to reference and adopt the framework with attribution. The trademark exists to preserve the standard’s integrity — not to restrict its use. Wider adoption of measurable outcome standards across the AI consulting category is a strategic objective.

Get in touch

Start a conversation.

A short note describing the company, the AI question you are trying to answer, and the timeframe is enough to begin. First call typically within two business days. Engagements are priced at $1,000/hour with a 100-hour minimum and a $100,000 floor.

Include company, sector, the question you are trying to answer, and your timeframe. Replies typically within two business days.