How to Measure the Realised Impact of a Digital Investment
Most organisations are diligent about projecting the value of an investment and negligent about measuring whether it arrived. The business case gets weeks of scrutiny; the post-implementation review gets a slide of screenshots six months later, if it happens at all. The result is predictable: nobody learns which assumptions were right, forecast quality never improves, and the next business case is greeted with the scepticism the last unmeasured one earned.
Measuring realised impact is not glamorous work, but it is the work that makes every future projection credible. This guide covers the six disciplines that separate honest measurement from retrospective storytelling: baselines, ramp windows, attribution, variance reporting, kill/scale decisions, and cadence.
Step 1: Baseline discipline — the measurement you cannot do later
The single most common measurement failure is not bad analysis after launch. It is a missing baseline before it. Once the new system is live, the "before" state is gone, and any baseline reconstructed from memory or partial exports will be challenged — rightly.
A usable baseline has four properties:
- Same units as the projection. If the business case promised "2,000 fewer calls per month", the baseline is calls per month from the telephony platform — not a proxy like "contact centre budget" that moves for a dozen other reasons.
- Long enough to show variation. A single month is an anecdote. Three to six months of pre-launch data lets you see seasonality, trend, and normal noise — which you need, because you will later have to distinguish your initiative's effect from that noise.
- From the system of record. Pull it from the source system, note the query and the date you pulled it. A baseline that cannot be re-derived is a baseline that can be disputed.
- Frozen. Write it down before go-live and do not revise it afterwards. A baseline adjusted after the fact — however good the reasons — looks like a moved goalpost.
If you did skip the baseline and are reading this after launch: reconstruct the best pre-period you can from system data, document exactly how it was built and what its weaknesses are, and label it as reconstructed. An honestly caveated baseline is recoverable; a silently reconstructed one is not.
Step 2: Ramp windows — do not measure the launch dip
No initiative performs at steady state on day one. Users are learning, processes are settling, edge cases are surfacing, and the team is still fixing things. Measure too early and you will conclude a good investment failed; wait indefinitely and you will never conclude anything.
The honest approach is to declare a ramp window in advance — before results start arriving, while you are still neutral. A practical structure:
- Weeks 0–4 (stabilisation): monitor for breakage only. No impact claims in either direction.
- The ramp (typically months 2–4, longer for behaviour-change initiatives): track the trend toward steady state. Adoption curves matter more than absolute numbers here.
- The measurement window (after ramp): a clean multi-month period at steady state. This is the period your variance report is built on.
Two rules keep this honest. First, declare the windows before you see results — a measurement window chosen after peeking at the data is indistinguishable from cherry-picking. Second, expect ramps to run longer than the pitch implied. Even practitioners' expectations have drifted this way: Deloitte's 2022 Intelligent Automation Survey found reported payback on automation pilots lengthening from 16 to 22 months between surveys. Building a realistic ramp into both the projection and the measurement plan is not pessimism; it is calibration.
Step 3: Attribution honesty — what else changed?
Your metric moved. The uncomfortable question is whether your initiative moved it. Between baseline and measurement window, other things changed too: pricing, marketing spend, seasonality, headcount, a competitor's stumble, the wider economy. Attribution honesty means asking "what else changed?" before claiming credit — because your CFO will ask it, and asking it first is the difference between analysis and advocacy.
A practical attribution routine:
- List concurrent changes. Every launch, campaign, price change, policy change, and staffing change that overlapped your measurement window. Five minutes with the marketing and ops calendars is usually enough.
- Use a comparison where one exists. A region, segment, product line, or cohort that did not get the initiative is the cheapest control you will ever find. If the treated group improved and the untreated group did not, your attribution is strong. If both improved equally, be suspicious of your own case.
- Check the mechanism, not just the outcome. If self-service was meant to reduce calls, you should see self-service sessions rising and calls falling, with the volumes roughly consistent. An outcome moving without its mechanism is a red flag that something else did the work. Be especially wary of assuming attempted usage equals success: Gartner's 2024 consumer survey found only 14% of customer service issues are fully resolved in self-service, even though most customers try it — usage dashboards flatter; resolution data tells the truth.
- Watch for quality offsets. Cost savings that quietly degrade quality have a way of returning as costs elsewhere — repeat contacts, complaints, churn. The cautionary example is public: Klarna reported its AI assistant handling two-thirds of service chats in its first month, a genuine best-case containment figure — and later rehired human agents after quality complaints. Measure the offsetting metrics (satisfaction, repeat-contact rate, escalations) alongside the headline saving, or your impact number is only half a number.
You will rarely achieve perfect attribution outside a controlled experiment, and you do not need to. What you need is to have looked, to say what you found, and to state your claimed share of the movement with a straight face.
Step 4: Variance reporting — impact against the original projection
Realised impact only means something against the projection that justified the investment. Reporting "we saved £180,000" is a headline; reporting "we projected £220,000 and realised £180,000 — here is the variance, line by line" is accountability, and it is what makes your next projection worth believing.
The mechanics:
- Report in the same lines as the projection. If the business case was volume × unit cost × improvement rate, report the realised value of each factor separately. This shows where the variance came from: perhaps volumes matched but the improvement rate ran at 80% of assumption. That is far more useful than a single blended miss.
- Treat variance as information, not verdict. An initiative that delivered 80% of projection is usually a success with a calibration lesson attached. The lesson — "we over-assumed adoption by a fifth" — feeds directly into the next business case's conservative scenario.
- Use a tighter band than the forecast. A projection deals in forecast uncertainty and deserves a wide scenario spread. A measurement deals in measurement noise and deserves a narrow one — our Impact reports run a ±10% band, framed explicitly as measurement variance rather than forecast uncertainty. Reporting realised impact with forecast-sized error bars reads as hedging.
- Report misses at the same volume as wins. The fastest way to destroy measurement credibility is to publish only favourable reviews. The second-fastest is to quietly re-baseline until the miss disappears.
This loop is built into the product: an Impact report links back to its original Projection, and the variance between projected and realised value is computed and displayed line by line, using the same formula and the same assumptions ledger. However you implement it, the principle stands — impact is only meaningful relative to the promise.
Step 5: When to call it — kill and scale criteria
Measurement exists to drive decisions, and post-implementation there are only three: kill, fix, or scale. The healthiest version of this decision is one whose criteria were written before launch, because criteria written afterwards bend towards whatever the sponsor already wants.
A workable decision rule set:
- Kill if the realised benefit is below the conservative scenario after the declared ramp window, and the sensitivity-driving inputs (adoption, resolution rate, whatever your model said mattered most) show no improving trend. Sunk costs are sunk; the only question is whether the next pound spent earns its keep.
- Fix if the mechanism is working but a specific, nameable input is underperforming — adoption is low but rising, or one segment resolves well while another does not. Set a second checkpoint and a threshold for it. "Fix" without a deadline is "kill" with extra spending.
- Scale if realised impact meets or beats the conservative case with attribution you would defend in front of finance. Scaling is a new business case — but one with the rare advantage of measured, not benchmarked, inputs. A projection built on your own realised unit economics is the strongest projection you will ever write.
The discipline that makes all three possible is the one from Step 2: pre-declared checkpoints. A decision date in the calendar, agreed at approval time, is what stops an underperforming initiative drifting on for a year because no meeting ever forced the question.
Step 6: Build the measurement cadence
One-off reviews decay; cadences persist. The aim is a lightweight rhythm that keeps measurement running without becoming a reporting industry:
- Weekly during ramp: the driving metrics only — adoption, usage, error rates. Purpose: catch breakage and track the ramp, not to declare victory.
- Monthly at steady state: the variance view — realised vs projected, same lines as the business case. Fifteen minutes if the data flows are automated; set them up once during implementation, when engineering attention is already on the systems involved.
- At each declared checkpoint: the kill/fix/scale decision, taken against the pre-agreed criteria, with the outcome recorded.
- Annually: a calibration review across all measured initiatives. Which assumptions ran hot, which ran cold, and by how much? This is where forecast quality compounds — a team that knows its own optimism bias can correct for it in the next case.
Keep the cadence proportionate. The point of measurement is better decisions per hour of effort, and a measurement programme that costs more than the insight it produces has failed its own test.
Close the loop
The projection-to-impact loop is the whole game: frame the case in measurable units, freeze the baseline, declare the windows, measure honestly, report variance against the promise, and let each measured result sharpen the next forecast. Teams that run this loop get faster approvals over time, because their projections carry a track record.
LeadersToolset is built around that loop — a Projection report to make the case, an Impact report linked back to it to measure what arrived, and the variance computed line by line on an itemised assumptions ledger. The methodology documents exactly how, and the sample report shows the output. Your first calculator is free with unlimited re-runs — start the loop with a projection.