Skip to content
Between the Events

Index  ·  Change

Measuring Whether It Worked

The same log before and after is an unusually clean comparison. The traps that still apply, and why null results must be reported.

Procedure

Process mining's main advantage over other improvement methods is that the measurement instrument does not change between before and after. That advantage is easily thrown away.

Keeping the comparison clean

Same query, literally, saved and re-run.

Same case boundary.

Same exclusion rules, which if changed invalidate the comparison silently.

Same activity mapping version, since a naming change alters everything.

Same period length, and the same season where seasonality exists.

Write all five down when the baseline is taken, because the person re-running it may not be you.

The baseline

Long enough to contain the normal variation, which is several weeks in most processes.

Not a single week, and not the week before the intervention, which is frequently unusual.

Documented: period, query, boundary, exclusions, mapping version, and the figures with their spread.

The traps that still apply

Regression to the mean. Intervening after a bad period, then observing an improvement that would have happened anyway.

Seasonality, comparing a quiet month with a busy one.

Volume effects, where cycle time falls because volume fell.

A concurrent change — a system release, a staffing change, a policy update — that explains the difference.

Selection, where the filter excludes incomplete cases and the improvement is that more cases are now incomplete.

The controls available

A comparison segment running unchanged: another region, another case type, another team. The difference in the differences removes seasonality and volume effects at once.

The historical series, which establishes what would have been expected.

Choose the control before the intervention, not afterwards, which prevents selecting a flattering comparison.

Reporting it

The whole series, not two numbers, so the reader can see whether the change coincided with the intervention or preceded it.

Median and ninetieth percentile, since they frequently move differently.

Volume alongside, so a fall in cycle time can be distinguished from a fall in demand.

The control, if there is one.

What else changed in the period, honestly.

The null result

A change that produced nothing must be reported.

Say why you think so, which is usually that the finding was not the constraint.

Say what you will try next.

A programme that only reports successful interventions is selecting its results, and the audience works that out eventually — at which point the successful results are discounted too.

The honest claim

"Median time from approval request to decision fell from 6.1 days to 1.9, comparing eight-week periods with similar volume and case mix, using the same query and boundary."

Specific, checkable, and it names the conditions.

Not: "we reduced cycle time by 30 percent", which nobody can verify.

The saved query

The mechanism that makes the before-and-after comparison actually comparable.

Save the query, literally, with the boundary, filters and exclusions embedded.

Version it alongside the mapping version.

Re-run it unchanged for the after-measurement.

Where it must change, re-run both periods with the new version and report both pairs.

This single discipline prevents the most common way improvement claims fall apart, which is a comparison between two differently computed figures.

Choosing the control before intervening

The provision that turns a weak before-and-after into a defensible result.

Pick a comparable segment running unchanged: another region, case type, or team.

Chosen before the intervention, which prevents selecting a flattering comparison afterwards.

Measure both, before and after.

The difference in the differences removes seasonality, volume effects and the novelty effect at once.

Report both series, so a reader can see whether the control moved too — and if it did, the intervention explains nothing.

Re-measuring properly

The step that separates a claimed improvement from a demonstrated one.

Same log definition, same boundary, same exclusions.

After a settling period, not immediately, because behaviour changes on observation and then decays.

Comparable period length and case mix.

Reported with the spread and against the normal range, so ordinary variation is not claimed as an effect.

With a control where one exists.

Including null results, which is what makes the positive ones believable to anyone checking.