Simulation and Prediction, Honestly
Both are sold heavily and both rest on assumptions the log cannot supply. Where they help and where the output is confident fiction.
Analysis
Simulation and prediction are the features that sell platforms. They are also where the gap between demonstration and reality is largest.
What simulation needs that the log does not have
Resource capacity: how many people are available, when, and how their time is split across processes.
Working time calendars, including holidays and site differences.
Arrival distributions, not just historical volumes.
Priority rules, which are frequently informal.
Service time distributions separated from waiting, which requires start and end timestamps.
Most logs supply none of the last three, which means the simulation is running on assumptions the analyst supplied.
Where simulation helps
Comparing two designs directionally. Whether removing a handover helps more than adding capacity.
Sizing a queue under a stated arrival rate.
Sensitivity analysis: which assumption the answer depends on, which is frequently the most useful output.
Communicating a bottleneck to people who do not read charts.
Where it does not
Producing a number to put in a business case. The output inherits every assumption and carries none of the uncertainty visibly.
Predicting the effect of a change to human behaviour, which no simulation models well.
Anything where the assumptions dominate the data, which is most of the time.
State the assumptions on the same page as the result, and if the result changes materially when one is adjusted, that is the finding.
Prediction claims
Predicting which cases will be late, or will rework, is a genuine and useful capability.
It works when the pattern is stable and the features are available early.
It fails silently when the process changes, which processes do, and a model trained on last year's paths degrades without announcing it.
What to check before trusting a prediction
What features it uses, and whether they are available at the point the prediction is needed. A model using data that only exists at case closure predicts nothing useful.
Whether it was validated on a later period, not a random split, since a random split leaks future information.
The base rate. A model predicting a rare outcome can be highly accurate and useless.
What it costs to be wrong, in both directions.
Whether anyone acts on it, which is the real test.
The automated decision boundary
A prediction that triggers a consequence for a person — a case reassigned, a customer deprioritised, a worker flagged — engages automated decision restrictions in several jurisdictions.
Human involvement and a route to contest are commonly required.
Which is an argument for using predictions to prioritise work rather than to judge people, and that is also where they are most useful.
The honest position
Discovery and conformance are the reliable core of this field.
Simulation is a thinking aid.
Prediction is useful in narrow, stable, well-validated cases.
A project justified on the last two will struggle; one justified on the first two and using the others as support will not.
Stating the assumptions on the same page
The discipline that makes a simulation honest.
List every input the log did not supply: capacity, calendars, arrival distribution, priority rules, service times.
Give the value used and its source, which is frequently the analyst's judgement.
Run the sensitivity: change each by a plausible margin and record how much the answer moves.
Report the assumption the answer is most sensitive to, which is usually the most useful output of the whole exercise.
A simulation result presented without its assumptions is a number with unknown error bars, and it will be quoted as though it had none.
Validating a prediction properly
Four checks before anyone acts on a model's output.
Validated on a later period, not a random split, since a random split leaks future information.
Features available at prediction time, not at case closure.
Compared against the base rate, because a model predicting a rare outcome can be accurate and useless.
Monitored for drift, since a model trained on last year's paths degrades without announcing it.
And one operational check: does anyone act on it. A prediction nobody uses is a maintenance cost.
What to ask before trusting a simulation
Simulation is bundled with most tools and its outputs are the least reliable thing they produce.
Where do the service time distributions come from? If from the log's elapsed times, they include waiting and are wrong.
How is resource availability modelled? Usually crudely, and it dominates the result.
Does it model the exceptions, or only the frequent paths?
Has any past prediction been checked against what actually happened? Almost never, and that is the answer.
Use it to compare options directionally, not to produce a number anyone will commit to.