Glossary

What is process mining? Definition and event-log method

Process mining is a versatile and effective extension of established business process management methods. In addition to simplifying and accelerating the documentation of processes, process mining also offers optimization and analysis.

Definition and event-log data basis

Process mining reconstructs actual process flows from event logs. The required data basis contains at least a case identifier, an activity and a timestamp; attributes such as resource, cost centre or site support deeper analysis and comparison with the intended process.

In some cases, this even takes place across several systems. Process mining is an important step on the way to the digital transformation of your company.

How does process mining work?

Process mining compares real processes with the ideal processes of theory or best-practice solutions. This provides insight and transparency. Especially in IT processes in indirect areas, it is much more difficult to recognize whether a process is running optimally compared to the physical processes in production. These IT processes are less tangible, which is why insight into the actual process is required. Without a process model, such an analysis is almost inconceivable. The process model results from the evaluation of all existing data in the IT systems. Data on business processes is stored in log files and databases. The large amount of data requires an approach similar to data analysis in data mining. The initial approach is to prepare the data and visualize it in a process diagram. This makes it possible to identify additional work and shortcuts, for example. In practice, deviations from processes, e.g. in throughput time, are identified using key performance indicators (KPIs). Process mining goes one step further and analyzes the process and identifies the reason for the delay, thus helping to eliminate it. Data mining thus goes beyond the mere visualization of bottlenecks.

How do deviations from the ideal process occur?

Process mining can be used to identify deviations from the ideal process. The reality often looks different from the ideal. This is also the case with business processes, e.g. in purchasing or sales. In reality, these are much more complex, unstructured and less clear-cut than described in documented processes. The way employees think processes work and the way they actually work also differs significantly in reality. Misinformation can be a critical problem in organizations. It is therefore very important to understand the difference between reality and the ideal. So how does this deviation from reality and the ideal state arise?

Additional work

Employees spend a lot of time on rework and additional steps that are not foreseen in the processes described.

Exceptions

Processes are described as they should run. Exceptions are ignored, even if they occur frequently.

Shortcuts

Processes are abbreviated or skipped completely.

Definition

Processes are often not clearly defined in the first place. Employees often carry out steps differently. This creates a difference. In addition, each employee only sees their own process steps and has no insight into the upstream or downstream processes.

Changes

Processes are frequently changed due to reorganization and restructuring and are therefore often not up to date.

What data is required for process mining?

At least the following data is required to analyze your business processes using process mining:

  • Transaction number(ID)
  • Timestamp
  • Status change
  • Further attributes (optional)

This data can be used to visualize the actual process using process mining software such as Celonis or Signavio. As a rule, this crystallizes a main process, the so-called happy path, as well as other deviating processes as they occur in reality. Process mining not only gives you transparency of the processes but also the visual preparation of key figures. This makes the percentage distribution, deviations and outliers of individual processes visible.

The three basic types

The discipline has a founding document: the Process Mining Manifesto, published in 2012 by the IEEE Task Force on Process Mining under the lead of Wil van der Aalst. It distinguishes three basic types, and the distinction is not academic — it determines which data a project needs in the first place.

  • Discovery. A technique takes an event log and produces a model without using any a-priori information. It is the most prominent technique — and for many organisations the most surprising, because real processes can be reconstructed from example executions alone.
  • Conformance. An existing model is compared with the event log of the same process. The check runs both ways: whether reality conforms to the model and whether the model matches reality. The reference is not limited to procedural models — organisational models, business rules, policies and laws qualify as well.
  • Enhancement. An existing model is extended or repaired using information from the event log. Timestamps, for instance, let bottlenecks, service levels, throughput times and frequencies be written into the model.
Input → technique → output Event logrecorded events Discoveryno prior information Process modelPetri net, BPMN, … Event log+ existing model Conformancecompared both ways Diagnosticsdifferences, commonalities Event log+ existing model Enhancementextend or repair Improved modele.g. with throughput times Conformance measures the gap, enhancement changes the model.
The three basic types of process mining — own illustration after the Process Mining Manifesto (2012), Figure 3.

The difference between the last two is where projects fail: conformance measures the gap between model and reality, enhancement changes the model. Anyone looking for deviations while adjusting the model ends up with no deviations — and nothing learned.

Three misconceptions the manifesto clears up

The document itself states what process mining is not:

  • Not limited to control-flow discovery. Discovery is one of three basic forms, and the scope reaches beyond control flow: the organisational, case and time perspectives matter as well.
  • Not a special case of data mining. The manifesto calls process mining the missing link between data mining and model-driven business process management. Most data mining techniques are not process-centric, and process models with concurrency cannot be expressed as decision trees or association rules.
  • Not limited to offline analysis. The techniques work on historical data, but the results apply to running cases — for example predicting the completion of a partially handled order.

What the data must look like for any of this to work is covered by the event log. Checking against a target model is conformance checking; releasing the rigid case view is object-centric process mining. How a project runs is described on our process mining consulting page.

Sources

van der Aalst, W. et al. (2012): Process Mining Manifesto. IEEE Task Force on Process Mining. In: Daniel, F. et al. (eds.): BPM 2011 Workshops, LNBIP 99, pp. 169–194, Springer. DOI 10.1007/978-3-642-28108-2_19. Further: van der Aalst, W.; Carmona, J. (eds., 2022): Process Mining Handbook. LNBIP 448, Springer, open access under CC BY 4.0.

Process Mining in practice: Our services overview brings together the relevant planning and consulting approaches.

FAQ

Frequently asked questions

What is process mining?

Process mining extracts knowledge about processes from the event data that information systems record anyway. The Process Mining Manifesto (2012) distinguishes three basic types: discovery produces a model from the event log, conformance compares an existing model with reality, and enhancement extends a model with information from the data.

How do conformance and enhancement differ?

Conformance measures the gap between model and reality and returns diagnostics. Enhancement changes the model — it is extended or repaired using information from the event log, for example throughput times and bottlenecks derived from timestamps.

Is process mining a form of data mining?

The Process Mining Manifesto rejects that classification explicitly and calls process mining the missing link between data mining and model-driven business process management. Most data mining techniques are not process-centric, and process models with concurrency cannot be expressed as decision trees or association rules.