← Back to list
AI & Data
#프로세스 마이닝#이벤트 로그#프로세스 발견#적합성 검사#XES#업무 프로세스 개선
Last updated · 2026-09-29

Process Mining for Business Process Analysis and Improvement

1. Overview

Definition: Process mining is a data-driven analysis technique that discovers real business flows from information-system event logs, checks deviations from reference models, and improves process models from time, resource, and data perspectives. [1][3]

Organizational work passes through many information systems, including ERP, CRM, electronic approvals, contact centers, and logistics. Each system records when work was executed and what result it produced, but isolated logs rarely show the full path of one case from beginning to end. Process mining connects events across systems at the case level and reconstructs actual execution paths.

Traditional process analysis uses interviews and workshops to draw an intended procedure and compare it with operations. It captures local knowledge and context well, but can be affected by recall bias, sample selection, and omitted exception paths. Process mining instead uses execution logs to quantify frequent paths, rework, bottlenecks, and approval bypasses.

However, a log is not reality itself. If events are missing, grouped under an incorrect case identifier, or recorded with inconsistent times, even a sophisticated algorithm may produce a misleading process. An information-management professional must therefore design not only algorithms but also process boundaries, log quality, privacy, model interpretation, and improvement controls.

1.1 Need and Objectives

First, process mining makes visible the difference between an actual flow and a documented standard procedure. For example, purchasing policy may require two approval steps, while logs may reveal approvals skipped by amount threshold or repeated at the same step. Together with process owners, the organization must decide whether each difference is a violation, a reasonable exception, or a system-design defect.

Second, it decomposes lead time and bottlenecks by activity. Knowing only that total processing time is long does not reveal which department or waiting stage needs improvement. Linking activity start and completion times with resource data separates working time from waiting time and rework.

Third, it supports ongoing compliance and operational-risk checks. A few audit samples may not represent the distribution of all process paths, whereas event logs can cover a large population or all recorded cases. Still, an automatically detected deviation must not be declared unlawful or fraudulent until its business context and log limitations are verified.

2. Core Concepts and Analysis Perspectives

An event log is the starting point for process mining. [1][3] It is a collection of process executions, each represented as a time-ordered group of events. The analysis result depends directly on how data was recorded and how case boundaries were defined.

flowchart LR
    B[Business goals and rules] --> M[Reference process model]
    S[ERP, CRM, business systems] --> E[Event log]
    E --> P[Process mining]
    M --> P
    P --> D[Actual flow, deviations, bottlenecks]
    D --> A[Business validation and diagnosis]
    A --> I[Process improvement and controls]
    I --> B

2.1 Event-Log Structure

A case is one unit of work execution, such as one order, inquiry, or insurance claim. An activity is a business step such as submission, review, approval, or payment; an event is one recorded occurrence of an activity within a particular case. The same activity may occur more than once, so activity names alone cannot distinguish the order or individual executions.

A basic log needs a case identifier, activity name, and timestamp. Adding resources (person, department, or system), activity start/completion state, case attributes such as amount or customer type, and outcomes broadens the analysis. Because timestamps determine event order, time zones, clock synchronization, and the difference between event time and ingestion time must be managed.

Element Example Analytical meaning
Case identifier Order O-301 Key for grouping one execution
Activity Order received, credit check, shipment A process step
Timestamp 2026-09-29 09:15 Order, waiting time, and duration
Resource Person, team, automation bot Work allocation and workload
Attribute Order value, channel, risk grade Path and performance by condition
Outcome Approved, rejected, canceled Relationship between path and result

The log must be normalized as event sequences for each case, rather than treated as an ordinary table. When ERP and approval systems record the same event, duplicates or different names for the same activity may arise. Unstable identifier mapping and activity taxonomies can split one case or merge different business processes.

2.2 Three Basic Types of Analysis

Process mining is commonly described through process discovery, conformance checking, and model enhancement. [1][3] They differ in how event logs and reference models are used and are connected analysis stages rather than substitutes for one another.

Type Input Core question Typical output
Process discovery Event log What does the actual flow look like? As-is process model
Conformance checking Log and reference model Does execution match the intended model? Deviation and conformance diagnosis
Enhancement Log and existing model How can time, resource, and data improve the model? Model enriched with performance, organization, or prediction

A. Process Discovery

Process discovery infers execution flows from logs without a predefined process model. The result describes the current As-is process and may expose alternative paths and repeated activities that were not previously known. The discovered model must be distinguished from the target To-be model so it is not mistaken for the approved standard.

For example, a reimbursement log may reveal not only Submit → Review → Approve → Pay but also repeated Review → Request information → Review loops. The model alone cannot tell whether a loop reflects a legitimate complex case or rework caused by inadequate initial guidance. The model is evidence for beginning an improvement discussion, not a replacement for business judgment.

B. Conformance Checking

Conformance checking compares a reference model with actual logs to diagnose how much behavior conforms and where deviations occur. For compliance analysis, first determine whether the model is a normative requirement or merely a descriptive picture of current work. If reality differs from a descriptive model, the model may not capture reality; if it differs from a normative model, approval bypasses or exception-control problems may exist.

Not every deviation is an error. Bypassing a mandatory control grounded in law or risk policy may warrant investigation, while an approved emergency path may be legitimate. Combine deviations with case attributes, authority, and business reason, and retain evidence and traceability for human review.

C. Model Enhancement

Model enhancement adds performance, resource, and data information from logs to an existing process model or corrects the model itself. Showing activity and waiting time with colors can reveal bottlenecks, and adding resource information helps examine queues concentrated in particular teams. When model and log do not fit well, deviation diagnostics can guide revision of missing conditions or branch rules.

Predictive analysis can estimate whether an in-progress case will complete or how much time remains. Past patterns may not persist after a policy change, so a prediction should not be the sole basis for automatic approval or rejection. Enhancement is not just better analytics; it also designs when people intervene and who is accountable.

3. Processing Structure and Analysis Methods

In practice, source logs should not be fed directly into a mining tool; business definitions and data pipelines must be designed together. Agree first on process start/end conditions and the case unit, then map system-specific events to common activities and time conventions. Next, validate log quality and iteratively perform exploration, modeling, and business review.

flowchart TD
    G[Define goals, scope, cases] --> X[Extract system events]
    X --> C[Link identifiers and map activities]
    C --> Q[Check time, duplicates, omissions]
    Q --> F[Filter and pseudonymize logs]
    F --> A[Discovery, conformance, performance analysis]
    A --> V[Business and audit validation]
    V --> R{Improvement approved}
    R -->|Yes| D[Improve process and system controls]
    R -->|Revise| C
    D --> O[Monitor metrics and reanalyze]

3.1 Log Extraction and Preparation

The first step is to define case boundaries for the analysis objective. Whether order fulfillment and refunds form one case or separate cases changes the discovered paths and lead times. When one event relates to several objects such as customer, order, and shipment, forcing it into one case key can lose relationships.

Next, map system event codes to a common activity dictionary. PAY_OK, PaymentSettled, and Payment completed may mean the same thing, but merging completion, cancellation, and failure removes important business distinctions. Manage mapping rules and change history, and verify whether historical data remains comparable when event definitions change.

Distinguish event-occurrence time, application-recorded time, and data-ingestion time. Clock differences across systems, delayed ingestion, and inverted timestamps can produce incorrect activity order or negative durations. Rather than deleting anomalous events automatically, document the correction, quarantine, and exclusion rules and record their effect.

3.2 Log Standards and Model Representations

XES (eXtensible Event Stream) is an IEEE standard for interoperability in event logs and event streams. [2] IEEE 1849-2023 specifies grammar and XML schemas for logs, streams, and extensions, and is listed as the active standard replacing the 2016 edition. Even with a standard format, business meanings and case-linkage rules do not become uniform automatically, so semantic mapping is still needed.

Choose process-model representations for the analysis objective and audience. Petri nets are useful for rigorously handling concurrency, token flow, deadlocks, and reachability, while BPMN readily communicates roles, events, and gateways to business owners. A Directly-Follows Graph (DFG) quickly explores activity frequency and direct succession but cannot always express branching and repetition precisely.

Representation Strength Caution and use
Petri net Concurrency, execution semantics, conformance analysis May be complex for non-specialists
BPMN Business roles, events, and branches Review mining output to restore business meaning
DFG Fast view of frequency and direct succession Concise, but may oversimplify control flow
Process tree Structured discovery and block-structured flows Manage model complexity when behavior is highly irregular

3.3 Discovery Algorithms and Quality

The alpha algorithm is an early representative technique that infers causal and concurrent relations from event logs to construct a Petri net. [3][5] It can explain simple logs, but may be sensitive to noise, rare paths, and short loops, making direct application to real logs difficult. An algorithm does not produce a single ground truth; it offers different trade-offs in simplification and explanatory power under its assumptions.

Heuristics Miner methods use frequency information to reduce the influence of rare behavior, while Inductive Miner can recursively discover process trees. [5] In practice, do not stop at filtering, parameter selection, and visualization; compare representative cases and deviations with what business experts know. An overly detailed model becomes hard to understand because it explains every exception, while an overly simple one can hide real risk paths.

Assess model quality through fitness, precision, generalization, and simplicity together. [3][5] Fitness measures how well the model reproduces observed executions; precision examines whether it allows excessive behavior absent from the log. Because improving one metric can sacrifice another, explain the balance according to business risk and decision purpose.

4. Performance Metrics and Interpretation

Process analysis addresses time, frequency, resources, and outcomes as well as the shape of the flow. Lead time is elapsed time between case start and completion; activity time is time spent performing an activity; waiting time is delay between activities. The recorded completion timestamp may not mean that actual work ended, so metric definitions and formulas must be agreed with process owners.

Perspective Representative measures Operational question
Flow Path frequency, variant count, rework count Which paths are standard and where do exceptions arise?
Time Case lead time, activity and waiting time, percentiles Which activity or queue causes delay?
Resources Work per person, handoffs, concentration Is work overly concentrated in a role or team?
Compliance Deviations, unapproved bypasses, missing mandatory steps Are mandatory controls reflected in actual execution?
Outcome Rejection, cancellation, rework, completion result How are paths related to outcomes?

The mean can hide a small number of very long cases, so examine the median and upper percentiles as well. Waiting between activities may arise from staff availability, approval policy, a system queue, or dependence on an external organization. Do not mistake observed correlation for causation; validate causes through improvement experiments or business investigation.

5. Case: E-Commerce Order Approval

This is a hypothetical scenario to explain analysis, not a claim about the performance of a real company. Assume an e-commerce business wants to reduce lead time from order receipt to shipment and verify controls for high-risk orders. The log contains order ID, receipt, payment, risk review, approval, and shipment events, timestamps, risk grade, and channel.

Discovery shows that ordinary orders follow Receipt → Payment → Shipment, while some high-risk orders require manual review and additional approval. Conformance checking separates cases approved by a user without sufficient authority from cases reordered because of an external payment delay. Performance analysis compares elapsed time from receipt to shipment with the review queue and approval wait separately.

The team does not immediately sanction deviation cases; it investigates risk grade, approval authority, system exceptions, and missing records. If the confirmed cause is a staffing bottleneck, it revises shift or assignment policy; if policy interpretation is unclear, it improves the authority matrix and system validation. After the change, it monitors both lead time and missing mandatory approvals to ensure that faster processing does not weaken controls.

The core lesson is not to optimize a single metric. If fraud-prevention steps are skipped to reduce shipment time, losses and customer harm may increase. Balanced measures should assess cost, speed, customer experience, and compliance together.

6. Comparison with Related Techniques

Interviews and BPMN modeling clarify stakeholder knowledge and target procedures, but may not directly show actual frequency or distribution of exceptions. Process mining quantifies the current state from logs, but cannot automatically recover unrecorded work or tacit knowledge. The approaches complement one another: interviews interpret model meaning, while log analysis validates evidence of execution.

General data mining focuses on patterns such as classification, clustering, and regression. Process mining centers on activity order and case flow, connecting data mining, BPM, and process modeling. BI dashboards aggregate outcome metrics, while process mining emphasizes the execution paths and deviations that produced those outcomes.

Dimension Business modeling and interviews Process mining General data mining and BI
Main evidence Expert knowledge and design rules Event logs and actual execution Structured data and metrics
Time order Expressed as required by model Case-level event order is central Selected according to analysis
Strength Clarifies goals, ownership, and rules Examines actual paths and deviations at scale Supports prediction, aggregation, and pattern discovery
Limitation Recall bias and omitted exceptions Depends on log quality and case linkage May lack process-flow context
Complementary role Defines To-be and business meaning Diagnoses As-is, conformance, and performance Supports outcome metrics and risk prediction

7. Advanced Topic: Log Standards and Data Governance

IEEE 1849-2023 XES supports interoperability by expressing event logs and streams in a common structure. The standard format enables data exchange; it is not a universal converter that resolves different activity names, case definitions, and time conventions across systems. Organizations should jointly manage data contracts, activity dictionaries, case-identification rules, time conventions, and quality ownership.

Process mining can analyze user and employee behavior in detail, creating privacy and worker-monitoring risks. Clarify the process-improvement purpose, use only necessary attributes, pseudonymize identifiers, and restrict secondary use and retention. Using resource-level performance analysis directly for discipline or automated personnel evaluation can cause bias, missing context, and poor explainability.

8. Considerations and Implications

8.1 Agreement on Process Boundaries and Case Identity

Before analysis, agree with process owners on start/end conditions, case units, reopening, cancellation, and split/merge rules. Since case boundaries change path counts and lead times, record them in metadata and analysis versions.

8.2 Event-Data Quality and Time Consistency

Validate how missing mandatory events, duplicates, inconsistent activity codes, and time-zone errors distort models and metrics. Preserve extraction and transformation rules, quality-check results, and excluded-case lists to make analysis reproducible.

8.3 Privacy and Purpose Limitation

Remove or pseudonymize identifiers and free-text content by default, and apply role-based access controls and audit logs. Legal, security, labor, and privacy stakeholders should review the analysis to prevent repurposing it for unrelated personnel evaluation or surveillance.

8.4 Model Complexity and Explainability

Increasing model complexity to explain every log can make it unintelligible to business users; simplifying it too much can hide rare but important risk paths. Separate filtering and detail levels by process risk, and disclose model assumptions, exclusions, and quality measures.

8.5 Measuring Improvement and Side Effects

For before-and-after comparisons, account for workload, seasonality, and system changes; where possible, use phased rollout or a control group. Measure error rates, customer complaints, compliance breaches, and employee burden alongside lead time and cost to prevent metric-optimization side effects.

8.6 Ongoing Operation and Accountability

When process definitions change, update activity dictionaries, log mappings, reference models, and rules through change management. Assign responsibilities to process owners, data engineers, security/audit roles, and process analysts, maintaining a closed loop from discovery to approved improvement and remeasurement.

References

  1. IEEE Task Force on Process Mining, Process Mining Manifesto: https://www.tf-pm.org/resources/manifesto
  2. IEEE Standards Association, IEEE 1849-2023: Standard for eXtensible Event Stream (XES): https://standards.ieee.org/ieee/1849/10907/
  3. Wil M. P. van der Aalst, Process Mining, Communications of the ACM: https://cacm.acm.org/research/process-mining/
  4. Process Intelligence Solutions, PM4Py: https://processintelligence.solutions/pm4py/
  5. Leemans, Fahland, and van der Aalst, Scalable process discovery and conformance checking: https://link.springer.com/article/10.1007/s10270-016-0545-x

In one line: Process mining discovers, checks, and improves real work from event logs, while requiring joint design of log quality, business context, privacy, and outcome controls.