← Back to list
Database
#CRUD#데이터모델링#상관분석#정합성검증#프로세스#133회
Last updated · 2026-09-27

CRUD Matrix

1. Overview

A. Definition

A correlation analysis tool that expresses the create (Create), read (Read), update (Update), and delete (Delete) relationships between an information system's processes (functions) and entities (data) as a matrix, in order to check the consistency between the data model and the process model.

The CRUD matrix is not merely a single table but a verification device that forces two different design viewpoints onto one plane so they can be seen together. The data viewpoint (what to store) and the function viewpoint (what to process) each look natural early in development, yet their misalignment surfaces only when they must actually work in tandem. The value of the CRUD matrix lies in exposing this misalignment before any code is written, during the analysis and design phase.

B. Background and Necessity of Its Emergence

In the Information Engineering (IE) methodology, the data model (ERD) and the process model (DFD, function decomposition diagram) are separate deliverables, often produced independently by different people. The problem is that the two deliverables must logically mesh with each other. If not a single process feeds values into a given entity, that data remains forever empty (creation omission); conversely, if a process is designed to reference data that does not exist, execution itself becomes impossible (breakdown of referential integrity).

Such holes in data-function consistency are easy to miss with the human eye no matter how carefully the documents are read, because the number of combinations explodes as entities grow to dozens and processes to hundreds. By forcibly aligning these relationships into a grid form that can be grasped at a glance, the CRUD matrix visually reveals omissions (empty cells), duplication (excessive marks), and isolation (entirely empty rows/columns). Furthermore, once written, the matrix does not end at verification but is reused as common input material for transaction boundary setting, distributed DB design, access-control (RBAC) design, and archiving/retention policy, making it highly reusable even among analysis deliverables.

C. Characteristics

The characteristics of the CRUD matrix can be summarized in three points. First, symmetric verification. It simultaneously checks the data side (whether an entity has its full life cycle) and the function side (whether a process handles real data). Second, it is a static deliverable with dynamic implications. The table itself is static, but the fact that several processes perform U on the same entity leads directly to the runtime issue of concurrency and lock design. Third, it is methodology-neutral. Although it originated in Information Engineering, it is applied as-is as an access-pattern analysis tool in object-oriented and microservice design.

2. Overall Structure and Position of the CRUD Matrix

Understanding first where this tool receives input and where it sends output makes clear why it is called the "hub" of the analysis phase. The structure diagram below shows the flow in which the data model and the process model converge into the CRUD matrix and then diverge again into various design activities.

flowchart LR
  ERD["Data model (ERD/entities)"] --> CM["CRUD matrix"]
  DFD["Process model (DFD/functions)"] --> CM
  CM --> V["Consistency verification (omission/isolation/phantom)"]
  CM --> TX["Transaction boundary design"]
  CM --> DIST["Distributed DB / partition design"]
  CM --> AUTH["Access control (RBAC) design"]
  CM --> ARC["Archiving / retention policy"]

As the diagram shows, the inputs to the CRUD matrix are two different models, and the outputs are five branches of subsequent design activity. The key here is the fact that the two inputs "were made independently." If they had been auto-generated from a single unified model, consistency verification would be meaningless. Precisely because two models made independently from different viewpoints are overlaid after the fact, inconsistencies that people failed to notice surface.

The matrix's coordinate system also follows a clear rule. By convention, rows hold processes (functions) and columns hold entities (data). This way, reading a row horizontally reveals the data-access scope of a transaction, i.e., "which data this process handles and how," while reading a column vertically reveals the life cycle of the data, i.e., "by whom and how this data is handled." Horizontal reading corresponds to functional design, vertical reading to data governance.

3. Notation and Writing Procedure

Each intersecting cell records the operations the process performs on that entity as C/R/U/D, and if one process performs several operations, they are written together. For example, "order cancellation" modifies an order and ultimately deletes it, so it is marked UD. In the shopping-mall example below, reading the "order registration" row horizontally reveals in one line that this function is a composite transaction that reads (R) member and product and creates (C) an order.

Process \ Entity Member Order Product
Sign-up C
Product lookup R
Order registration R C R
Order cancellation UD

Writing it is not a matter of filling cells at random but follows a set procedure. The procedure diagram below shows the flow from entity/process derivation to verification and reuse.

flowchart TD
  A["1. Derive entity list (ERD-based)"] --> B["2. Derive process list (decomposition-based)"]
  B --> C["3. Build matrix frame (rows=processes, cols=entities)"]
  C --> D["4. Mark C/R/U/D in each cell"]
  D --> E["5. Verify row/column consistency"]
  E --> F{"Omission/isolation/phantom exists?"}
  F -->|"Yes"| G["Supplement model, then rewrite"]
  G --> D
  F -->|"No"| H["Reuse in subsequent design"]

The first step, deriving entities and processes, determines 80% of the matrix's quality. As a rule, the entity list comes from a normalized ERD and the process list from the lowest-level (atomic) processes of the function decomposition diagram. If the granularity of entity definition (conceptual/logical/physical) and process definition (major function/unit task) is not aligned, the matrix becomes too sparse or, conversely, too dense, rendering verification meaningless. In practice, they are often aligned at the level of logical entities and unit processes (corresponding to one screen or API).

Second, in the cell-marking step, the team needs to agree on "what counts as an R." For instance, whether an existence-check lookup for foreign-key validation is marked R, or only display-purpose lookups count as R, changes the matrix's density. If this rule is unclear, two people drawing the same system will produce different matrices.

Third, verification and iteration. When a defect is found in verification, the key is not merely to fix the matrix but to go back to the source models (ERD/DFD) and supplement them. The matrix only shows the symptom of a defect; the real problem lies in the data model or the process model. This feedback loop makes the CRUD matrix a verification process rather than a mere document.

4. Consistency Verification Rules

The core of verification is to symmetrically confirm "whether every entity has its full life cycle" and "whether every process handles real data." Each rule corresponds to a specific defect to be suspected upon violation.

Each entity column must have at least one C. An entity with no creating process has no path for data to flow in, so unless a separate inflow path such as batch loading or external interfacing is specified, it is a clear design omission. In the earlier example, the fact that the product column has no C tells us that the "product registration" process is missing.

Each entity column must have at least one R. Data that nobody reads is likely unnecessary data with no reason to be stored. However, since there are cases like log/audit data that are "accumulated now for later even if not read," rather than immediately deciding to delete when there is no R, treat it as a review signal to reconfirm the retention purpose.

The presence or absence of U/D checks the need for management processes. An entity with no U at all may be immutable data, and an entity with no D may be a policy managed only by a status flag without physical deletion (logical deletion). That is, the absence of U/D is not a defect in itself but a question that makes you distinguish whether it is a policy choice or an omission.

Every row and column must have at least one mark. A column with no marks at all is an isolated entity that no function uses, and a row with no marks at all is an empty (phantom) process that touches no data; both are error candidates. An isolated entity raises suspicion of over-design or a missing process, and a phantom process raises suspicion of a shell function that does not actually operate.

Check rule Meaning Suspicion upon violation
C exists for each entity Secures data inflow path Missing creation process
R exists for each entity Confirms data utilization Unnecessary data or retention-purpose review
Presence of U/D Confirms state-change/deletion policy Missing management process vs. immutable/logical-delete policy
At least one per row and column Prevents isolation/phantom Isolated entity / empty (phantom) process

5. Areas of Use and Real Cases

The CRUD matrix does not stop at consistency verification but is widely used as input for subsequent design. In transaction analysis, when one row (process) performs C/U/D across several entities, that scope becomes one logical unit of work, i.e., a transaction boundary. For example, if the "order registration" row includes both order C and inventory U, these must be committed atomically, so the design rationale that they should be bound into one transaction comes straight from the table.

The rationale for concurrency control design is also read from columns. If many processes' Us are concentrated in a particular entity column, that data is a high-contention hotspot, so a lock strategy and whether to use optimistic concurrency control must be decided. In practice, entities where U concentrates, such as inventory or balances, are the representative case; missing this point leads to lost updates or deadlocks during operation.

In distributed/partitioned DB design, the locality of access patterns is the key clue. If only a particular group of processes accesses a particular group of entities, splitting (sharding/partitioning) the DB along that boundary improves performance by increasing traffic locality. Conversely, if many processes span many entities broadly, distributed-transaction costs grow upon splitting, so caution is needed. In access-control design, one starts from this matrix to expand into an RBAC permission matrix that maps which CRUD permission each role has on each entity. For example, the "general member" role is given only C/R on orders, while the "administrator" role also has U/D.

Area Way of use Clue read from the table
Business/data verification Requirements-data consistency check Empty cells, isolated rows/columns
Transaction analysis Transaction boundary/concurrency design A row's C/U/D scope, a column's U concentration
Distributed DB design Access-pattern-based split/placement Process-group to entity-group access locality
Access-control design Basis for role-based CRUD permission matrix Role × entity × operation mapping

6. Comparison with Similar Techniques

The CRUD matrix differs in purpose from other correlation analysis tools. Distinguishing them lets you describe them without confusion in an exam answer. The Entity-Function correlation table (E-F Matrix) marks only the presence or absence of a relationship, without distinguishing CRUD, and is used at the derivation stage, whereas the CRUD matrix subdivides down to the operation type and is used for consistency and transaction design. A DFD focuses on the flow of data (where it moves), while the CRUD matrix focuses on the operations on data (what is done to it), making the two complementary.

The reason this difference arises is that each tool targets a different design question. The E-F Matrix asks about existence, "is there a relationship," the DFD about the movement path, "how does it flow," and the CRUD matrix about consistency, "is the life cycle completed." Therefore, in practice, the three tools are not chosen exclusively but used together sequentially according to the analysis stage. Early in derivation, the E-F Matrix captures large relationships; in the detailing stage, CRUD fills in the operations to verify; and the DFD confirms the flow.

7. Deep Dive — The CRUD Matrix in the MSA Era and Expected Exam Directions

The modern reinterpretation of the CRUD matrix stands out in microservice architecture (MSA). Whereas in monolithic systems CRUD pattern analysis was the rationale for DB partitioning, in MSA the same analysis becomes the rationale for defining the data boundary a service should own, i.e., the bounded context. To uphold the data-ownership principle that "only the service that owns the data changes (C/U/D) it," one must first grasp which process performs writes (C/U/D) on which entity. Since the CRUD matrix is exactly the map of that write pattern, a cell showing that several services perform writes on one entity is a warning sign that the service boundary has been drawn incorrectly.

From here, the linkage with CQRS and event sourcing is naturally derived. If reads (R) span many services/processes broadly while writes (C/U/D) concentrate in a few, CQRS that separates the read model and the write model is advantageous. That is, the distribution of R versus C/U/D in the CRUD matrix itself becomes the basis for deciding whether to apply CQRS. Recently, from a data-governance perspective, there is also a trend of extending the CRUD matrix by connecting it with data lineage and personal-information processing records to clarify "which personal-information entity is collected (C), used (R), corrected (U), and destroyed (D) by which function," as compliance evidence material.

From the perspective of the Professional Engineer exam, this topic is frequently used not only as a standalone short answer but also as partial supporting evidence in data-modeling, transaction, and MSA design questions. When composing an answer, a progression that (1) presents the definition and notation with an example matrix, (2) describes the four consistency verification rules together with violation cases, and then (3) expands step by step into transaction, distribution, permission, and MSA gives the grader a sense of completeness. In particular, describing each rule with the suspicion upon violation attached, rather than "just listing tables," is highly distinguishing.

8. Considerations and Implications

  • Need for automation and tooling: In large-scale systems with hundreds of entities and processes, manual management is impossible. Use CASE tools and metadata-repository-based auto-generation to synchronize model changes with the matrix, and have the matrix update automatically when the model changes—this is the crux of maintaining consistency.
  • Granularity alignment and notation-rule standardization: Align the definition levels of entities and processes, and codify notation rules such as "what counts as an R" as a team standard, so a reproducible matrix results. Without rules, different people produce different tables, and the reliability of verification collapses.
  • Foundational deliverable for data governance and compliance: It is the starting point for managing data ownership, quality, and life cycle, and can be connected with personal-information processing records and data-lineage management to serve as evidence material for GDPR and privacy-law compliance.
  • Extension to MSA/CQRS design: CRUD pattern analysis becomes the basis for deciding a service's bounded context and whether to separate reads and writes (CQRS). However, in MSA, owned data must be divided and managed per service rather than in one physical matrix, so hierarchical management of the enterprise-wide integrated matrix and per-service matrices is needed.
  • Recognizing the limits of static verification: The CRUD matrix verifies only static consistency at design time and does not guarantee actual runtime load, contention, or performance. Therefore, it must always be paired with dynamic verification such as load testing on U-concentrated entities and transaction-isolation-level verification.

In one line: The CRUD matrix is a correlation analysis tool that expresses the create, read, update, and delete relationships between processes and entities as a matrix to symmetrically verify data-function consistency (omission, isolation, phantom), and extends to transaction, concurrency, distribution, and access-control design as well as the basis for MSA bounded contexts and CQRS decisions.