← Back to list
Infrastructure & Cloud
#멀티테넌시#SaaS#테넌트격리#RLS#NoisyNeighbor
Last updated · 2026-10-08

Multi-tenancy Architecture

1. Overview

Multi-tenancy is a software architecture approach in which a single application instance and shared infrastructure serve many customers (tenants) simultaneously, while each tenant experiences a logically isolated, seemingly dedicated environment.

The rise of SaaS (Software as a Service) shifted the software delivery model from "selling a product" to "selling a subscription," and the core technology underpinning the economics of this shift is multi-tenancy. In the traditional on-premises/hosted model, each customer received a separate server, database, and application instance, so when customers grew to 1,000 the number of things to operate grew to 1,000 copies as well, causing maintenance, patching, and monitoring costs to grow linearly. Multi-tenancy, by contrast, pools resources so a single deployment or patch updates all customers at once, dramatically lowering the marginal cost per additional customer. Salesforce opening the SaaS market in the early 2000s by serving hundreds of thousands of organizations from a single codebase is the archetypal example.

What distinguishes multi-tenancy from mere "server sharing" is that it must simultaneously achieve two conflicting goals: tenant isolation and resource efficiency. Pursue isolation to the extreme and each tenant ends up with dedicated resources, hurting efficiency (effectively single-tenancy); pursue efficiency to the extreme and the risk grows that one tenant's failures, overload, or data spill over to others. Therefore the essence of multi-tenancy design is deciding, per layer — data, compute, operations, and billing — where to strike the balance between isolation and efficiency. This article covers the isolation models, the core design challenges (tenant identification, Noisy Neighbor, security), a comparison of data isolation models, real-world cases, and considerations from a professional engineer's perspective, all at an advanced level.

Multi-tenancy has the following characteristics. First, a single logical instance serves all tenants, securing operational simplicity. Second, every request carries a Tenant Context, so data, configuration, and permissions branch by tenant. Third, configuration-based customization provides per-tenant screens, workflows, and fields without code branching. Fourth, because resources are shared, elastic scaling and cost allocation (Chargeback/Showback) become part of the design.

2. Overall Structure of Multi-tenancy

A multi-tenant system is built as a vertical isolation pipeline that identifies the tenant the moment a request arrives, propagates that context across all layers, and ultimately accesses per-tenant isolated data. The structure diagram below shows the entire flow.

graph TD
    U1["Tenant A user"] --> GW["API Gateway / router"]
    U2["Tenant B user"] --> GW
    GW --> RES["Tenant Resolver<br/>subdomain, token, header"]
    RES --> CTX["Tenant Context injection"]
    CTX --> APP["Shared application instance"]
    APP --> POL["Isolation & authz policy engine<br/>(RLS, filter, quota)"]
    POL --> DL["Data access layer"]
    DL --> DBA[("Tenant A data")]
    DL --> DBB[("Tenant B data")]
    APP --> CFG["Per-tenant config store<br/>(settings, metadata)"]
    POL --> QOS["Resource quota, Rate Limit"]

The first element to act in the structure above is the Tenant Resolver. If the system cannot reliably determine which tenant a request belongs to, all subsequent isolation is meaningless. Identification methods include subdomain (tenant-a.service.com), URL path (/t/tenant-a/...), the authentication token (a tenant_id claim in the JWT), and HTTP headers; in practice it is safest to treat the JWT claim as the primary source of trust and use the subdomain only for routing convenience. Making a client-mutable value such as a URL or header the sole basis exposes the system to Tenant Impersonation attacks.

The identified tenant ID is elevated to the Tenant Context and propagated across the entire request-handling path. Java's ThreadLocal, Spring's request-scoped beans, and Node.js's AsyncLocalStorage are used, and the propagation mechanism must be strictly managed so the context is not lost or cross-contaminated with another request's context in asynchronous/thread-pool environments. In reality, many of the most catastrophic multi-tenancy incidents are of the type "request A is processed under B's context, leaking data," which originates from connection-pool reuse, missing cache keys, or misuse of global variables.

A. The Isolation Level Spectrum

Multi-tenancy is not a dichotomy of "isolated or not"; it is a design combined along a spectrum between shared (Pool) and dedicated (Silo) for each resource layer. The diagram below shows where the three representative models sit on the isolation and efficiency axes.

graph LR
    subgraph POOL["Pool model: fully shared"]
        P1["Shared DB<br/>shared schema<br/>tenant_id column"]
    end
    subgraph BRIDGE["Bridge model: partial isolation"]
        B1["Shared DB<br/>per-tenant schema"]
    end
    subgraph SILO["Silo model: full isolation"]
        S1["Per-tenant<br/>dedicated DB/instance"]
    end
    POOL -->|"stronger isolation →"| BRIDGE
    BRIDGE -->|"stronger isolation →"| SILO
    SILO -->|"← stronger efficiency"| BRIDGE
    BRIDGE -->|"← stronger efficiency"| POOL

In the Pool model, all tenants use the same DB, schema, and tables, and each row carries a tenant_id column to mark ownership. Resource efficiency and operational simplicity are highest, but every query must apply the tenant filter without omission, and one tenant's large data set can bloat shared indexes, so the risk of interference is high. The Silo model gives each tenant a dedicated DB (or dedicated instance/account), providing the strongest isolation and making regulatory compliance, performance guarantees, and per-tenant backup/restore easy, but operational targets grow with the number of tenants, so patching, monitoring, and schema migration costs surge. The Bridge (Hybrid) model distinguishes tenants by per-tenant schema (or table prefix) within a single DB, striking a compromise between the two extremes. In practice, a tiered/mixed strategy — "accommodate the many small tenants in a Pool and large/regulated customers in a Silo" — is most common.

B. Comparison of Data Isolation Models

The choice of data-layer isolation model is the most far-reaching decision in multi-tenancy design. The table below organizes the trade-offs of the three models, while the "why" of the choice is explained in prose after the table.

Category Pool (shared schema) Bridge (per-tenant schema) Silo (dedicated DB)
Isolation level Low (logical) Medium High (physical)
Resource efficiency Very high High Low
Cost per tenant Very low Low High
Customization Limited Possible per schema Fully free
Backup/restore (per tenant) Hard Medium Easy
Schema migration One for all Repeated per tenant Repeated per tenant
Regulatory/data sovereignty Hard Moderate Easy
Suitable scale Thousands–millions of small tenants Mid scale Few large/regulated customers

The reason the Pool model is overwhelmingly favorable in cost and scalability is that the operational unit stays fixed at "1," independent of the number of tenants. When a schema change is needed across 100,000 tenants, Pool finishes with a single ALTER TABLE, whereas Silo must orchestrate 100,000 migrations, and if some of them fail, the schema versions diverge across tenants — the drift problem. Conversely, the reason Silo is preferred in regulated industries (finance, healthcare, public sector) is that tenant data is physically separated, making it easy to prove by audit that "it is not mixed with other customers' data," and naturally satisfying Data Sovereignty requirements to keep data in a specific country or requests for per-tenant selective deletion (the GDPR right to erasure). For example, keeping EU customer data in a Frankfurt-region dedicated DB while placing US customers in a Virginia region is implemented cleanly in a Silo.

The key mechanism for safely operating the Pool model is Row-Level Security (RLS). The application code's WHERE tenant_id = ? filter becomes a single point of failure that exposes all tenants' data if a developer omits it in even one place. Enforcing the tenant filter automatically at the DB-engine level based on a session variable (SET app.tenant_id), as with PostgreSQL's RLS policies, lets the DB serve as the last line of defense even when the application has a bug. This is an application of the Defense in Depth principle: "do not trust the application; enforce isolation at the data layer."

C. Tenant Onboarding and Lifecycle Management

Multi-tenancy holds up operationally only when lifecycle automation of the tenant creation → operation → termination backs up runtime isolation. On new-tenant onboarding, tenant record creation, dedicated-resource provisioning (DB/schema in a Silo), injection of the initial admin account and default configuration, and application of isolation policies must be performed through an idempotent, automated pipeline. Manual onboarding becomes a bottleneck the moment tenants exceed a few hundred.

Tenant off-boarding is often overlooked but legally important. GDPR and privacy laws mandate data deletion and portability requests, so on termination the process must completely and verifiably delete that tenant's data (or render it inaccessible by destroying the encryption key) and keep deletion evidence. Deleting only a specific tenant in the Pool model causes a bulk DELETE on a huge table, affecting performance, so Crypto-shredding — rendering data effectively unrecoverable by destroying the per-tenant encryption key — is used as a practical alternative.

3. Core Design Challenges

A. The Noisy Neighbor Problem

The most frequent operational issue when sharing resources in multi-tenancy is the Noisy Neighbor. When one tenant abnormally consumes a lot of CPU, memory, I/O, or connections, response latency and errors for the other tenants sharing the pool surge together. For example, if one tenant repeatedly runs a heavy report query scanning millions of rows, the DB connection pool is exhausted and requests for all tenants pile up in the queue. In several SaaS outage postmortems in the 2010s, the case "a specific large customer's batch job degraded the entire service" was reported repeatedly.

The response is layered. At the application layer, per-tenant Rate Limiting and concurrency quotas cap the request rate and connection count a tenant can consume. At the resource layer, a Cell-based architecture — distributing tenants into multiple independent cells (self-contained resource bundles) so that one cell's failure does not spread to others — is effective. Also, the Bulkhead pattern, allocating a separate connection pool/thread pool per tenant group, can block failure propagation. The important point is that Noisy Neighbor looks like a "performance problem" but is in essence a problem of Fairness and isolation, and the art of the design lies in guaranteeing fairness without giving up sharing efficiency.

B. Preventing Cross-Tenant Data Leakage

The biggest security threat in multi-tenancy is an incident where one tenant gets to see another tenant's data. The causes are mainly: (1) a missing tenant filter in a query, (2) a cache key without the tenant ID (a result queried by tenant A remains in the cache and is returned to B), (3) context contamination (context confusion during asynchronous processing), and (4) an IDOR (Insecure Direct Object Reference) with guessable identifiers. OWASP classifies such tenant-isolation defects as a serious access-control violation.

Defense is composed of "multiple layers of netting." First, enforce the tenant filter at the data layer with the DB RLS explained earlier. Second, always include the tenant ID in the keys/indexes of all secondary stores — caches, search engines, message queues (e.g., the Redis key tenant:{id}:user:{uid}). Third, use unguessable UUIDs as object identifiers and re-verify the owning tenant on access. Fourth, continuously include cross-tenant access attempt cases in automated tests to regression-verify that "requesting B's resource with A's token returns 403." In multi-tenancy, these negative tests are as important as functional tests.

C. Per-tenant Configuration and Customization

SaaS customers each want different screens, fields, workflows, and branding, yet multi-tenancy must maintain a single codebase. Therefore customization is absorbed not by code forking but by a configuration- and metadata-driven approach. Per-tenant settings are kept in an external store and loaded at runtime to branch the UI, validation rules, and feature flags, while custom fields are accommodated without schema changes via EAV (Entity-Attribute-Value) or JSONB columns. The secret to how Salesforce provides different objects, fields, and screens to hundreds of thousands of organizations while maintaining a single platform is precisely this metadata-driven architecture.

Combining this with Feature Flags and pricing Tiers makes it possible to expose a different feature set per tenant from the same code and open advanced features only to higher tiers — a differentiated provisioning of resources and features. This is an architectural element directly tied to the business model (price differentiation) beyond mere customization.

D. Tenant-aware Observability and SLA Differentiation

In multi-tenancy, the tenant ID must be attached as a dimension to all logs, metrics, and traces. If you cannot tell which tenant's request an error on the shared instance originated from, or whether response latency is concentrated on a specific tenant, failure response and root-cause analysis become impossible. Therefore Tenant-aware Observability — always including tenant_id in structured logging, attaching a tenant label to metrics, and propagating a tenant attribute on distributed-tracing spans — becomes a prerequisite of operations. However, when tenants reach the hundreds of thousands, the cardinality explosion of per-tenant metrics drives up monitoring cost, so a selective strategy is needed: collect fine-grained metrics only for higher-tier/large tenants and manage the rest with aggregate metrics.

Observability also connects to SLA (Service Level Agreement) differentiation. In multi-tenancy, promising all tenants the same SLA makes it hard to meet top customers' expectations due to the limits of shared resources. Therefore response-time and availability targets are differentiated by tier, and to back them up, higher-tier tenants are placed on dedicated cells/dedicated connection pools, aligning architecture and contract. For example, a structure where the free tier runs best-effort in a shared Pool while the enterprise tier guarantees 99.95% availability on a dedicated shard is typical.

4. Application Cases and Comparison

A multi-tenancy isolation strategy diverges greatly with an industry's regulatory intensity and customer composition. Collaboration SaaS (Slack, Notion, Salesforce, etc.) must accommodate hundreds of thousands to millions of small tenants at low cost, so it defaults to the Pool model while using a mixed strategy that provides dedicated shards/dedicated regions to large enterprise customers. By contrast, finance and healthcare SaaS often choose isolation close to a Silo due to regulatory/audit requirements, emphasizing physical separation of tenant data and dedicated encryption keys.

Cloud providers are also a giant real-world case of multi-tenancy. AWS, Azure, and GCP are themselves multi-tenant platforms that isolate millions of customers on shared hardware via virtualization/hypervisors, and here, if a customer selects "dedicated hardware (Dedicated Host)," it is a structure trading one more level of stronger isolation for cost. In Kubernetes environments, there is a spectrum of namespace-based soft multi-tenancy (distinguished by namespaces, RBAC, ResourceQuota, NetworkPolicy), cluster-separation-based hard multi-tenancy (a dedicated cluster per tenant), and intermediate approaches such as virtual clusters (vCluster). Because containers share the kernel, the isolation of soft multi-tenancy is weaker than that of virtual machines, and when strong isolation is needed, it is reinforced with sandbox runtimes such as Kata Containers or gVisor.

Viewing economies of scale in numbers makes the motivation for multi-tenancy clear. An organization operating a dedicated stack per tenant (Silo) must patch and monitor 1,000 copies of applications/DBs to accommodate 1,000 customers, and because the resources of small customers with low average utilization mostly stay idle, the infrastructure cost per unit customer stays fixed at a high level. Pool multi-tenancy, by contrast, lets the entire customer base time-share the same resource pool, raising average utilization so that the same hardware can accommodate tens of times more customers. In fact, large collaboration SaaS accommodate millions of tenants in a small set of shared cells, and this density is the source of subscription pricing dramatically lower than on-premises. However, this efficiency rests on the statistical-multiplexing assumption that "one tenant's peak offsets another tenant's slack," so in time windows where all tenants' peaks overlap (e.g., end-of-month settlement), you must also design for the correlated load risk of the entire resource being pressured simultaneously.

The fundamental reason a difference arises in the comparison with single-tenancy (a dedicated stack per tenant) lies in "who bears the operational complexity." Single-tenancy is favorable in isolation, customization, and version independence, but in return the provider bears an operational burden proportional to the number of tenants. Multi-tenancy unifies operations to gain economies of scale, but in return takes on the design difficulty of having to guarantee isolation and fairness in software. Therefore, for "a few large/highly regulated customers," single/Silo is reasonable, while for "many small customers at low cost," Pool multi-tenancy is reasonable, and most mature SaaS combine the two as tiers.

5. Advanced: The SaaS Maturity Model and Serverless Multi-tenancy

The SaaS Maturity Model proposed by Microsoft explains the evolution of multi-tenancy in four levels. Level 1 (Ad-hoc/Custom) is effectively a hosted model with separate code/instances per customer; Level 2 (Configurable) customizes a single codebase by configuration but separates instances; Level 3 (Configurable + Multi-tenant) accommodates multiple tenants on a single instance; and Level 4 (Scalable Multi-tenant) distributes tenants dynamically across many load-balanced instances, aiming for effectively unlimited scaling. This model suggests that "multi-tenancy is not a finished form but a goal reached in stages as the organization matures," and it is useful in a professional-engineer answer as a framework for diagnosing the current level and presenting a target architecture.

The recent trend is moving toward serverless/cell-based multi-tenancy. AWS, in its SaaS design guidance such as the SaaS Lens and reference architectures, recommends "applying a mix of pool/silo/bridge per tenant tier," and with serverless resources such as Lambda and DynamoDB, the resources themselves auto-scale, so Noisy Neighbor is largely absorbed at the infrastructure level. Still, even in serverless, tenant-context propagation and data isolation (including the tenant_id in the DynamoDB partition key, conditional access in IAM policies, etc.) remain the designer's responsibility. Moreover, Cell-based Architecture distributes tenants into self-contained cell units so the blast radius is confined to a cell, drawing attention as a pattern that raises both availability and isolation of large-scale multi-tenant services simultaneously. In generative-AI SaaS, isolating per-tenant embeddings/vector indexes and guaranteeing a data boundary so that tenant data is not mixed into shared model training are emerging as new multi-tenancy challenges.

6. Considerations and Implications

Because multi-tenancy is a strategic choice that determines the economics of SaaS, from a professional engineer's perspective it should be approached not as a single technology but as a design decision frame spanning data, security, operations, and business.

  • Layer-by-layer separated design of the isolation-efficiency trade-off: Avoid the uniform choice of "all Pool" or "all Silo," and adopt a tiered strategy that differs the isolation level per layer — data, compute, network, billing. Lower cost for small customers with Pool, guarantee isolation for large/regulated customers with Silo, and design in advance a tenant promotion (Pool → Silo migration) path to respond to growth.

  • Security: mandating defense in depth and negative tests: Do not rely on a single application filter for isolation; overlay DB RLS, cache-key separation, and non-guessable identifiers. In particular, place "cross-tenant access blocking" as a constant regression test in the CI pipeline so the build fails the moment a code change breaks isolation. Design on the premise that a tenant-isolation defect is a fatal incident that, once it occurs, loses the trust of all customers.

  • Operations: schema evolution and blue-green/progressive deployment: The advantage of a single codebase is "fix once, apply to all," but inverted it is "break once, fail for all." Therefore decompose schema migration into backward-compatible stages (expand → dual-write → switch → clean up), and control risk with canary/ring deployment that rolls out a new version to a few tenants first. Include per-tenant backup/restore (point-in-time) and the possibility of per-tenant rollback in the operations design.

  • Cost visibility and billing linkage (FinOps): Because resources are shared, it is hard to measure "which tenant incurred how much cost." Metering per-tenant resource usage and linking it to Showback/Chargeback and pricing tiers lets you identify cost-anomalous tenants early and design pricing policy based on data. The cost efficiency of multi-tenancy turns into business value only when a measurement/allocation system backs it up.

  • Regulatory and data-sovereignty response: Regulations such as GDPR, privacy laws, network separation, and CSAP require data location, deletion, and separation. Per-region tenant placement, per-tenant encryption keys and crypto-shredding, and retention of deletion evidence must be reflected at the start of design; retrofitting is more costly the more it is a Pool model. As an outlook, guaranteeing the data boundary of AI SaaS and cell-based/serverless multi-tenancy will emerge as next-generation challenges demanding isolation and scaling simultaneously.

References


In one line: Multi-tenancy is the core SaaS architecture that serves many tenants on shared infrastructure while pursuing isolation and efficiency at once; success hinges on combining Pool/Bridge/Silo isolation levels by tier across the data, security, operations, and billing layers and controlling data leakage and the Noisy Neighbor with RLS, Rate Limiting, and cell-based design.