Distributed Unique ID Generation Strategies (UUID, ULID, Snowflake ID)
1. Overview
A. Definition
A Distributed Unique ID is an identifier-generation technique designed to guarantee global uniqueness without relying on a centralized number issuer, in environments where many nodes, services, and data centers create records simultaneously. Representative examples are the random-based UUID, the time-sortable ULID and UUIDv7, and the bit-composition-based Snowflake ID.
Distributed unique IDs arise because the most basic question in database design—"what primary key should I assign to newly created data?"—resurfaces as a hard problem once we move past the single-DB era into microservices, sharding, and multi-region environments. In the past, an RDBMS AUTO_INCREMENT or sequence object was effectively the standard, but today, when dozens of horizontally partitioned shards and geographically distributed write nodes perform inserts concurrently, a single issuing point itself becomes a bottleneck and a single point of failure (SPOF). An identifier strategy is therefore treated not as a mere implementation detail but as an architectural decision that governs scalability, availability, index performance, and security.
B. Background and Necessity
The most familiar issuing scheme, the DB auto-increment key, is concise in a single-node environment and offers excellent index locality. But the moment sharding is introduced, a fatal limit appears. If shard A and shard B each number from 1, keys collide; to avoid that, placing a central sequence server forces every insert through one point, so throughput is bound to that server's limit and network round-trip latency is added on top. Extending to multi-region makes cross-region coordination costs explode, growing write latency to tens of milliseconds.
This problem recurs frequently on the Professional Engineer exam, transformed into "data identifier design for MSA / large-scale distributed systems" or "primary-key strategy for sharded environments," and it demands not rote memorization but an essay that explains trade-offs. One must therefore be able to explain each scheme's pros and cons starting from its uniqueness-guarantee mechanism.
In addition, auto-increment keys carry the security weakness of predictability. Since it is obvious that /orders/1002 follows /orders/1001, lax authorization exposes the system to IDOR (Insecure Direct Object Reference) attacks that sequentially change the identifier to read others' resources, and it also enables enumeration-based information leakage such as a competitor estimating sales volume from "this month's order number minus last month's." Distributed unique IDs emerged to satisfy three demands simultaneously: (1) coordination-free uniqueness, (2) high generation throughput, and (3) (when needed) unpredictability—with (4) time sortability (index locality) being the key challenge of modern designs.
2. Classification of Generation Schemes and Requirements
Distributed ID design splits into three branches by "how uniqueness is secured." The first is random/hash-based, represented by UUIDv4, which lowers collision probability to a negligible level using a sufficiently large random space. The second is time + random combination, placing a millisecond timestamp in the high bits so that IDs sort roughly by creation order while the low bits are filled with randomness—ULID and UUIDv7. The third is time + node + sequence composition, assigning each node a unique number to achieve uniqueness without coordination—the Snowflake family.
flowchart TB
Root["Distributed Unique ID Strategies"]
Root --> A["Random/hash-based(no coordination)"]
Root --> B["Time+random(sortable)"]
Root --> C["Time+node+seq(semi-coordinated)"]
A --> A1["UUIDv4(122-bit random)"]
A --> A2["UUIDv1(time+MAC)"]
B --> B1["ULID(48-bit time+80-bit random)"]
B --> B2["UUIDv7(RFC 9562)"]
C --> C1["Twitter Snowflake(64-bit)"]
C --> C2["Instagram/Sonyflake variants"]
The root reason these branches diverge is the difference in the uniqueness-guarantee mechanism. The random-based approach relies on a statistical guarantee that "probabilistically they almost never overlap," while the composition-based approach relies on a structural guarantee that "if node numbers differ, they can never overlap." This difference leads directly to differences in the need for node coordination, key length, and sortability.
The requirements a good distributed ID should meet can be summarized as follows; they stand in mutual trade-off, so one scheme rarely satisfies all of them.
| Requirement | Meaning | Especially important when |
|---|---|---|
| Global uniqueness | No collision across all nodes/times | Sharding, multi-region writes |
| Generation throughput | IDs generatable per second | Bulk orders, logs, messages |
| Time sortability | Roughly increasing by creation order | B-tree index insert performance |
| Unpredictability | Next value cannot be guessed | Public URLs, security identifiers |
| Compactness/length | Storage/transfer cost | Index size, network |
3. UUID and ULID — Coordination-free Schemes
A. Characteristics by UUID Version
A UUID (Universally Unique Identifier) is a 128-bit value, usually written as a 36-character string of 32 hexadecimal digits plus 4 hyphens. It was standardized in 2005 as RFC 4122, and in May 2024 RFC 9562 revised and replaced it while adding new versions. The most widely used, UUIDv4, fills 122 bits with randomness excluding the 6 bits used for version/variant identification. This space reaches about 5.3×10³⁶, so even generating a billion per second for 85 years, the probability of a single collision stays extremely low. Because no inter-node communication is needed, coordination cost is zero, and the decisive advantage is that it can be generated instantly even on an offline client.
Another reason UUIDv4 is widely used is the maturity of the ecosystem. Almost every language's standard library and database provides a native UUID type and generation function, so it can be adopted instantly with no extra infrastructure, and it suits especially well identifiers that are "short-lived and impossible to coordinate," such as distributed trace IDs, idempotency keys, and event IDs. In short, unless it is a high-volume primary key where performance is extremely critical, v4 remains the safest and most unremarkable default choice.
By contrast, UUIDv1 combines a 100-nanosecond-unit timestamp with the network card's MAC address. Because it contains time information, it is in theory sortable, but since the bit layout is not time-ordered from the high bits, creation order is not revealed by string sorting, and the privacy problem that the generating device can be identified through MAC-address exposure has long been noted (the past tracing of the Melissa virus is a representative case). For this reason v4 is used for public identifiers, while v1 is used cautiously for internal purposes where traceability is not needed.
To look at collision probability a bit more concretely, UUIDv4's uniqueness is gauged by the birthday problem. For a collision to first occur with 50% probability in a 122-bit space, roughly 2^61—about 2.3×10¹⁸—must be generated. This is a scale hard to reach even if the whole world churns out IDs explosively for centuries, so in practice UUIDv4 collisions are designed around as "not happening." However, this guarantee depends entirely on the quality of the entropy source, so one must note that generating with a weak pseudo-random generator, or in an entropy-starved state right after a virtualized environment boots, sharply raises the actual collision risk.
UUIDv4's hidden weakness is the absence of index locality. Because the value is fully random, each insert into a B-tree index lands at a random position in the tree, triggering page splits and buffer cache misses. In engines like MySQL InnoDB where the primary key of a clustered index determines physical ordering, a random UUID primary key is well known from measurements to greatly degrade insert performance and cache hit rates, so in large tables it is avoided or replaced by the sortable schemes described later. Many benchmarks have reported that on tens-of-millions-of-rows tables a random UUID primary key dropped insert throughput several-fold versus a sequential key, so this problem is treated as a real field bottleneck, not theory.
B. ULID and UUIDv7 — Time-sortable IDs
What emerged to target this index problem is ULID (Universally Unique Lexicographically Sortable Identifier). ULID composes 128 bits as a high 48-bit millisecond timestamp + low 80-bit randomness and encodes it in Crockford Base32 as a 26-character string. Because time comes first, sorting the string lexicographically yields creation-time order, and within the same millisecond the 80-bit randomness secures uniqueness. As a result, it keeps UUIDv4's advantage of coordination-free generation while inserting into indexes in a nearly monotonically increasing form, so page splits plummet.
UUIDv7, newly introduced by RFC 9562, implements essentially the same design philosophy as ULID within the standard UUID format, placing a Unix millisecond timestamp in the high 48 bits and filling the rest with randomness. Because it can obtain time sortability while using the existing UUID storage and library ecosystem as-is, it is quickly establishing itself as the default alternative to v4 in new designs. However, since time information is exposed, it is unsuitable for security identifiers that must hide "when it was created," in which case the fully random v4 is still chosen. In other words, sortability and unpredictability cannot be simultaneously maximized, and the version must be chosen by use case.
Even time-sortable IDs like ULID/UUIDv7 have a pitfall to watch. Because the mutual order of IDs generated within the same millisecond is determined by the low-order randomness, sub-millisecond creation order is not guaranteed. If strict monotonic increase is required, one must use a variant like ULID's "monotonic mode," which within the same millisecond adds 1 to the previous value instead of randomness. Also, a 48-bit millisecond timestamp can represent about 8,900 years, so there is virtually no lifetime concern, but the fact that time sortability depends on the client clock is a weak link in distributed generation. The global order of IDs generated on different machines can be trusted only as much as each machine's clock accuracy, so sorting must not be taken as a strong premise of business logic.
4. Snowflake ID — 64-bit ID Based on Central Coordination
A. 64-bit Bit Structure
Snowflake ID was devised by Twitter in 2010 because it needed an identifier that was sortable yet compact at 64 bits in a distributed environment. To solve both problems at once—128-bit UUIDs being long and incurring high index/network cost, and DB sequences being a central bottleneck—it splits a 64-bit integer into four regions as below.
flowchart LR
S["Sign 1 bit(always 0)"] --> T["Timestamp 41 bits(ms, ~69 yrs)"]
T --> W["Worker ID 10 bits(1024 nodes)"]
W --> Q["Sequence 12 bits(4096/ms)"]
From the top it consists of 1 sign bit (0 to ensure positivity), a 41-bit timestamp, a 10-bit worker ID, and a 12-bit sequence. The 41-bit millisecond timestamp can represent about 2^41 ms—i.e. about 69.7 years—from a reference point (epoch), so setting a custom epoch at service start covers several decades. The 10-bit worker ID distinguishes up to 1024 nodes, and the 12-bit sequence lets one node issue up to 4096 (2^12) IDs within the same millisecond. As a result, in theory the whole system can generate about 1024×4096 ≈ 4.19 million per millisecond, roughly 4.2 billion IDs per second, without coordination.
B. Generation Procedure and Failure Handling
The generation process works as follows. The node reads the current time and fills the timestamp field; if it is the same millisecond as the previous issuance, it increments the sequence by 1; and if the sequence exceeds 4096, it briefly busy-waits until the next millisecond arrives. When the millisecond changes, it resets the sequence to 0. Because time is in the high bits, generated IDs increase roughly in time order globally, and within the same node they are fully monotonically increasing. Thanks to this, it enjoys index locality like ULID while being half the size of a UUID at 64 bits.
Snowflake's core premises are the uniqueness of worker IDs and the monotonicity of the clock. Since two nodes using the same worker ID collide, at node startup a non-duplicate worker ID must be allocated from ZooKeeper, etcd, a DB, etc. This is the point that, unlike the fully coordination-free UUID, requires lightweight coordination at startup, so it is classified as a "semi-coordinated" scheme. Also, if NTP clock regression or a leap second moves the timestamp backward, duplicates can occur, so practical implementations place defensive logic that stops ID issuance, or waits within an error margin, when the clock falls behind the last issuance time.
C. Flexibility of Bit Allocation and Operating Environment
It is also important that the bit allocation itself is a design parameter. How many bits to give to timestamp, worker, and sequence is a choice reflecting a "lifetime vs. node count vs. issuance-per-second" trade-off, and Twitter's 10/12 allocation is a balance point that "1024 nodes are enough and per-node issuance is maximized." Conversely, in an environment with tens of thousands of nodes, one must reallocate by increasing worker bits and decreasing sequence bits, or lowering time resolution. In other words, Snowflake is accurately understood not as a single formula but as a design framework that adjusts bit widths to the organization's traffic profile. In environments like Kubernetes where nodes constantly appear and disappear, fixed worker-ID allocation is hard, so variants that dynamically borrow the worker ID from the sequence space or assign it by hashing pod information, or hybrid designs combined with a central segment allocator (e.g. Leaf-segment), are considered together.
5. Comparison and Cases
A. Comparison of Characteristics by Scheme
The choice among the three schemes is a question of "what to give up." As the comparison below shows, one cannot simultaneously have full coordination-freedom and compact 64-bit sortability.
| Category | UUIDv4 | ULID / UUIDv7 | Snowflake |
|---|---|---|---|
| Length | 128-bit / 36 chars | 128-bit / 26 chars(ULID) | 64-bit |
| Uniqueness guarantee | Probabilistic(random) | Probabilistic(time+random) | Structural(node+seq) |
| Sortability | None | Time-sorted | Time-sorted |
| Node coordination | Not needed | Not needed | Worker ID at startup |
| Unpredictability | Very high | Low(time exposed) | Low(time·node exposed) |
| Clock dependence | None | Yes | Strong(needs regression defense) |
Real-world cases illustrate this trade-off well. First, Instagram in its early sharding environment generated 64-bit IDs via a PostgreSQL stored procedure, designing a combination of a high 41-bit time, a middle 13-bit logical shard ID, and a low 10-bit per-shard sequence, thereby gaining the routing advantage that "you can tell which shard it is on just from the ID." This is a representative case of applying Snowflake thinking to shard identification. Second, Discord adopted Twitter's Snowflake but used 2015-01-01 as a custom epoch, and uses the time sortability of message IDs to page "messages before/after a given time" with ID-range queries rather than a separate index. Third, Sony (Sonyflake) redesigned the bit allocation to represent about 174 years by lowering time resolution to 10 ms units while expanding node count to up to 2^16—a variant fit for environments with "very many nodes but relatively low issuance per second." It is in the same vein that many services at home and abroad run in-house libraries such as Baidu UidGenerator and Meituan Leaf that combine Snowflake with segment allocation and clock correction.
Fourth, MongoDB's ObjectId is an interesting compromise case in which all three design philosophies are fused into one identifier. It consists of 12 bytes (96 bits): the high 4 bytes are a second-granularity Unix timestamp, the middle 5 bytes are randomness distinguishing process/machine, and the low 3 bytes are an incrementing counter. Because time is in the high part it achieves rough time sorting (the advantage of ULID/Snowflake), the middle randomness avoids collisions between different instances without node coordination (the advantage of UUID), and the low counter guarantees uniqueness within the same second. Though shorter than a 128-bit UUID and longer than a 64-bit Snowflake, in that it has no central coordination at all, it is widely used as a practical design realizing "full coordination-freedom + time sorting" within 96 bits. This shows that real-system identifiers are usually not a single pure type but a mixture of several philosophies.
6. Deep Dive — Index Locality and the UUIDv7 Standardization Trend
The hottest issue in recent identifier design is database index locality. In write-heavy large tables, a random UUIDv4 primary key causes inserts at random points in the B-tree, noticeably lowering TPS through disk write amplification and cache misses. Conversely, a monotonically increasing key inserts only at the right end of the index for high cache efficiency, but multiple nodes crowding the same hot page can cause lock contention (hotspot). So the design balance point is "neither fully random nor fully sequential, but a time prefix + random tail," and this is the technical basis for the attention ULID and UUIDv7 receive.
Around this issue, various variants besides UUID/Snowflake have been proposed. KSUID combines a 32-bit second-granularity timestamp with 128-bit randomness to make 160 bits, encoded in Base62 to be URL-friendly and time-sortable, while NanoID is a random ID focused on short, URL-safe strings rather than uniqueness, popular in frontend/short-URL areas. All of them share the common skeleton of "time prefix + sufficient randomness," but choose different points in length, encoding, and the degree of sortability guarantee. In other words, identifier design is a problem of choosing coordinates on a few axes (sortability, length, unpredictability, coordination frequency), and whichever library one picks, one must first check whether those coordinates match one's requirements.
An advanced issue from an operations view is the observability of ID generation itself, which cannot be left out. If a Snowflake node's clock regresses and ID issuance stops, or if a particular millisecond exhausts its 4096 sequence values and the wait lengthens, that itself becomes a cause of latency spikes. Therefore mature operations organizations collect per-node issuance QPS, sequence-saturation counts, clock-regression detection events, and worker-ID allocation collisions as metrics and raise early alerts. Also, the point at which the custom epoch exhausts its 41 bits (about 69 years after service start) may look far off now, but it is a design debt that "will surely come someday," so explicitly noting the exhaustion year in the design document is a responsible design attitude.
On the standardization side, RFC 9562, published in May 2024, replaces the old RFC 4122 and formalizes UUIDv6 (v1 rearranged to be sortable), v7 (Unix-time-based), and v8 (custom). This is a trend to absorb the "time-sortable ID" area—where non-standards like ULID, KSUID, and Snowflake had proliferated—into the standard UUID system, and PostgreSQL, major ORMs, and language standard libraries are rapidly adding UUIDv7 support. From a Professional Engineer perspective, the key outlook is "whether the default primary key for new distributed systems will shift from v4 to v7, and whether Snowflake will keep coexisting in areas where 64-bit compactness matters." In addition, it is becoming a security best practice to use v4 for public API identifiers so that time/node information is not exposed, or to wrap an internal Snowflake once more in a hashed (HMAC) public token or Hashids to block enumeration attacks.
7. Considerations and Implications
From a Professional Engineer perspective, when selecting and adopting a distributed unique ID strategy, the following should be examined comprehensively.
- Adopting multiple strategies by use case: Since no single scheme satisfies all demands, it is realistic to separate the design—internal primary keys as Snowflake/UUIDv7 with good index locality, and externally exposed public identifiers as highly unpredictable UUIDv4 or separate tokens. Actively consider a dual-identifier pattern in which a single entity holds both an internal key and a public key.
- Index/storage cost trade-off: A 128-bit UUID uses twice the space of a 64-bit Snowflake in the primary key, all foreign keys, and indexes, and incurs higher join/network costs. For tables reaching billions of rows, the benefit of 64-bit compactness is substantial, so the uniqueness-guarantee scheme and length must be weighed together.
- Clock reliability and failure preparedness: Snowflake, ULID, and UUIDv7 all depend on the system clock, so NTP synchronization quality, clock-regression defense, and the time of custom-epoch exhaustion (41-bit depletion) must be reflected in operational design. Do not forget that the availability of the worker-ID allocation system (ZooKeeper, etcd, DB) is also a prerequisite for ID generation.
- Security/privacy impact assessment: Time/node/MAC/shard information embedded in an identifier can itself be metadata leakage. To defend against IDOR/enumeration attacks, do not let authorization depend solely on identifier unpredictability, and when necessary randomize/hash public identifiers.
- Migration strategy: When converting a system already operating on auto-increment keys to distributed IDs, a strangler approach—keeping the existing key while introducing a new identifier column in parallel and gradually migrating references—is safe. Because a key-type change ripples through all foreign keys and application contracts, a staged transition plan is essential.
- The double edge of sorting/paging usage: Time-sortable IDs give the powerful benefit of implementing time-series paging with just "ID-range queries," but overusing this to make the ID the sole basis for sorting/time comparison falls into the pitfalls of clock error and the lack of sub-millisecond order guarantee. If creation order matters to the business, it is more robust to keep a separate creation-time column or a logical sequence alongside, separating the identifier from the ordering basis.
- Standard convergence and long-term compatibility: Considering the trend of RFC 9562 absorbing time-sortable IDs into the UUID standard, new systems may find adopting standard UUIDv7 advantageous over in-house custom formats for ecosystem compatibility and long-term maintenance. However, since the Snowflake family remains valid in areas where 64-bit compactness is decisive, rather than blindly following standard convergence, choose by re-evaluating storage cost, sortability, and security requirements.
References
- RFC 9562, "Universally Unique IDentifiers (UUIDs)", IETF, 2024. https://www.rfc-editor.org/rfc/rfc9562
- Twitter Engineering, "Announcing Snowflake", 2010. https://blog.twitter.com/engineering/en_us/a/2010/announcing-snowflake
- Instagram Engineering, "Sharding & IDs at Instagram". https://instagram-engineering.com/sharding-ids-at-instagram-1cf5a71e5a5c
- ULID specification (spec repository). https://github.com/ulid/spec
In one line: A distributed unique ID is a design for guaranteeing global uniqueness without a central issuer; representative schemes are the fully coordination-free, unpredictable UUIDv4, the ULID/UUIDv7 that gain index locality through time sorting, and the Snowflake that provides 64-bit compactness with semi-coordinated uniqueness—and the key is to separate the design by use case according to the trade-offs of uniqueness-guarantee method, sortability, length, and security.