← Back to list
Networking
#시간동기화#NTP#PTP#TrueTime#분산시스템
Last updated · 2026-10-04

Clock Synchronization in Distributed Systems (NTP·PTP·TrueTime)

1. Overview

A. Definition

Clock synchronization is the collective term for the protocols and algorithms that align the physical (wall-clock) clocks of many independently running computers and network devices to a single reference time (e.g., UTC), or that keep the time difference (offset) between nodes within a known bound, so that a distributed system can treat time as a trustworthy resource. Representative technologies are NTP (Network Time Protocol), which aligns the whole Internet to millisecond level; PTP (IEEE 1588), which aligns LANs and industrial networks to microsecond–nanosecond level; and Google Spanner's TrueTime, which exposes the uncertainty of time as a number and uses it to order transactions.

Clock synchronization looks like a simple matter of "setting clocks," but it is in fact a hidden foundation of infrastructure that underpins everything from correlating distributed logs, the validity of security certificates·OTP·Kerberos tickets, the execution order of financial trades, radio-frame alignment of 5G base stations, to phasor measurement in the power grid. If log timestamps disagree, root-cause analysis becomes impossible, certificate validation misfires, and the order of trades is reversed, causing regulatory violations.

B. Background and Necessity

Every computer has an internal clock based on a quartz oscillator, but the frequency of that oscillator varies slightly with temperature, voltage, and aging. As a result the clock runs a little fast or slow relative to the reference time; this rate of change is called drift, and a typical commercial quartz clock drifts by several seconds per day — on the order of tens of parts per million (tens of ppm). With no synchronization at all, within a few days the inter-node time difference reaches several seconds, shaking every decision that relies on time.

The reason the problem is not merely "the clock is wrong" is that in distributed systems time is used as a premise of correctness and security. TLS·code-signing certificates validate their validity period by time, Kerberos·TOTP reject authentication when the time difference exceeds the allowed window (typically ±5 minutes), and a distributed database's LWW (Last-Write-Wins) decides the winner by timestamp magnitude. When the clock is off, valid certificates are rejected or expired tickets pass, and freshly written data is overwritten by an older value — a lost update. Thus "how precisely and how reliably we align time across nodes" becomes a design problem that simultaneously governs availability, security, and data consistency.

Concrete damage cases are not rare either. During the 2012 leap-second insertion, a kernel bug sent CPUs on many Linux servers into a runaway, and large services such as Reddit·LinkedIn failed simultaneously; in 2016, a misconfiguration on a public NTP server rolled some devices' time back into the past, so that certificates were judged "not yet valid" and connections dropped. Time thus has the character of a hidden single point of failure (hidden SPOF) — invisible in normal times, but the moment it slips it surfaces as a sweeping, simultaneous outage.

C. Key Characteristics

The nature of clock-synchronization technology can be summarized in three points. First, convergence to a reference time — hierarchically aligning to a primary reference clock such as GNSS·atomic clocks·radio signals that track UTC. Second, estimation and correction of delay — measuring network round-trip delay to estimate one-way delay, and accepting the asymmetry inherent in a message round trip as a correction error. Third, a precision-cost trade-off — software-timestamp NTP is cheap but stays at millisecond level, whereas hardware-timestamp·dedicated-equipment PTP achieves nanosecond level at a high cost. TrueTime in particular is decisively different in that, through the conceptual shift of "exposing time not as a single point but as an uncertainty interval [earliest, latest]," it does not hide the error but lets the algorithm handle it explicitly.

The common challenge running through all three properties is preserving monotonicity. If a synchronization correction rolls time back into the past, "a time smaller than a timestamp just taken" appears, distorting log order·expiry decisions·cache TTLs. So in practice the convention is to correct only by slewing so that time never runs backward, and to provide the application a separate monotonic clock that never goes back, for measuring elapsed time. In the end, good clock synchronization must satisfy two demands at once: "accurate absolute time" and "relative time that never runs backward."

2. The Principle of Physical-Clock Error and Synchronization Models

The starting point for understanding clock synchronization is to distinguish the three elements that make up physical-clock error. Offset is the difference from the reference time at a given instant, drift/skew is the speed at which the clock diverges (error per second), and jitter is the short-term fluctuation of measured values. Synchronization is, in the end, a feedback-control process that periodically measures and undoes the offset (step/slew) and estimates the drift to discipline the oscillator frequency. The step method, which jumps time all at once, breaks monotonic increase and causes timestamp reversal, so operational systems usually prefer slew, gently speeding up or slowing down the clock.

Synchronization is divided, by whether a reference exists, into external synchronization and internal synchronization. The former aligns to an absolute reference such as UTC — NTP·PTP belong here — while the latter, with no external reference, has nodes average among themselves to reduce only their mutual difference, the Berkeley algorithm being representative. Also, as a delay-correction method, Cristian's algorithm assumes half of the round-trip time RTT to be the one-way delay and adds it to the server time; this "half the round trip = one way" assumption breaking down in the face of path asymmetry (up/down delays differ) is the fundamental error source of all network-based synchronization.

flowchart TB
    subgraph Ref["Reference time hierarchy"]
      UTC["Coordinated Universal Time UTC"] --> ATOMIC["Atomic clock·GNSS(GPS)"]
      ATOMIC --> PRIM["Primary reference clock(Stratum 0)"]
    end
    subgraph Err["Physical-clock error factors"]
      OSC["Quartz oscillator"] --> DR["Drift(temperature·voltage·aging)"]
      DR --> OFF["Offset(difference from reference)"]
      OSC --> JIT["Jitter(short-term fluctuation)"]
    end
    subgraph Sync["Synchronization methods"]
      EXT["External sync(UTC reference): NTP·PTP"]
      INT["Internal sync(mutual average): Berkeley"]
      EST["Delay estimation: Cristian(RTT/2 assumption)"]
    end
    PRIM --> EXT
    OFF --> EXT
    OFF --> INT
    EXT --> ASY["Path asymmetry → correction error"]
    EST --> ASY

The Berkeley and Cristian algorithms are the two axes showing how this model was historically implemented. The Cristian method is a client-server model that queries one trusted time server and receives its time corrected by half the RTT — the direct ancestor of today's NTP. The Berkeley method, by contrast, is an active internal-synchronization scheme in which, in an environment with no reference server, a coordinator (master) collects every node's time, averages it, excludes outliers, and returns to each node a correction amount of "speed up or slow down by this much." The former suits cases that need absolute time, the latter closed networks that need only relative consistency among nodes with no external reference.

As this figure shows, the quality of every synchronization technology is ultimately determined by how accurately·symmetrically it measures the delay of the network path. The difference among NTP·PTP·TrueTime reduces to the difference in where this measurement is done (software vs hardware), under what assumption (average vs actual measurement), and against what reference (a single point vs an interval).

3. NTP (Network Time Protocol) — the Internet-Standard Synchronization

NTP, designed by David Mills in 1985 and standardized as RFC 5905 (NTPv4), is the most widely used time-synchronization protocol on the Internet. NTP organizes time servers into a hierarchy called stratum. Stratum 0 is the reference device itself, such as an atomic clock·GPS; Stratum 1 is a server directly connected to it; Stratum 2 and below are servers that reference an upper server, and error accumulates as the layers descend. A client cross-validates the responses of several upper servers to filter out falsetickers and selects the most trustworthy combination.

The core of NTP is computing offset and delay from four timestamps. If the client sends a request at T1, the server receives it at T2 and responds at T3, and the client receives it at T4, then the round-trip delay δ = (T4−T1)−(T3−T2) and the offset θ = ((T2−T1)+(T3−T4))/2 are estimated. This formula stands on the symmetric assumption that the up/down one-way delays are equal, so on an asymmetric path an error of half that difference remains. This is why public-network NTP typically stays at several to tens of milliseconds.

sequenceDiagram
    participant C as Client
    participant S as NTP server(Stratum 2)
    Note over C,S: Derive offset·delay from four timestamps
    C->>S: Send request(record T1)
    Note right of S: Received at T2
    Note right of S: Response generated at T3
    S-->>C: Response(includes T1,T2,T3)
    Note over C: Received at T4
    Note over C: Delay δ = (T4-T1)-(T3-T2)
    Note over C: Offset θ = ((T2-T1)+(T3-T4))/2
    Note over C: Gently correct the clock by θ(slew)

The more fundamental reason NTP stays at millisecond level lies in where the timestamps are taken. Because the request·response times (T1–T4) are recorded in the application·operating-system software layer, all of the time a packet sits in a kernel queue or interrupt handling is delayed gets mixed in as measurement error. This software-stack delay is variable from hundreds of microseconds to several milliseconds and is asymmetric up/down, so no matter how often one measures it is hard to break the millisecond wall. Even so, NTP operates over the existing network across wide areas with no dedicated hardware and filters out false sources with algorithms proven over decades, so it remains the de facto standard for Internet time.

NTP's long-standing weakness was security. Being plaintext-UDP-based, a man-in-the-middle manipulating time can disrupt certificate expiry·Kerberos authentication, and in the past the monlist feature was abused for large-scale reflection·amplification DDoS. To resolve this, NTS (Network Time Security) was standardized in 2020 as RFC 8915, exchanging keys over TLS and attaching an authentication tag to NTP packets to verify the integrity·origin of time. In the public·financial domains, applying NTS and multiplexing NTP servers have become de facto mandatory security controls. On the operational side too, rather than relying only on the public pool.ntp.org, it is recommended for both reliability and security to place a national time-standards body (KRISS in Korea) or an own Stratum 1 server as the primary and configure cross-referencing of multiple uppers.

4. PTP (IEEE 1588) — the Precision Time Protocol

What emerged for the industrial·telecom·financial domains, for which milliseconds are not enough, is PTP (Precision Time Protocol, IEEE 1588). PTP uses the same message-exchange principle as NTP but achieves microsecond–nanosecond precision through two decisive differences. First, the timestamp is taken by the hardware of the NIC·switch at the instant the packet leaves the physical layer (hardware timestamping), not by OS software — removing the largest error source, the queuing·interrupt delay of the software stack. Second, switches on the path act as a Transparent Clock, adding the residence time for which they held the packet into a correction field, thereby removing the variable delay from network congestion.

PTP-capable network equipment handles delay in two roles. The Transparent Clock measures the residence time inside a switch and accumulates it into a correction field, canceling the variable delay from congestion, while the Boundary Clock becomes the slave of one segment and the master of the next, cutting error accumulation off segment by segment. Passing through an ordinary switch that lacks these two mechanisms leaves queuing delay as error directly, so PTP's nanosecond precision is realized in full only when every device on the path supports PTP.

PTP automatically elects the best reference in a domain, the Grandmaster, with the BMCA (Best Master Clock Algorithm), and a Boundary Clock relays it to lower segments. Sync·Follow_Up·Delay_Req·Delay_Resp messages separately measure the master-slave offset and path delay, and thanks to hardware timestamps tens of nanoseconds of precision is possible within a LAN. In return, PTP yields its best performance only when every switch on the path supports PTP, so dedicated-hardware investment and network-design constraints follow.

In numbers, the gap between the two protocols is clear. NTP on the public Internet typically stays at 1–50 ms, and even on a well-managed LAN at hundreds of μs–1 ms, whereas PTP with hardware timestamps and transparent clocks achieves tens to hundreds of ns on the same LAN. This nearly 10,000-fold difference comes not from the algorithm but from where the timestamp is taken and whether the path equipment supports it. That is, PTP's precision is a product of investment in the entire hardware ecosystem of NIC·switches, not just protocol design.

Real application cases clearly show the value of the technology. 5G mobile communications require synchronization within ±1.5 μs to align TDD frames among base stations and so use PTP (ITU-T G.8275 profile); stock exchanges adopted PTP as Europe's MiFID II regulation mandated UTC tracking within 100 μs and timestamp recording for high-frequency trading (HFT). In the power grid, PMUs (Phasor Measurement Units) use PTP·GPS synchronization to measure the phase of 60 Hz AC at several-microsecond precision.

5. TrueTime and Time in Distributed Databases

Google Spanner's TrueTime is an original approach that couples clock synchronization with database consistency. Where traditional synchronization "tells you time as a single point but hides the error," TrueTime places a GPS receiver and an atomic clock together in each data center, and the API returns time as an uncertainty interval TT.now() = [earliest, latest]. That is, it exposes the error ε as a number, saying "the true time right now is somewhere in this interval." The reason for using GPS and atomic clocks together is that GPS is weak to antenna failure·radio jamming while atomic clocks have long-term drift, so they cross-compensate each other's weaknesses. In an operational environment ε is usually in the 1–7 ms range.

What matters here is that the size of ε directly relates to performance. A large ε lengthens the commit-wait explained below, increasing write latency, while a small ε raises throughput accordingly. So Google invested heavily in densely deploying GPS·atomic clocks in each data center and synchronizing at a short period to hold ε down to a few milliseconds. This demonstrates a design philosophy in which hardware and algorithms interlock — "capital investment in precision-time infrastructure is itself the performance of the distributed DB."

The device that leverages this uncertainty is commit-wait. After assigning a transaction the timestamp s, Spanner deliberately waits to commit until TT.now().earliest > s, that is, until the uncertainty interval has fully passed s (at most 2ε). Then any transaction that starts later is guaranteed to receive a larger timestamp, guaranteeing external consistency (a strong form of linearizability) even in a physically distributed environment. In the end, instead of trying to eliminate clock error, TrueTime is a strategy that buys consistency by making the error measurable and waiting out exactly that much, and in that investment in precision-time infrastructure (the smaller ε, the shorter the wait) leads directly to performance, it serves as a bridge joining physical clocks and distributed algorithms.

This approach later influenced the open-source camp as well. CockroachDB·YugabyteDB, with no dedicated atomic clock, set an NTP-based maximum error bound (max offset) as a configuration value and approximate TrueTime's idea on commodity hardware by retrying reads (read restart) in the uncertain interval that exceeds that range. This is a good contrasting case showing the trade-off between "precision-time investment" and "algorithmic correction," clearly revealing that the poorer the time infrastructure, the more retry·wait cost the algorithm has to bear.

6. Comparison — NTP·PTP·TrueTime

The three technologies fundamentally differ in "what they guarantee and to what extent." NTP aims at low cost·wide area, PTP at high precision·short range, and TrueTime at explicit use of error. The table below is a supplementary summary; the essence of the choice lies in the balance of the precision requirement and the infrastructure cost it incurs.

Category NTP (RFC 5905) PTP (IEEE 1588) TrueTime (Spanner)
Typical precision several ms ~ tens of ms tens of ns ~ several μs ε several ms(interval exposed)
Timestamp software hardware(NIC·switch) GPS+atomic clock
Scope Internet·WAN LAN·industrial·telecom global distributed DB
Key assumption/device RTT symmetry, stratum transparent clock·BMCA uncertainty interval·commit-wait
Cost low high(dedicated HW) very high(dedicated time infra)
Security NTS(RFC 8915) per-profile authentication internal control

It is also important that the three technologies coexist hierarchically rather than competing. Even when a global distributed DB uses TrueTime, its reference GPS·atomic clocks are physical time, data-center servers still keep microsecond sync via PTP, and office·development environments keep millisecond sync via NTP. Practical architecture is generally designed as a precision pyramid that places a GNSS·atomic-clock reference at the top, PTP in the core network·data center, and NTP outside that. Thus a more accurate design question than "which to choose" is "what precision of time to supply to each layer and how to secure resilience."

The reason the difference arises lies in where the error source is removed. NTP's largest error source is the variable delay of the OS software stack, which PTP strips away with hardware timestamps. TrueTime, rather than reducing the network-measurement error itself, places the reference directly in each data center with GPS·atomic clocks, cutting off dependence on the network path. Thus "logs·authentication for which milliseconds suffice" resolve to NTP, "telecom·finance·power where microseconds are life-or-death" to PTP, and "a global strongly-consistent DB" to TrueTime-style investment.

7. Deep Dive — Latest Trends and PNT Resilience

The biggest topic in the clock-synchronization field lately is the risk of GNSS (satellite navigation) dependence and its alternatives. As finance·telecom·power have come to depend broadly on GPS time, concern has grown that if time is contaminated by jamming·spoofing, it could spread into a wide-area outage. In response the United States, through a PNT (Positioning, Navigation, Timing) resilience executive order, requires GPS-independent backup (terrestrial eLoran, etc.) and anomaly detection, and the EU too has begun to address the time resilience of critical infrastructure by regulation. From a design viewpoint, holdover (self-sustaining when the reference is lost) design — multiplexing GPS·NTP·PTP·atomic clocks and isolating false sources through cross-validation among time sources — is becoming the standard.

The risk of GNSS spoofing is not hypothesis but reality. Cases of GPS position·time being manipulated on ships·aircraft are reported steadily, centered on conflict zones, and since telecom·finance share the same GPS time, a wide-area time disruption can spread into secondary damage. For this reason critical infrastructure uses different satellite constellations (GPS·Galileo·BeiDou) together with terrestrial references, and is designed to automatically isolate a source when the time between sources diverges sharply.

At the protocol level, security hardening and ultra-precision advance simultaneously. NTP secured integrity with NTS (RFC 8915), and inside data centers there is active movement to spread sub-microsecond sync to commercial data centers, as in the large-scale PTP-based operation case published by Facebook (Meta). White Rabbit (a PTP extension), which started on particle-physics research networks, realized sub-nanosecond sync and was reflected into the next-generation PTP high-precision profile, and 5G-Advanced·6G are discussing the concept of TaaS (Time as a Service), which provides network-wide time as a service. Likely exam directions are "the principle behind the precision difference between NTP and PTP and each one's application fields," "the mechanism by which TrueTime's commit-wait guarantees external consistency," and "GNSS-dependence risk and time-infrastructure resilience design" being treated in essay form.

8. Considerations and Implications

  • Define the precision requirement first and invest accordingly. Not every system needs nanoseconds. Log correlation·authentication is served by millisecond-level NTP, and only real-time control in telecom·finance·power justifies PTP·dedicated hardware. Introducing high-precision infrastructure with no requirement only raises cost, while under-investing conversely leads to regulatory violations·outages, so work-impact-based precision-grade design must come first.

  • Treat time as a security-control target. Time manipulation is a high-risk attack surface that collapses certificate validation·Kerberos·OTP·log integrity all at once. Apply NTS to NTP, multiplex·cross-validate external time sources, and monitor sudden time changes as anomaly signs. In particular, segments that receive time from outside the trust boundary must have their origin verified by signing·authentication.

  • Break single-reference (GNSS) dependence and design for resilience. Relying on GPS alone makes jamming·satellite failure spread immediately into a wide-area time outage. Combine sources of different principles (GNSS·atomic clocks·terrestrial) and secure availability by placing an atomic clock capable of self-sustaining (holdover) when the reference is lost as a holdover source. This ties directly to the business-continuity (BCP) requirements of critical infrastructure.

  • Do not hide synchronization error; handle it explicitly. TrueTime's lesson is that "an algorithm that assumes error is zero is the most dangerous." In distributed DBs·event ordering, do not decide order by physical time alone; defensively design by combining uncertainty intervals·logical clocks ([[logical-clock]])·consensus algorithms so that time error does not break consistency.

  • Include operations·observability in the design stage. Time sync is easily left alone after configuration, but it quietly drifts from upper-server failure·changes in network asymmetry·oscillator aging. An observability system that constantly collects each node's offset·jitter·stratum·source state, alerts when a threshold is exceeded, and cross-validates sudden time changes against logs governs long-term reliability.

References


In one line: Clock synchronization is the technology that aligns the physical clocks of independently flowing nodes to a reference time or keeps the error within a known bound; it divides into software-timestamp, low-cost·wide-area NTP, PTP that achieves nanosecond level with hardware timestamps, and TrueTime that exposes error as an uncertainty interval and buys external consistency with commit-wait — with the precision-cost balance, GNSS-dependence resilience, and time security as the core design challenges.