Socket Communication
1. Overview
A. Definition
A socket is an endpoint for inter-process communication over a network, identified by the combination of an IP address and a port number. Socket communication is the way an application sends and receives data through this socket; it is the network extension of inter-process communication (IPC), provided by the operating system by wrapping the TCP/IP stack in a standard API.
The key to understanding sockets is that a socket is 'an abstracted window through which an application deals with a complex network'. The internal behavior of TCP/IP (packet fragmentation and reassembly, routing, ordering, retransmission, flow control, congestion control) is extremely complex, but a developer can communicate without knowing this complexity as long as they know the standard interface of socket(), bind(), connect(), send(), and recv(). This is the power of abstraction, much like using a telephone without needing to know how the switching network works internally.
A socket identifies the communication counterpart with two coordinates. The IP address specifies 'which computer (host) on the Internet' and the port number specifies 'which program (process) on that computer'. For example, a web server usually opens sockets on ports 80 (HTTP) and 443 (HTTPS) to wait for client connections. IP:port is often likened to 'building address:room number', and in TCP communication a single connection is uniquely identified by the 5-tuple of (source IP, source port, destination IP, destination port, protocol). Thanks to this 5-tuple concept, a single server can distinguish and simultaneously handle tens of thousands of client connections even on the same port 80.
B. Background and Necessity
The origin of the socket lies in BSD Unix introducing the Berkeley Socket API for network programming in the early 1980s. Before that, communication code had to be rewritten for each network hardware and protocol, but by abstracting network communication with a unified descriptor model similar to file I/O (read/write), the socket made portable network application development possible. For programs on different devices to communicate, they need a standard method that ① specifies the counterpart, ② exchanges data reliably, and ③ works regardless of operating system or language; the socket became the de facto standard that satisfies all three requirements. Today almost all network applications—web, messengers, games, IoT, streaming—run on top of sockets.
C. Characteristics
Socket communication is bidirectional, and in the case of connection-oriented sockets, once a connection is established it remains stateful until it is explicitly closed. In addition, because the operating system kernel handles buffering, retransmission, and flow control, the application layer can focus solely on the meaning of the data. On the other hand, since the characteristics of the transport-layer protocol (TCP/UDP) are exposed as-is beneath this abstraction, the developer is directly responsible for the reliability-versus-performance trade-off depending on which socket type is chosen.
2. Socket Structure and Identification Scheme
flowchart TB
subgraph APP["Application Layer"]
A1["Application"]
end
subgraph OS["OS Kernel"]
S["Socket (Socket Descriptor)"]
SB["Send Buffer / Receive Buffer"]
TCP["Transport Layer (TCP / UDP)"]
IP["Network Layer (IP)"]
end
NIC["NIC (Physical Layer)"]
A1 -->|"socket() / send() / recv()"| S
S --> SB --> TCP --> IP --> NIC
NIC -->|"network"| NET(("Internet"))
style S fill:#e8f0fe,stroke:#2f6fed
A socket is the boundary located between the application and the network stack of the operating system kernel. When an application calls socket(), the kernel creates the data structures needed for communication and returns an integer handle called the socket descriptor (a kind of file descriptor). Thereafter the program handles network communication like a file using this single integer. This is the result of extending the Unix philosophy that "everything is a file" to the network, and thanks to this uniformity, multiplexing functions such as select/poll can treat files and sockets identically.
The fact that socket descriptors draw from the same resource pool as file descriptors shows up as a concrete constraint in practice. On Linux, the number of descriptors a process can open simultaneously is often set to 1024 by default (ulimit -n), so unless this value is raised, accept() fails with an EMFILE error the moment concurrent connections exceed 1024. Therefore a high-connection server must raise kernel limits such as ulimit and fs.file-max to the hundreds of thousands and monitor for connection leaks (unclosed sockets). This shows that a socket is not merely a concept but an entity directly tied to operating-system resource management.
What is especially important in the socket identification scheme is the existence of the send buffer and receive buffer. The return of send() does not mean the data has reached the counterpart; it has merely been copied into the kernel's send buffer, while the actual transmission, retransmission, and ordering are performed asynchronously by the kernel's TCP implementation. For this reason, socket programming must always account for "partial write" and "partial read". For example, even if you send() 10KB, only 4KB may be processed depending on the buffer situation, so the application must write a loop that checks the return value and repeats the transmission. Neglecting this point causes bugs that look like data loss in large transfers.
| Component | Role | Practical implication |
|---|---|---|
| Socket descriptor | Integer handle pointing to a communication endpoint | The number a process can open (ulimit -n) governs the concurrent-connection ceiling |
| IP address | Identifies the host (computer) | Binding address must be chosen in NAT/multi-interface environments |
| Port number (0-65535) | Identifies the process (service) | 0-1023 are well-known ports (require privileges) |
| Send/receive buffer | Kernel's temporary data store | Buffer size (SO_RCVBUF etc.) affects throughput |
| 5-tuple | Unique identifier of a connection | Basis for handling many connections simultaneously on one port |
3. Socket Communication Procedure and Types
sequenceDiagram
participant C as Client
participant S as Server
Note over S: socket() → bind() → listen()
S->>S: wait for connection at accept() (blocking)
C->>C: socket()
C->>S: connect() (TCP 3-way handshake)
S-->>C: accept() returns (connection socket created)
loop Data exchange
C->>S: send() / recv()
S->>C: send() / recv()
end
C->>S: close() (4-way handshake)
S->>S: close()
The TCP socket communication procedure consists of four stages: server preparation → connection establishment → data exchange → termination. The server first creates a socket with socket(), binds its own IP and port with bind(), transitions to a state where it can accept connection requests (creating a wait queue) with listen(), and waits for client connections at accept(). An important design point here is that accept() returns a new, separate connection socket. That is, the listen socket remains as the 'reception window' that accepts connections, while the actual data communication is handled by a newly created connection socket for each client. This structure is the basis on which one server can serve many clients simultaneously.
The client requests a connection to the server with connect() after socket(), and during this process TCP's 3-way handshake (SYN → SYN/ACK → ACK) takes place. Once the connection is established, both sides exchange data with send()/recv(), and when communication ends, close() tears down the connection through the 4-way handshake (FIN/ACK exchange). At this time a TIME_WAIT state remains on the server side for a certain period (usually 2*MSL), which is a safeguard for handling late-arriving packets; however, on a server that makes a large number of short connections, TIME_WAIT sockets can accumulate and cause port exhaustion. In practice this is mitigated with the SO_REUSEADDR option or connection reuse (keep-alive).
Sockets are divided into two types depending on the transport protocol used. A stream socket (TCP) establishes a connection, guarantees the order and reliability of data, and provides a boundary-less byte stream. Because of this stream nature, the message boundaries the application sent are not preserved, so multiple messages may arrive glued together (or split apart). To handle this, a self-defined framing convention such as a length prefix or delimiter is needed. A datagram socket (UDP) sends individual messages quickly without a connection and does not guarantee order or arrival, but message boundaries are preserved.
The way sockets operate has another important axis: the distinction between a blocking socket, which halts the thread until data arrives (as in recv()), and a non-blocking socket, which returns immediately and only reports readiness. The blocking approach yields intuitive code but occupies a thread per connection, so scalability is poor; the non-blocking approach can handle a large number of connections with a few threads when combined with an event loop, but state management becomes complex. This choice is directly tied to the I/O model of large-scale servers discussed later. In addition, socket options such as SO_KEEPALIVE (idle-connection liveness check), TCP_NODELAY (disabling the Nagle algorithm to remove latency for small data), and SO_REUSEADDR (address reuse) allow latency, throughput, and resource use to be finely tuned on top of the same socket API. For example, a real-time game server turns on TCP_NODELAY to prevent small packets from being gathered in the buffer and sent late.
| Type | Protocol | Characteristics | Representative use |
|---|---|---|---|
| Stream socket | TCP | Connection-oriented, guarantees reliability/order, byte stream | Web/file transfer/DB/messenger |
| Datagram socket | UDP | Connectionless, fast but unreliable, preserves message boundaries | Real-time video/games/DNS/VoIP |
| RAW socket | IP/ICMP etc. | Direct access to lower layers | ping/traceroute/security tools |
4. TCP/UDP Sockets and WebSocket — Comparison and Cases
If a TCP socket is low-level communication that deals with the transport layer directly, WebSocket is a higher-layer technology for the web browser environment. Because browsers cannot open arbitrary TCP sockets for security and policy reasons, WebSocket (RFC 6455) emerged as the standard for real-time bidirectional communication. WebSocket starts with an ordinary HTTP GET request, then switches the protocol (handshake) with the Upgrade: websocket header, after which the server and client freely exchange frame-based messages (full-duplex) over a single persistent connection (the same TCP connection). Thanks to this, 'server push', in which the server sends data to the client first, becomes possible, implementing real-time web services such as chat, real-time notifications, stock quotes, and collaborative editing (Google Docs-style).
Comparing with concrete cases: online games feel even tens of ms of latency, so they commonly use a hybrid structure that sends position/movement data over UDP (compensating for some loss with the next packet) and handles accuracy-critical data such as payments and item trading over TCP. Video conferencing (WebRTC) likewise splits media over a UDP base (SRTP) and signaling over TCP/WebSocket because signaling needs reliability. A stock quote dashboard must have the server push thousands of quote updates per second, so WebSocket greatly reduces latency and server load compared to polling. Thus the latency sensitivity and accuracy requirements of the data govern the protocol choice.
| Category | TCP socket | WebSocket |
|---|---|---|
| Layer | Transport layer (TCP) directly | Application layer (upgrade over HTTP) |
| Connection | socket→connect→3-way handshake | Upgrade after HTTP handshake |
| Communication | Bidirectional byte stream | Bidirectional full-duplex (frame-based) |
| Environment | Server/native apps | All areas including web browsers |
| Framing | Implemented by the application | Frames provided by the protocol |
| Use | General network apps | Web real-time (chat/notifications/quotes) |
5. Comparison of Socket Communication and HTTP Communication
Socket communication and HTTP differ fundamentally in how they maintain connections. HTTP is a stateless-oriented approach that completes processing after a request-response, so in principle the server cannot speak to the client first. When real-time is needed, it relies on polling or long polling in which the client repeatedly requests, which is inefficient because it causes unnecessary requests, header overhead, and latency. That said, HTTP also runs on top of TCP sockets internally, and HTTP/1.1 keep-alive and HTTP/2 multiplexing have reduced this overhead through connection reuse. The key difference is that socket communication (especially WebSocket) naturally supports bidirectional, real-time communication by keeping the connection itself alive.
In practice, intermediate alternatives among the three are also actively used. If the server only needs one-way push to the client, SSE (Server-Sent Events) works economically over HTTP; if true bidirectional, low-latency is needed, WebSocket is chosen; and for simple request-response, REST (HTTP) is chosen. This judgment must consider not only performance but also firewall/proxy traversability and the complexity of reconnection and authentication handling.
Quantitatively the difference is clear. For example, providing data that updates every second to 10,000 people via polling generates 10,000 HTTP requests per second, with a few hundred bytes of headers repeated on each request; WebSocket, by contrast, pushes only a few-byte frame when needed after a single initial handshake, cutting network and server load by tens of times or more. Conversely, for simple information that is queried only a few times a day, the cost of maintaining a persistent connection is a waste, so HTTP request-response is the right answer. In other words, the key is not 'which is superior' but matching the workload characteristics of 'update frequency, bidirectionality, and connection count'.
| Category | Socket communication (WebSocket) | HTTP communication |
|---|---|---|
| Connection | Persistent connection (stateful) | Terminated after request-response (stateless) |
| Direction | Bidirectional (server push possible) | Unidirectional (client-initiated) |
| Real-time | High | Low (polling needed) |
| Overhead | Minimal frame after initial handshake | Headers repeated on every request |
| Use | Real-time/bidirectional | Web documents/REST API |
6. Deep Dive: Large-Scale Real-Time Services and High-Performance Socket Handling
In large-scale real-time services, the real challenge of sockets is how to efficiently maintain and process a huge number of concurrent connections. This was formalized as the classic C10K problem (the problem of one server handling 10,000 concurrent connections), and today it is discussed up to the C10M (ten million) level. The early 'thread/process per connection' model had limits to scaling because memory and context-switching costs grew linearly with the number of connections. What solved this is event-based I/O multiplexing. Linux's epoll, BSD/macOS's kqueue, and Windows' IOCP let a single thread efficiently receive notifications from the kernel about state changes on tens of thousands of sockets and process them. This event-loop model is exactly the secret behind Nginx, Node.js, and Redis handling large numbers of connections with a few threads. Recently Linux's io_uring (5.1+) is an asynchronous I/O interface that reduces even system-call overhead, and its adoption is expanding in high-performance servers and proxies.
The power of the event-loop model is starkly revealed in resource usage. A model that uses a thread per connection consumes hundreds of KB to several MB per connection just for the thread stack, so 10,000 connections require several GB of memory; the event loop, however, manages a connection with only a socket descriptor and a small state object, handling far more connections on the same hardware. However, since an event loop stalls entirely if a callback occupies the single thread for too long, a design that separates CPU-intensive work into worker threads/processes must accompany it. In this way, the choice of socket-handling model determines the server's cost structure and failure characteristics.
When scaling out a WebSocket-based service across multiple servers, a new problem arises. Because a persistent connection is fixed to a specific server instance (sticky session), propagating a message received by server B to a user connected to server A requires inter-server message propagation. In practice, a message broker such as Redis Pub/Sub or Kafka is placed as a backplane to fan out events between instances, and connections are distributed at the load balancer via sticky sessions or consistent hashing. In addition, periodic ping/pong (heartbeat) is sent so that idle connections are not dropped by firewalls/NAT, and dropped connections are reconnected with exponential backoff—such connection lifecycle management governs service quality. Large-scale messengers such as KakaoTalk, Slack, and Discord are representative cases of this architecture.
On the security side, TLS must be applied to socket communication (TLS for TCP, wss:// for WebSocket) to prevent eavesdropping and tampering, and unauthorized connections and CSWSH (Cross-Site WebSocket Hijacking) must be prevented through Origin validation and token-based authentication (JWT, etc.) during the WebSocket handshake. In addition, prepare for resource-exhaustion DoS with per-connection resource caps, message size limits, and rate limiting.
Recently the very landscape of socket communication is changing. QUIC (the transport foundation of HTTP/3) re-implements connection, reliability, and encryption (built-in TLS 1.3) on top of a UDP socket, merging TCP's 3-way handshake and the TLS handshake into one and supporting connection migration (keeping the connection alive even when the network switches). This breaks the long-standing premise that "reliability must be TCP" and shifts the center of gravity of application development toward using UDP sockets for both real-time and reliability. Also, as serverless and edge environments spread, a pattern is growing in which, instead of traditional WebSocket that presumes long-lived persistent connections, the server side minimizes state and delegates connection management to a managed WebSocket gateway (e.g., API Gateway's WebSocket support).
7. Considerations and Implications
Choosing the communication method that fits the requirements is the starting point of architecture. Simple request-response is straightforward with HTTP/REST, real-time bidirectional suits WebSocket, one-way server push suits SSE, and latency-sensitive with tolerance for small loss suits UDP. From a professional engineer's perspective, a comprehensive judgment is needed that includes not only performance but also development/operational complexity and proxy traversability.
Concurrent-connection scalability is determined by the I/O model. Large-scale services should consider an event loop based on
epoll/kqueue/IOCP rather than thread-per-connection, and further io_uring; connection-keeping services must also carry a message-broker, sticky-session, and backplane design.Applying the reliability-versus-speed trade-off separately per data unit is realistic. As in games and video conferencing, a hybrid design that splits media over UDP and control/transactions over TCP even within one service is common, and this connects to the recent trend of re-implementing reliability on top of UDP, as in QUIC (HTTP/3).
Connection lifecycle and failure resilience must be reflected in the design. TIME_WAIT/port exhaustion, zombie connections, NAT timeouts, and partial transfers do not surface at small scale but spread into failures as traffic grows, so heartbeat, reconnection, backpressure, and timeout policies must be standardized.
Security by design is essential. Avoid plaintext sockets and default to TLS/
wss, and embed authentication, Origin validation, and rate limiting in the communication layer so that the real-time channel does not become an attack surface.Securing observability is a prerequisite for operations. Because persistent connections, unlike request-response, accumulate problems silently, metrics such as active connection count, connection lifetime, reconnection rate, message latency, and buffer occupancy must be continuously measured and threshold alarms set to catch anomalies early. This connects naturally with SLO management from the SRE perspective.
References
- IETF RFC 6455, "The WebSocket Protocol" — https://datatracker.ietf.org/doc/html/rfc6455
- Dan Kegel, "The C10K problem" — http://www.kegel.com/c10k.html
- Linux man-pages, socket(2)/epoll(7) — https://man7.org/linux/man-pages/man2/socket.2.html
- MDN Web Docs, "The WebSocket API" — https://developer.mozilla.org/en-US/docs/Web/API/WebSockets_API
In one line: A socket is a communication endpoint identified by IP and port, with TCP (reliable) and UDP (fast) sockets and real-time bidirectional WebSocket; persistent, bidirectional socket communication contrasts with request-response, stateless HTTP and suits real-time services, while large-scale scaling hinges on event-based I/O such as epoll and io_uring plus message-broker and security design.