Real-Time Operating System (RTOS)
1. Overview
A. Definition
An RTOS (Real-Time Operating System) is an operating system designed to honor a timing constraint—responding within a defined deadline without fail. Its core value is not 'how fast (throughput)' but 'can we predict when it will finish (determinism).'
The reason an RTOS emerged lies in the very nature of embedded control. In systems that operate coupled to the physical world—airbag deployment, motor control, aircraft flight control, industrial PLCs—what matters for safety is not an 'on-average fast response' but 'a response that finishes within a fixed deadline even in the worst case.' Even if an airbag fires on average in 10ms, if it occasionally takes 50ms the occupant may die. A general-purpose operating system (GPOS) is designed to maximize overall throughput and fairness and cannot guarantee an upper bound on response time, so an RTOS that makes determinism its primary goal became necessary to meet such requirements.
The starting point for understanding an RTOS, therefore, is to distinguish 'fast' from 'predictable.' Whereas a GPOS pursues statistical performance (average latency, throughput), an RTOS is designed and validated against a bounded latency and a WCET (Worst-Case Execution Time). No matter how excellent the average performance, if the upper bound cannot be guaranteed it is a failure as a real-time system.
B. Classification and characteristics of real-time behavior
The 'real-time behavior' an RTOS must guarantee is divided into three kinds according to the consequence of missing a deadline. This distinction is not mere classification but a design criterion that decides how much verification and margin to invest in.
| Type | Consequence of a missed deadline | Example |
|---|---|---|
| Hard | System failure, loss of life | Airbag, aircraft flight control, reactor control |
| Firm | Result loses value (no danger) | Industrial vision inspection, real-time trade execution |
| Soft | Quality degradation (tolerable) | Video streaming, VoIP |
Because a missed deadline in hard real-time is a disaster, 100% deadline compliance must be mathematically proven (schedulability analysis), whereas soft real-time tolerates occasional violations as quality degradation. For this reason, the more a system is hard real-time, the more cost it spends on WCET estimation, on removing the indeterminacy of caches and pipelines, and on guaranteeing an upper bound on interrupt latency. How much the fluctuation of latency, called jitter, is suppressed is another yardstick of RTOS quality.
A point to note here is that real-time behavior does not mean 'must be absolutely fast.' If a system has a deadline of 1 second, a response of 0.9 seconds is 'real-time,' while if the deadline is 1ms and the average is 0.5ms but the worst case is 1.2ms, it is 'not real-time.' In other words, the criterion for real-time is not the absolute value of speed but whether the declared deadline is relatively guaranteed. This shift in perspective is the starting point of RTOS design and analysis, and in an exam answer, describing it as a 'fast OS' is deemed to have missed the essence.
2. Core concepts and overall structure of an RTOS
The determinism of an RTOS is realized through several interlocking mechanisms. The most important is preemptive priority scheduling. It always runs immediately the highest-priority task among those in the ready state, and when a higher-priority task wakes up it immediately preempts the running task. As long as this rule holds, one can analytically compute when a given task will obtain the CPU.
The second is fast, upper-bounded context switching and interrupt response. A GPOS may disable interrupts for long stretches for throughput, but an RTOS minimizes interrupt-disabled intervals and specifies their upper bound, reacting to external events on the order of a few microseconds. The third is prevention of priority inversion. By applying the Priority Inheritance and Priority Ceiling protocols to the mutexes that protect shared resources, it bounds the phenomenon in which a low-priority task blocks a high-priority one indefinitely (see [[priority-inversion]]).
graph TD
subgraph APP["Application layer"]
T1["Task A(high)"]
T2["Task B(medium)"]
T3["Task C(low)"]
end
subgraph KERNEL["RTOS kernel"]
SCH["Scheduler(preemptive priority)"]
IPC["IPC(semaphore·queue·mutex)"]
MEM["Memory mgmt(fixed block)"]
TMR["Timer·tick service"]
end
subgraph HW["Hardware"]
CPU["CPU/MPU"]
INT["Interrupt controller"]
end
T1 --> SCH
T2 --> SCH
T3 --> SCH
SCH --> CPU
INT -->|ISR| SCH
IPC --- SCH
TMR --> SCH
MEM --- CPU
In the structure diagram above, application tasks are allocated the CPU only through the kernel's scheduler. External events pass through the interrupt controller and are delivered to an ISR (Interrupt Service Routine), and the ISR does only the minimum processing and then takes a deferred processing structure that wakes the relevant task through a semaphore or queue. If everything is handled inside the ISR, the interrupt-disabled interval lengthens and the response to other events is delayed, so 'short ISR + delegation to a task' is the canonical form of RTOS design.
Memory management also differs from a GPOS. A general-purpose OS's variable-size dynamic allocation (malloc) has uneven allocation times due to fragmentation, which harms determinism. So an RTOS guarantees constant-time allocation with a fixed-size memory pool scheme, or forbids dynamic allocation altogether (via safe-coding rules such as MISRA C) and uses only static allocation.
In the end, the determinism of an RTOS is not 'a single trick' but the result of eliminating or bounding multiple sources of non-determinism one by one. The non-determinism of scheduling is tamed by preemptive priority, that of interrupts by minimizing disabled intervals, that of memory by fixed blocks, and that of resource contention by priority inheritance. If even one of these collapses, the upper bound of the overall response time breaks, so RTOS design demands conservative thinking that traces the 'worst case' all the way through.
3. RTOS architecture and components
An RTOS kernel is generally microkernel-oriented. It keeps only minimal functions—scheduling, IPC, timers—in kernel mode and separates the file system, network stack, and so on into user-space servers or middleware. This keeps the kernel small and narrows the scope of verification, which is advantageous for obtaining safety certification (aviation DO-178C, automotive ISO 26262). Below is a detailed diagram showing the relationship between task state transitions and kernel services.
stateDiagram-v2
[*] --> Ready: create
Ready --> Running: dispatch highest-priority
Running --> Ready: preempt
Running --> Blocked: resource wait P-op·delay
Blocked --> Ready: resource release V-op·timeout
Running --> Suspended: explicit suspend
Suspended --> Ready: resume
Running --> [*]: delete
A. Tasks and the scheduler
A task (or thread) is the execution unit of an RTOS and each has its own priority, stack, and context (register set). At every scheduling point (tick interrupt, blocking call, return from interrupt) the scheduler picks the highest-priority task from the ready queue and dispatches it. If there are several tasks of the same priority, it may share fairly via round-robin (time slicing). Making this state transition occur in O(1) constant time using a bitmap-based priority table is a common technique of FreeRTOS, μC/OS, and others.
B. Inter-task communication and synchronization (IPC)
For multiple tasks to cooperate, they must exchange data and align their execution order. An RTOS provides semaphores (synchronization, resource counting), message queues (data transfer), mutexes (mutual exclusion), and event flags (waiting on multiple conditions). In particular, the mutex has priority inheritance built in to prevent priority inversion. For example, a FreeRTOS mutex lets the low-priority task that acquired it temporarily inherit the priority of the waiting high-priority task so it exits the critical section quickly.
Here, distinguishing a semaphore from a mutex is a practical pitfall. Both can be used for mutual exclusion, but the concept of ownership differs. A mutex can be released only by the task that locked it, so its owner can be identified and priority inheritance is possible; a binary semaphore can be released by any task, so its owner is unknown and inheritance cannot be applied. Hence the principle 'mutex for protecting a resource, semaphore for event signaling (ISR→task notification)' holds, and confusing the two—protecting a resource with a semaphore—produces a defect that reopens priority inversion.
C. Time management and interrupts
An RTOS counts time with a system tick timer. At each tick it wakes tasks whose delay has expired and refreshes the time slice. However, if the tick frequency is too high the overhead grows, and if too low the time resolution drops, so recently a tickless mode that stops the timer when no tick is needed pursues both low power and precision. Guaranteeing an upper bound on the task response time, which is the sum of interrupt latency and scheduling latency, is the core goal of kernel design.
4. Comparison of scheduling techniques
Real-time scheduling splits on whether priority is 'fixed' or changed 'dynamically.' The difference between the two representative algorithms is not a mere implementation difference but arises from the upper bound of schedulability.
| Technique | Priority assignment | Schedulable bound (single CPU) | Characteristics |
|---|---|---|---|
| RMS (Rate Monotonic) | Higher for shorter period (fixed) | About 69.3% (n→∞, n(2^(1/n)-1)) |
Static, simple, easy to predict |
| EDF (Earliest Deadline First) | Higher as deadline nears (dynamic) | 100% | Optimal, but cascading collapse on overrun |
RMS assigns fixed high priority to tasks with short periods (that run frequently). It is simple to implement and easy to analyze, so it is widely used in industry, but if CPU utilization exceeds its theoretical bound (about 69.3% as the number of tasks grows), deadlines cannot be guaranteed. That is, one must leave about 30% of the CPU 'as margin' to be safe, which is a cost but the price paid to buy the value of predictability.
EDF gives the highest priority dynamically to the task whose deadline is most imminent at that moment. It is an optimal algorithm that can meet all deadlines up to 100% utilization on a single processor, but when overrun occurs it is hard to predict which task will miss its deadline first, posing a risk of cascading deadline failure (domino effect). So safety-critical hard systems choose RMS with its clear analysis, while systems that must squeeze out utilization choose EDF—thus the trade-off splits. In fact, the OS of automotive AUTOSAR uses a fixed-priority preemptive scheme by default, valuing the analyzability of the RMS family.
A. RMS schedulability — numerical example
Concretely, suppose there are three tasks. Task 1 has a period of 50ms and execution time of 10ms, Task 2 a period of 100ms and 20ms, Task 3 a period of 200ms and 40ms. The CPU utilization of each task is 10/50=0.2, 20/100=0.2, 40/200=0.2, summing to 0.6. The RMS utilization bound for three tasks is 3(2^(1/3)-1)≈0.78, so a total utilization of 0.6 is below the bound 0.78, and thus it is guaranteed that all three tasks can meet their deadlines. If Task 3's execution time is increased to 60ms, making total utilization 0.7, it is still below the bound and safe, but increasing it to 90ms (0.85) exceeds the Liu-Layland bound and can no longer be guaranteed by this simple test alone, requiring a precise RTA (Response-Time Analysis). In this way, the fact that RMS can 'prove safety by calculation' is the reason it is preferred in industry.
5. Comparison with a GPOS, and application cases
The difference between an RTOS and a GPOS arises from the difference in design goal—'what is being optimized.' The comparison below is not a mere feature list but shows why it is hard to satisfy both requirements with a single OS.
| Perspective | RTOS | GPOS (Linux, Windows) |
|---|---|---|
| Primary goal | Determinism (response bound) | Throughput·fairness |
| Scheduling | Preemptive priority, bound guaranteed | Fairness-centric such as CFS |
| Kernel latency | A few μs, upper bound specified | A few ms, non-deterministic |
| Memory | Fixed block·static allocation | Virtual memory·paging |
| Size | A few KB to a few hundred KB | Tens of MB or more |
The fundamental reason it is hard to satisfy both requirements with a single OS is that the optimization targets conflict. To raise throughput and fairness, the scheduler reorders execution to chase cache hits, batches interrupts, and gathers I/O with deferred write-back—all techniques that improve the 'average' but make the 'worst case' unpredictable. Conversely, an RTOS sacrifices some average performance to bound the worst case. So traditionally the roles were split and different OSes were used, and only recently has the approach of coexisting the two on one chip via multicore and a hypervisor become feasible.
Looking at actual cases, the difference in design goals becomes clear. First, an automotive electronic control unit (ECU) uses an OSEK/VDX-based AUTOSAR Classic OS for engine and braking control. Because injection timing must be controlled on the order of hundreds of μs according to crank angle, it is impossible with a GPOS whose upper bound is not guaranteed, no matter how good the average performance. Second, Mars exploration rovers have carried Wind River's VxWorks; the 1997 Mars Pathfinder reboot fault was due to priority inversion, and the anecdote of fixing it by remotely patching in priority inheritance remains a textbook case of RTOS design. Third, small IoT and wearables use FreeRTOS (acquired by AWS in 2017) or Zephyr running in a few KB of RAM to meet the timing of sensor sampling and wireless communication with milliwatt-level power.
6. Deep dive — latest trends and mixed criticality
The biggest change in the RTOS domain recently is that the boundary between a general-purpose OS and real-time behavior is breaking down. Linux's real-time patch PREEMPT_RT, after a long period as an external patch, was largely merged into the mainline in the Linux 6.12 kernel in 2024, so Linux too can obtain a response close to limited hard real-time. This reflects industrial and robotics demand to run a rich application (Linux) together with real-time control on a single SoC.
That said, PREEMPT_RT does not replace all hard real-time. Linux still has a large kernel and it is hard to fully remove the non-determinism of virtual memory and caches, so for control loops with strict microsecond-level deadlines a dedicated RTOS or a separate core is still used alongside. In other words, it is accurate to understand it not as 'Linux replacing an RTOS' but as a reorganization of role division in which loose real-time is absorbed by Linux and strict real-time is handled by an RTOS.
Interlocking with this, mixed-criticality systems are rising. On a single multicore chip, functions of different safety grades (e.g., safety control and infotainment in autonomous driving) run together, but cores, memory, and peripherals are isolated with a hypervisor (e.g., Xen, PikeOS, QNX Hypervisor) so that a malfunction in a low-grade function cannot intrude on the timing of a high-grade one. Here, how to bound the interference of shared resources such as cache and memory bandwidth remains a hard problem for analysis.
On the standards and ecosystem side, the Linux Foundation's Zephyr Project is growing into a de facto standard for open-source RTOSes, supporting hundreds of boards, and there is a clear direction of combining an RTOS with TSN (Time-Sensitive Networking), which guarantees determinism in industrial networks, to pursue end-to-end real-time behavior 'from chip to network.' As likely exam directions, (1) the hard/soft distinction and schedulability analysis (computing the RMS utilization bound), (2) priority inversion and the inheritance/ceiling protocols, (3) mixed criticality and hypervisor isolation, and (4) RTOS vs GPOS (including PREEMPT_RT) comparison may often appear combined.
7. Considerations and implications
From a professional engineer's perspective, the following should be considered holistically when adopting or designing an RTOS.
- Application strategy (matching requirement grades): Designing every function as hard incurs excessive margin and cost. A criticality-based resource allocation that distinguishes hard/firm/soft per function and concentrates WCET analysis and spare CPU only on the hard paths is cost-effective.
- Trade-off (utilization vs predictability): RMS leaves about 30% of the CPU empty but has clear analysis, while EDF uses utilization up to 100% but carries the risk of collapse on overrun. One should choose according to the system's safety grade and load characteristics, and design the overrun handling (a degraded-performance mode on a missed deadline) together.
- Verification and certification perspective: In hard systems, the reliability of WCET estimation is the foundation of overall safety. Because caches, branch prediction, and multicore interference make WCET estimation difficult, safety certification (DO-178C, ISO 26262) requires a design that disables hardware features harming determinism or conservatively upper-bounds their effect.
- Outlook and related technologies: The mainlining of PREEMPT_RT and multicore hypervisors are accelerating the 'integration of RTOS and GPOS.' In linkage with TSN, edge AI, functional safety (ISO 26262), [[priority-inversion]], and [[race-condition]], the capability to design architectures that guarantee end-to-end determinism grows ever more important.
References
- FreeRTOS Documentation, https://www.freertos.org/features.html
- Zephyr Project (Linux Foundation), https://www.zephyrproject.org/
- Linux 6.12 release notes (PREEMPT_RT), https://kernelnewbies.org/Linux_6.12
- Liu & Layland, "Scheduling Algorithms for Multiprogramming in a Hard-Real-Time Environment", https://dl.acm.org/doi/10.1145/321738.321743
In one line: An RTOS makes its primary goal not throughput but determinism that guarantees a response within a deadline, realizing predictability with preemptive priority scheduling, upper-bounded interrupts, priority inheritance/ceiling, and fixed-block memory; it is designed for hard/soft requirements on the trade-off between RMS (easy to analyze, 69.3% utilization) and EDF (optimal, 100%), and is evolving in a direction where the boundary with a GPOS breaks down through PREEMPT_RT, mixed criticality, and TSN.