Silicon Photonics and Optical Interconnects
1. Overview
A. Definition
Silicon Photonics is a technology that reuses the very CMOS semiconductor process once used to build transistor circuits that carry electrons, to integrate optical devices (modulators, waveguides, detectors, etc.) that use light (photons) instead of electrons as the signal medium onto a silicon substrate. In other words, its essence is to fabricate Photonic Integrated Circuits (PICs) that generate, transmit, and receive data as optical rather than electrical signals, using cheap, high-volume silicon manufacturing.
Traditionally, optical communication components (lasers, modulators, photodetectors) were built and assembled individually from compound semiconductors such as indium phosphide (InP) and gallium arsenide (GaAs), making them expensive and hard to integrate at scale. Silicon photonics integrates these optical devices at wafer scale on mature silicon CMOS fab infrastructure, seeking to capture both the performance of optical communication (bandwidth, low loss) and the economics of semiconductors (yield, integration density, low cost) at once. Intel, Cisco (Acacia), GlobalFoundries, and Tower Semiconductor operate commercial platforms, and recently NVIDIA and Broadcom have adopted the technology in AI datacenter switches, propelling it into an industry-central technology.
B. Background and Necessity
Three converging pressures drove the rise of silicon photonics. The first is the explosion of AI/HPC traffic and the physical limits of electrical wiring. Copper (electrical) wiring suffers surging signal attenuation, crosstalk, and skin effect as data rates rise, so beyond 100 Gbps the effective transmission distance shrinks abruptly to tens of cm to a few m. In large AI clusters connecting tens of thousands of GPUs, this "distance wall of electricity" becomes the fundamental bottleneck to scaling. Light is far freer from such distance and bandwidth constraints, making optical transmission inevitable.
The second is power and heat. A substantial share of datacenter network power is spent converting and driving signals between electrical and optical domains. As data rates climb, failing to lower the energy per bit (pJ/bit) causes power and cooling costs to explode. Integrating optical devices right next to compute and switch chips minimizes the electrical section and greatly lowers energy per bit.
The third is cost and integration density. Manually aligning and assembling individual compound-semiconductor components does not scale in volume. Applying the high-volume, self-aligned advantages of the silicon CMOS process to optical devices lowers unit cost and raises integration density. As these three pressures intertwined, silicon photonics became the core infrastructure of Data Movement in the AI era.
C. Technical Characteristics
Silicon photonics can be characterized in three points. First, CMOS process reuse. Because it leverages decades of mature silicon fab and design assets directly, optical devices can be produced at wafer scale at low cost and high volume. This creates a fundamental cost and integration-density advantage over traditional optical components that were assembly-centric. Second, high bandwidth density and low-loss long-distance transmission. Unlike electricity, light has no skin effect or crosstalk, and WDM multiplies channels per wavelength to maximize bandwidth per width and area. Third, the heterogeneity of electronics and photonics. Because of the physical duality that silicon is ideal for compute and logic but unsuitable for light emission, a hybrid structure that relies on III-V materials for the light source alone is unavoidable, and this is both the key constraint and a differentiator of design and supply chain.
2. Operating Principle and Core Devices
A silicon photonics link operates as a chain that converts an electrical signal into light (transmit) → carries it over a waveguide → converts it back to electricity (receive). The diagram below shows the relationships among the devices that make up a single optical link.
flowchart LR
DATA["Electrical data(Tx)"] --> MOD["Optical modulator<br/>(Modulator)"]
LASER["Laser source<br/>(usually external InP)"] --> MOD
MOD --> WG["Silicon waveguide<br/>(Waveguide)"]
WG --> MUX["Wavelength mux(WDM)"]
MUX --> FIBER["Fiber transmission"]
FIBER --> DEMUX["Wavelength demux"]
DEMUX --> PD["Photodetector<br/>(Photodetector)"]
PD --> TIA["Amplify/recover(Rx)"]
TIA --> OUT["Electrical data"]
Let us examine each device by principle. The light source (Laser) is the starting point that generates light. Silicon is an indirect-bandgap material with a fundamental weakness in emitting light efficiently, so most designs bond a III-V laser such as InP onto the silicon chip (hybrid integration) or attach an External Laser Source. This physical constraint that "silicon cannot emit light well" is the greatest challenge in silicon photonics design, and the choice of light-source integration method (hybrid bonding, flip-chip, external source) determines performance, yield, and cost.
The optical modulator (Modulator) is the device that imprints 0/1 onto the continuously emitted laser light by loading electrical data onto it. In silicon, a Mach-Zehnder Modulator (MZM) or a Micro-ring modulator (Micro-ring) that changes the refractive index by varying carrier density is mainly used. The micro-ring is small and low-power but temperature-sensitive and needs precise thermal control, while the MZM is large but stable — this trade-off is central to design choices.
Modulation performance directly governs the link's data rate. Recently, higher-order modulation that loads multiple bits per symbol, such as PAM4, is combined to achieve over 100 Gbps per wavelength, which is again multiplied by WDM to obtain hundreds of Gbps to Tbps-class bandwidth per waveguide. That is, "modulation order × number of wavelengths" are the two axes that determine link bandwidth, and layered with low-power and low-latency demands they create a complex device-selection optimization problem.
The waveguide (Waveguide) is a silicon channel that light passes through without leaking, confining light via the large refractive-index difference between silicon and its insulating layer (SiO₂). Thanks to this large index contrast, light does not escape even when the waveguide is bent into sub-micrometer curves, so complex optical circuits can be densely integrated in a small area — a strength of the silicon platform. The photodetector (Photodetector) converts arriving light back into current; because silicon does not absorb communication wavelengths (1310/1550 nm) well, germanium (Ge) is grown on silicon to secure absorption efficiency. This heterogeneous combination, in which silicon, germanium, and III-V each contribute their strengths on a single chip, is the typical composition of silicon photonics. Finally, WDM (Wavelength Division Multiplexing) loads data onto light of different wavelengths and sends them simultaneously over one fiber, multiplying bandwidth by the number of wavelengths.
| Device | Role | Silicon's limit and remedy |
|---|---|---|
| Light source (Laser) | Light generation | Silicon emits inefficiently → III-V (InP) hybrid integration |
| Modulator | Electrical→optical data imprint | MZM (stable) · micro-ring (small, low-power, temp-sensitive) |
| Waveguide | Light transmission channel | Confines light via Si/SiO₂ index difference |
| Photodetector (PD) | Optical→electrical conversion | Poor Si absorption → remedied by Ge growth |
| WDM | Wavelength multiplexing | Multiplies bandwidth by number of wavelengths |
A. Light-Source Integration — Three Paths Past Silicon's Emission Limit
The success or failure of silicon photonics design effectively hinges on "where and how the light is generated and injected". Because of silicon's physical inability to emit light efficiently, light-source integration strategies split broadly into three paths.
The first is hybrid/hetero-integration, in which a III-V laser wafer such as InP is bonded onto the silicon wafer and then processed to couple the light source with the silicon waveguide. It offers high integration density and low optical coupling loss, but the process difficulty and yield of heterogeneous material bonding are challenges. The second is flip-chip/micro-optical coupling, in which a separately fabricated laser chip is precisely aligned and placed atop the silicon PIC. It can use light-source chips of verified quality, but alignment precision governs performance. The third is the External Laser Source (ELS), which keeps the laser in a low-temperature, stable environment outside the package and supplies light via fiber. It can avoid the problem of laser reliability degrading at the high temperatures beside a hot CPU/GPU, so it is preferred in CPO, but it adds an external light-source module and optical distribution structure.
Each method differs greatly in the precision and loss of the optical coupling that joins the light source to the silicon waveguide. If coupling alignment drifts, light is lost and eats into the link budget, so coupling techniques such as self-aligned structures or spot-size converters are designed alongside.
This choice is not merely a component problem but a fundamental trade-off among performance, yield, reliability, and cost. For example, when placing the optical engine inside a hot package as in CPO, including the laser too (on-package) raises integration density but increases the risk of thermal wavelength drift and shortened lifetime, so many commercial designs separate the light source as an ELS. Conversely, for small links where latency and integration density are paramount, hybrid integration is advantageous.
B. WDM and the Principle of Bandwidth Expansion
The most powerful means of increasing the bandwidth of a single fiber/waveguide is WDM (Wavelength Division Multiplexing). Because light of different wavelengths (colors) operates as independent channels without interference even when passing through the same waveguide, loading data onto each of N wavelengths yields N times the bandwidth over one physical wire. For example, loading 100 Gbps onto each of 8 wavelengths transmits 800 Gbps over a single waveguide. This principle of "dividing channels by color" is the basis for silicon photonics' overwhelming advantage over electrical wiring in bandwidth density (bandwidth per width and area).
Realizing WDM requires multiple light sources of precise wavelengths and a multiplexer (MUX) and demultiplexer (DEMUX) that combine and separate wavelengths. In silicon these are implemented with micro-ring filter arrays or an Echelle grating; because a ring's resonant wavelength shifts with temperature, precise thermal-tuning circuits are included so that each wavelength stays exactly aligned. This temperature-stabilization burden is the key design challenge of micro-ring-based high-density WDM, and it is directly tied to the thermal-management challenge mentioned earlier.
3. The Evolution of Optical Interconnect Architecture — From Pluggable to CPO
The industrial value of silicon photonics varies greatly with how close the optical devices are placed to the compute/switch chip. Below is the evolution in which the optical-electrical conversion point moves ever closer to the chip.
flowchart TB
subgraph P["Pluggable optical modules"]
S1["Switch ASIC"] -->|"long electrical wiring"| M1["Pluggable optical transceiver(front port)"]
end
subgraph L["LPO(linear drive)"]
S2["Switch ASIC"] -->|"DSP removed, linear drive"| M2["Optical transceiver"]
end
subgraph C["CPO(Co-Packaged Optics)"]
S3["Switch/GPU ASIC"] -->|"short electrical path"| E3["Optical engine(integrated in same package)"]
E3 --> F3["Direct fiber attach"]
end
The traditional Pluggable approach drags data from the switch/compute chip over long electrical wiring to a pluggable optical transceiver (QSFP, etc.) on the front of the equipment, converting to optical there. It is easy to maintain and replace, but the long electrical section from chip to port is the main culprit of power and loss, and maintaining that section becomes harder as data rates rise. For example, toward 800G/1.6T class, the DSP and equalizer power for compensating electrical-channel loss surges.
LPO (Linear-drive Pluggable Optics) removes the power-hungry DSP inside the transceiver and lets the switch ASIC drive the optical device directly and linearly, lowering energy per bit and latency as a compromise. As an approach that reduces power while keeping the pluggable form (ease of replacement), it draws attention as a realistic alternative before CPO.
CPO (Co-Packaged Optics) integrates the optical engine (the silicon-photonics-based optical-electrical conversion unit) side by side inside the same package as the switch ASIC or GPU, extremely shortening the electrical path to the level of a few mm. Because the electrical section is short, the loss-compensation burden disappears, energy per bit drops sharply, and fiber is pulled directly out of the package. Instead, if the optical engine fails, the entire package must be handled, so maintainability, yield, and thermal management become new challenges. Thus the essence of architectural evolution is the question of "how tightly the optical-electrical conversion point is pressed against the compute silicon to shorten the electrical section."
This evolution also affects the network layer structure. Traditionally optical was used only for long inter-rack distances, but as bandwidth demands rise, optical is penetrating even intra-rack and GPU-to-GPU connections within a package. That is, the distance range that optical covers is descending from kilometers to cm and mm, and this is enlarging the role of optical interconnects in both scale-up (tight coupling within a node) and scale-out (expansion across nodes).
| Category | Pluggable | LPO | CPO |
|---|---|---|---|
| Conversion location | Front port of equipment | Front port (DSP removed) | Same package as chip |
| Electrical path | Long (tens of cm) | Long (linear drive) | Very short (mm class) |
| Power efficiency (pJ/bit) | Low | Medium | High |
| Maintenance | Easy (pluggable) | Easy | Hard (package-integrated) |
| Maturity | Commercial standard | Early adoption | Full commercialization in 2025 |
A. Packaging/Integration Tiers and Thermal Management
In CPO, the optical engine is placed atop and interconnected on an advanced package (2.5D interposer, silicon bridge) together with the switch/GPU ASIC. In exchange for shortening the electrical path to mm class, the optical engine, compute die, and (in some cases) HBM come to share heat on a single substrate. The hetero-integration and 2.5D packaging techniques covered earlier in [[chiplet]] are reused as-is for optical-engine integration, which is why the commercialization of silicon photonics is strongly tied to advanced packaging capacity (e.g., CoWoS).
Heat is CPO's greatest challenge. The switch ASIC consumes hundreds of watts and that heat conducts to the adjacent optical engine, while micro-ring modulators/filters shift their resonant wavelength with temperature at the level of tens of pm/°C. If wavelength drifts, channels misalign and the link breaks, so a micro-heater is attached to each ring to actively correct the wavelength, but this heater consumes power in turn. Ultimately, minimizing the paradox of "choosing CPO for low power yet spending power again on thermal correction" is central to design optimization, and athermal device design or separation of the external light source are studied as remedies.
B. Optical Link Performance Metrics and Design Budget
The quality of a silicon photonics link is evaluated by several key metrics. Insertion Loss indicates how much light attenuates from source to detector, summing coupling loss (fiber↔chip), waveguide propagation loss, and modulator loss. The Link Budget is the design that examines whether this total loss can be absorbed between the source output and the detector's minimum receive sensitivity; if the budget is tight, a stronger laser or a more sensitive detector is needed, raising power and cost.
As performance/efficiency metrics, energy per bit (pJ/bit), bandwidth density (Tbps/mm), BER (bit error rate), and latency are managed together. Especially in AI scale-up, latency and energy per bit directly affect the entire cluster's training efficiency, so what is required is not individual device performance but end-to-end optimization spanning link, package, and system. This shows that silicon photonics is not a mere component technology but a problem of system-design competency.
4. Comparison and Trade-offs — Electrical vs. Optical, and the Price of CPO
When optical interconnects replace electrical wiring is determined by the break-even of distance, bandwidth, and power. At short distances and low bandwidth, copper electrical wiring remains simple, cheap, and advantageous. However, as data rates rise and distances lengthen, electrical loss and power surge, and beyond a certain threshold optical becomes overwhelmingly advantageous. As datacenter inter-rack and intra-rack GPU connections climb from 100G→800G→1.6T, that threshold moves ever closer to the chip, which is why CPO is forecast to have "optical dominate AI datacenter interconnects within five years."
Concretely, in numbers: copper electrical wiring shrinks its effective transmission distance to roughly 1 m at 200 Gbps-class signals, whereas optical transmits tens of m to a few km with low loss. Also, in terms of power, a pluggable transceiver uses roughly 15–30 pJ/bit, but CPO, which extremely shortens the electrical section, has the potential to lower this to less than half, so in large AI clusters with tens of thousands of links this difference accumulates into megawatt-class power savings. This structure, in which "a small saving on one link" is amplified into "an enormous saving across the whole cluster," is the fundamental motivation for hyperscalers to hasten CPO adoption.
Conversely, in domains where these gains are small, optical is actually disadvantageous. For short board wiring inside a server or for small, low-speed links, copper wiring leads in component count, cost, and simplicity, so it is rational to deploy optical selectively only in sections where bandwidth, distance, and scale exceed the threshold. Ultimately electricity and optics coexist not as substitutes but as a complementary relationship that divides roles by tier according to distance, bandwidth, and power demands, and the key point is that this boundary is moving ever closer to the chip as technology advances.
CPO's gains are not free.
First, the thermal coupling problem. When the optical engine (especially temperature-sensitive micro-rings/lasers) sits beside a switch/GPU emitting hundreds of watts, it is exposed to high heat and wavelength drifts or laser efficiency drops. Sophisticated temperature control and cooling design are therefore essential, and this is directly tied to the thermal-tuning power paradox discussed earlier.
Second, reduced maintainability. A pluggable module can be swapped part-by-part on failure, but with CPO an optical-engine defect can lead to disposal/replacement of the entire expensive package, so a KGD (Known Good Die selection) and repair strategy matters. From an operations perspective this expands into a service-design problem of "to what unit is replacement done on failure."
Third, laser reliability. Putting the light source inside the package makes lifetime and reliability in a high-temperature environment the crux, so a design that separates it as an External Laser Source (ELS) is sometimes preferred. In other words, CPO is a design that trades "power/performance gains" for "thermal/repair/reliability costs," and is adopted first where power and bandwidth gains are overwhelming, such as large AI clusters. For this reason, early commercialization tends to concentrate not across the access layer as a whole but in high-density, high-bandwidth domains such as GPU scale-up and switch fabrics.
5. Deep Dive — Recent Trends and the Link to AI Semiconductors
Silicon photonics reached an inflection point in the shift from a research technology to mass commercial infrastructure around 2025. Notably, at GTC 2025 NVIDIA unveiled the network switches Quantum-X (InfiniBand, 1.6T class) and Spectrum-X (Ethernet, 3.2T class) applying silicon-photonics-based CPO. NVIDIA stated that this CPO approach provides about 3.5× power efficiency and about 10× reliability improvement over conventional pluggable transceivers, and it is reported to name its optical-engine platform COUPE (Compact Universal Photonic Engine) and mass-produce it on TSMC's silicon photonics process. This supports the industry outlook that CPO will be a core driver of bandwidth increases in scale-up networking. The industry projects that within the next several years a substantial portion of AI datacenter interconnects will shift to optical, and a structural change is underway in which the center of gravity moves from pluggable transceivers to co-packaging.
Also, in the market, rather than CPO replacing everything at once, a structure is forming in which the LPO (linear drive) examined earlier coexists as a transitional practical alternative. Because LPO removes the DSP while retaining the operational convenience of being pluggable to reduce power, it becomes a short-term solution for operators who cannot bear CPO's thermal/repair risks. That is, a multi-tier coexistence of "pluggable (present) → LPO (transition) → CPO (long term)" is expected to continue for a while, and mixing the three methods according to tier, scale, and power targets becomes the realistic architectural choice.
On the standards/ecosystem side, the trend of treating the optical engine like a chiplet is strengthening. Combined with the [[chiplet]] architecture examined earlier, the direction is to hetero-integrate the compute die, HBM, and optical engine into a single advanced package (2.5D/3D), and discussions to have standards such as UCIe encompass optical I/O are also underway. On the standardization side, industry bodies including the OIF (Optical Internetworking Forum) are organizing the electrical/optical interface and reliability specifications of CPO/co-packaging, laying a foundation that eases vendor lock-in and enables multi-sourcing. As an industry application case, the AI training clusters of hyperscale cloud operators are representative: in scale-up fabrics that bind tens of thousands of GPUs at low latency and high bandwidth, optical interconnects are being adopted as the key to relieving bottlenecks. Domestically as well, SK hynix and Samsung Electronics are pursuing research combining HBM with optical interconnects, and telecom/datacenter operators are investing in advancing optical networks. That said, since CPO's concrete product roadmaps, performance figures, and supply capacity may vary with each company's announcements and market conditions, in actual deployment it is advisable to verify the latest specifications and validation results. (Power-efficiency and reliability figures are per manufacturer announcements and may vary by condition.)
6. Considerations and Implications
From a professional engineer's perspective, adopting silicon photonics/optical interconnects requires comprehensive judgment on the following.
- Application strategy (phased transition): A wholesale shift to CPO carries large thermal, repair, and maturity risks, so at present it is realistic to approach it via a phased roadmap of pluggable → LPO → CPO. A mixed strategy that applies it first to sections where bandwidth/distance demands clearly exceed electrical limits (large-scale GPU scale-up) while keeping proven pluggables in the access layer is effective.
- Break-even from a power/TCO view: Optical's gains come from reduced energy per bit (pJ/bit) and lower cooling cost. Break-even should be analyzed not only on initial adoption cost but on TCO including lifecycle power, cooling, and maintenance, and the higher the data rate and the larger the scale, the greater the optical gain.
- Thermal/reliability/repair strategy: Since temperature-sensitive devices (micro-rings/lasers) sit beside high-heat compute chips, one must prepare in advance, alongside precise temperature control and cooling design, a KGD selection, module replacement, and redundancy strategy to compensate for CPO's maintainability weakness. For the light source, consider a design that lowers reliability risk via external-light-source (ELS) separation.
- Supply-chain/standard dependence: Silicon photonics depends on InP light sources, specialty fabs, and advanced packaging (CoWoS, etc.) capabilities, concentrating the supply chain in a few vendors. Ease dependence via adoption of open standards (optical I/O standardization) and multi-sourcing, and weigh the trade-off between required performance and standard maturity.
- Reorganization of workforce/design competency: Silicon photonics is a domain converging semiconductors, optics, packaging, and systems, so organizations accustomed to electrical design must secure end-to-end design competency that handles optical link budget, heat, and reliability together. In adoption decisions, evaluate both in-house competency and the partner (fab, packaging, light-source) ecosystem.
- Outlook (expansion into compute-in-photonics): Optical may expand beyond interconnects into computation itself, such as optical computing and optical neural networks, and must be recognized as a strategic technology that, combined with [[cxl-compute-express-link]]-based memory disaggregation and composable infrastructure, could reshape datacenter architecture.
References
- Yole Group, "Silicon photonics and co-packaged optics at the heart of next-generation AI-driven data infrastructure", https://www.yolegroup.com/press-release/silicon-photonics-and-co-packaged-optics-at-the-heart-of-next-generation-ai-driven-data-infrastructure/
- SemiEngineering, "All AI Data Center Interconnects Will Be Optical Within 5 Years", https://semiengineering.com/all-ai-data-center-interconnects-will-be-optical-within-5-years/
- EDN, "Where co-packaged optics (CPO) technology stands in 2026", https://www.edn.com/where-co-packaged-optics-cpo-technology-stands-in-2026/
In one line: Silicon photonics is a technology that integrates optical devices onto silicon via the CMOS process to move data with light instead of electrons; it overcomes the distance and power limits of electrical wiring under exploding AI/HPC traffic and, evolving toward CPO (co-packaged optics) that presses the optical-conversion point against the compute chip, is establishing itself as core infrastructure for AI datacenter interconnects.