Cloud infrastructure

Where photonics fits in an AI data centre

By OPU Cloud. Published .

An AI data centre must calculate, store, and move data. Photonics can contribute to more than one of those jobs, but the technologies and evidence differ. Optical networking carries information between devices. Optical I/O moves conversion closer to computing chips. Photonic accelerators investigate using optical transformations to perform selected calculations.

Understanding those layers makes it easier to judge a new announcement and to identify which bottleneck it might address.

Map the layers before choosing technology

Start with a simplified system rather than a vendor’s product names.

LayerJobPossible role for photonics
Links between network devicesCarry messages over the facilityOptical transceivers and fibre
Communication near processorsMove data into or out of chipsOptical I/O and integrated optical engines
Numerical compute stageTransform inputs into resultsA task-specific photonic accelerator
Host and orchestrationPrepare jobs, schedule work, manage stateElectronics and software coordinate the complete pipeline

These layers are not interchangeable. Cisco’s transceiver explanation concerns signal conversion and communication. The photonic tensor-core paper by Feldmann and colleagues concerns a specific computational accelerator. Read each as evidence about its own function. Cisco optical-link explanation, tensor-core research.

An installation can adopt better optical networking without changing the arithmetic hardware. It can also evaluate an optical accelerator while retaining a conventional host, memory system, and network.

Why communication can become part of the calculation time

A distributed job often needs several devices to exchange intermediate results. Moving those results is part of completing the application, even though the link itself may perform no model arithmetic.

NVIDIA’s NCCL library exposes collective communication operations across GPUs and nodes, including all-reduce and all-gather. Its AllReduce example combines contributions from participants and makes the reduced result available to each participant. NCCL overview, AllReduce example.

As a small original example, suppose two workers produce scalar contributions 7 and 5. A sum reduction produces 12; an all-reduce makes 12 available to both workers. The application needs a communication operation to coordinate that result. A faster accelerator on one worker cannot remove the other worker’s contribution or the exchange itself.

Real collective performance depends on message size, topology, software, congestion, and overlap with computation. A link’s rated bandwidth is therefore one input to application performance, not a direct prediction of it.

Work through a bottleneck example

Consider a hypothetical training step with no overlap between stages:

Arithmetic             60 ms
Communication          30 ms
Other work             10 ms
Total                 100 ms

If a networking improvement halves communication time, the new total is 60 + 15 + 10 = 85 milliseconds. The step speedup is 100 / 85, approximately 1.18 times.

If a compute improvement halves arithmetic time instead, the total is 30 + 30 + 10 = 70 milliseconds, or about 1.43 times faster.

If both improve as assumed, the result is 30 + 15 + 10 = 55 milliseconds, about 1.82 times faster. None of these is a forecast for photonic hardware. They demonstrate why you need the time breakdown before estimating value.

In a real system, stages may overlap. Then adding their independent durations overstates elapsed time. Measure the critical path and wait time rather than using this simple model unchanged.

Optical I/O moves the conversion boundary

Placing optical communication closer to a processor targets a different part of the data path from replacing a numerical core. Intel’s optical I/O chiplet demonstration shows this integration around a CPU. It is evidence of a communication architecture, and its original announcement must be read with its demonstration and development context. Intel optical I/O announcement.

For an operator, useful questions include the supported protocol, reach, energy per delivered bit, error behaviour, thermal requirements, and replacement procedure. For a developer, many of those changes may remain below the application interface.

A photonic compute accelerator creates different questions: which operators run on it, how input and output conversion work, what precision is supported, and whether the software can map the intended model. A transport upgrade and a compute upgrade therefore need separate validation.

Separate infrastructure from cloud access

A cloud provider using optical links does not automatically offer users optical computation. A user-facing service needs an actual access route, supported workloads, documentation, and a way to distinguish hardware execution from simulation.

When assessing a service, ask for an example job that completes on identified hardware. Check what can be controlled, whether access is public or restricted, what result is returned, and how failures are reported. A simulator may be useful for learning an API without proving that remote hardware is available.

For an infrastructure announcement, ask where the technology sits and which measured problem it addresses. For a compute announcement, ask which task it completes and how the whole system is evaluated. Photonics is easiest to understand when each claim is attached to a layer, a bottleneck, and a verifiable result.

Sources

Related reading