Optical computing

How to evaluate photonic-computing performance claims

By OPU Cloud. Published .

A photonic chip can perform an impressive experiment without being the best system for your application. To evaluate a performance claim, first identify what was calculated, then ask what the measurement includes. A large operations-per-second number is useful only when you can connect it to a correct result and the resources needed to produce that result.

This guide is a reading method, not a ranking of products. The examples below are hypothetical calculations, not measurements of commercial hardware.

Start with the complete task

Write the proposed comparison as a sentence: “System A completes this model, on this input, at this quality target, faster or with less energy than system B.” If the sentence has blanks, the headline comparison is incomplete.

A matrix multiplication, one neural-network layer, a full inference, and a training step are different tasks. Ask whether inputs begin as digital data in host memory or as light from a scene. Ask whether outputs are usable predictions, analog voltages, or intermediate features awaiting further processing.

The ACCEL research paper illustrates why the input boundary matters: it investigates an optical and analog-electronic pipeline for vision, including direct processing of optical inputs. Its demonstrated vision tasks should not be treated as a general benchmark of arbitrary digital workloads. Read the original ACCEL paper.

Also distinguish measured results from simulations and scaling projections. A measurement on a small device can support that experiment; a projected larger device needs additional assumptions about fabrication, control, bandwidth, and power.

Understand the operation count

TOPS means trillions of operations per second. It does not tell you what those operations are. Some reports count a multiply and an addition separately; others report multiply-accumulate operations, or MACs. State the convention before comparing numbers.

Consider a hypothetical 128-by-128 matrix multiplied by a vector. It has 16,384 scalar multiplications. Counting one multiplication and one accumulation per term yields about 32,768 operations under a two-operations-per-MAC convention. At one million vectors per second, that convention gives approximately 0.0328 TOPS. Counting MACs instead gives 0.0164 tera-MACs per second.

Neither number changes the work completed. Their numerical difference comes from bookkeeping. Padding, sparsity, reused weights, and unused channels can also change the relationship between peak arithmetic capacity and useful work.

The photonic tensor-core research by Feldmann and colleagues focuses on a computationally specific convolution accelerator. That specificity is part of the result, rather than a reason to extrapolate its headline throughput to every application. Read the authors’ paper.

Compare accuracy before speed

An analog output has measurement error. A nominal “8-bit” interface does not automatically mean every computed result has the numerical behaviour of an 8-bit digital calculation.

For your comparison, record both numerical error and task quality where relevant. A classifier needs an agreed dataset, preprocessing, and accuracy measure. A scientific solver may need a residual tolerance or an error bound. A fast answer outside the required tolerance is not a completed task.

Check whether the photonic and electronic baselines use the same model. If a hardware-aware retraining procedure changes the model, explain that change and measure its quality. It may be a valuable engineering choice, but it answers a different question from replacing hardware under an unchanged model.

MLPerf Inference is a useful example of disciplined comparison: its benchmarks define datasets and quality targets, and distinguish testing scenarios with different request patterns. Citing its principles does not imply a photonic device has submitted a valid MLPerf result. MLCommons benchmark guidance.

Draw the power boundary

Use a simple table to make omitted resources visible.

Measurement boundaryResources to ask about
Optical coreOptical energy passing through the computation
Accelerator subsystemLasers, drivers, detectors, converters, calibration and control
Complete serverHost CPU, memory, accelerator, power-supply losses and other active devices
ServiceNetworking, idle capacity, scheduling and any separate preprocessing

Compare like boundaries. Optical energy alone and wall-plug server energy are different measurements. MLCommons describes its validated power metric as measured AC energy or power for the entire system during the accompanying benchmark, rather than a component’s rated power. Power measurement explanation.

For an original hypothetical example, suppose an accelerator completes 2,000 accepted inferences per second while drawing 300 watts at the chosen system boundary. Its energy is 300 / 2,000 = 0.15 joules per inference. If including the host changes the total to 450 watts, the same throughput costs 0.225 joules per inference. Reporting the boundary changes the conclusion.

Inspect latency and reproducibility

Throughput describes how much work finishes over time. Latency describes how long an individual job waits. A large batch can raise throughput while making a single request slower. Include data transfer, weight programming, conversion, and required calibration in the timing boundary.

Look for enough detail to reproduce the experiment: hardware configuration, model, precision, software versions, input shapes, batch size, warm-up, trial duration, and variability. Check whether the electronic baseline uses an appropriate optimised implementation and similar quality constraints.

Before accepting a claim, keep five items together: completed task, quality target, operation convention, measurement boundary, and evidence level. If one is missing, record the claim as incomplete rather than guessing the missing value.

Sources

Related reading