Research Noteresearch note
Heterogeneous scientific computing across CPU, GPU and QPU systems
Most "quantum computing" work is, in practice, a classical pipeline with a small quantum subroutine somewhere in the middle. The interesting engineering question is therefore not quantum vs. classical but placement: given a workload, which parts belong on a CPU, which on a GPU, which on a simulator, and which — if any — on a QPU.
The layers I keep coming back to
- Data movement dominates. The cost model that matters is bytes moved across the CPU ↔ GPU ↔ QPU boundaries, not raw FLOPs or gate counts.
- Orchestration should be declarative: describe the DAG of tasks and their device affinities, and let a scheduler place them. Hand-wiring device transfers does not scale.
- Reproducibility is harder across heterogeneous hardware — seeds, driver versions, and QPU calibration snapshots all have to be captured, or a result cannot be re-run.
Where quantum subroutines belong
A quantum subroutine earns its place only when it changes the asymptotics of a bottleneck sub-problem and the surrounding classical cost does not swamp the gain. That is a high bar, and stating it plainly keeps the rest of the architecture honest.
This note will grow as the orchestration prototype does.
heterogeneous computingorchestrationreproducibilityQPU
Discussion