← all posts
// hardware · hardware

Cornelis wants to put compute inside the network, because your GPUs are waiting on it

Cornelis Networks raised $205 million, led by IAG Capital, and at the same time introduced something it calls Active Compute Fabric. The company is an Intel spin-off from 2020, so this is not a garage startup discovering networking. The pitch fits in one sentence: put programmable compute inside the network fabric itself, so chips can compute and send data at the same time instead of taking turns.

I care less about the funding round than about the sentence behind it. Cornelis says GPUs spend a lot of their time waiting for data. That is the claim the whole product hangs on, and it is worth taking seriously even though the source gives me no numbers for it.

Where the idle GPU comes from

Every large training or inference job has a phase where accelerators stop computing and exchange results. Gradients get summed across the cluster, activations get shuffled between pipeline stages, expert outputs travel between nodes. The classic arrangement is strictly sequential: compute, then communicate, then compute again. During the communicate step the expensive silicon sits there.

You can hide part of it with overlap tricks, and everyone does. But overlap only works as far as the software can find independent work to schedule. Past a certain cluster size the network stops being plumbing and becomes part of the critical path.

A toy calculation, and I want to be clear it is an illustration and not a Cornelis figure. If a training step spends 30% of its wall-clock time blocked on the network, then 1,000 GPUs deliver the useful throughput of about 700. Cut the blocked share to 20% and you have effectively bought 100 GPUs without buying them. That is why anyone pays attention to interconnect at all: the arithmetic is brutal even when the improvement is modest.

What compute in the fabric would change

The idea of doing arithmetic inside the switch is older than Cornelis. If the network can already touch every packet passing through, it can also do the reduction on the way instead of hauling every partial result to an endpoint and back. Fewer bytes cross the wire, fewer round trips, and the GPU is released earlier.

What the brief tells me is narrower than that. Cornelis describes the architecture as integrating programmable compute into the fabric, with the goal that chips compute and send simultaneously. It does not tell me which operations run in the network, how programmable it really is, or what a developer has to change in their collective communication code to benefit. Those three answers decide whether this is a drop-in upgrade or a research project you adopt at your own risk.

Once the network does arithmetic, it stops being a cable choice and becomes part of the compute budget you have to plan for.

Shipping status and the competition

The status is easy to state. CN5000 is shipping. CN6000 is sampling, with broader availability targeted for Q4 2026. The announcement came together with a collaboration with Qualcomm, and Cornelis positions itself against Nvidia, Cisco and Arista.

Sampling means a small number of customers have hardware and nobody outside their NDAs has published results. So there are no benchmarks in this article, on purpose. I haven't tested any of it, and I'd distrust any vendor number that arrives before independent ones.

The competitive picture also deserves one honest caveat. Nvidia sells the GPU and the fabric as a matched pair, and buyers of that stack get one vendor to blame and one tuning guide. A challenger has to beat that on total cluster throughput, not on a switch spec sheet, and it has to do it with the collective libraries people already run.

What I would ask before caring

If you plan racks rather than read press releases, the question is not whether interconnect matters. It does. The question is which fraction of your steps is actually network-blocked today. You can measure that with profiler traces from your own jobs: look at the gaps between kernels during collectives. If that gap is 5% of step time, a smarter fabric is a rounding error for you. If it is 25%, you have a reason to ask for CN6000 samples in Q4 and to write your own test plan before the vendor writes one for you.

Do that measurement first. Then the announcement becomes a hypothesis you can check instead of a story you have to believe.

#hardware#interconnect#gpu#infrastructure