Need GPUs Now?

Tell us who you are and we'll bring up the agreement for you to sign right here.

400 / 800G
400 / 800G
PER PORT
~1 µs
~1 μs
SMALL-MESSAGE LATENCY
Rail-optimized
Rail-optimized
NON-BLOCKING
NCCL
NCCL
VALIDATED AT HANDOVER

The Constraint

Most Underperforming
AI Fabrics Aren't Badly
Chosen. They're Badly
Tuned.

The received wisdom that Ethernet is dramatically slower than InfiniBand for training mostly describes misconfigured networks rather than a ceiling in the technology. The catch is that configuring one correctly demands expertise most networking teams haven't needed before.

Lossless behavior on Ethernet has to be engineered

Lossless behavior on Ethernet has to be engineered

RoCEv2 emulates losslessness using priority flow control and congestion notification. Configured well, it performs. Configured carelessly, it collapses under burst and produces tail latency that ruins collective operations — and the symptom looks like a slow cluster, not a network fault.

01
Collectives punish the worst link, not the average

Collectives punish the worst link, not the average

Synchronous gradient exchange means every GPU waits for the slowest participant. A single interface flapping for microseconds stalls a distributed job, and the failure surfaces as a hung training run rather than an alert on a dashboard.

02
Topology decides throughput more than port speed

Topology decides throughput more than port speed

Oversubscription ratios and rail alignment set the real ceiling. A fast fabric wired as a conventional oversubscribed campus network will not deliver its rated bandwidth to a training job, no matter what the optics are rated for.

03

How we do it

How We Arrive
At A Fabric,
Then Prove It.

The decision is made against your workload profile in the open, and the build is validated under collective load before anyone calls it finished.

Profile the collectives

Profile the collectives

We characterise what your jobs actually do on the wire: message-size distribution, all-reduce and all-to-all patterns, expert routing if you're running mixture-of-experts, and how tightly coupled the work is. Loosely coupled fine-tuning and large synchronous pre-training have genuinely different fabric requirements.

OUTPUT COLLECTIVE AND TRAFFIC PROFILE
Build the comparison matrix

Build the comparison matrix

InfiniBand, RoCEv2 on merchant-silicon Ethernet, and AI-optimized Ethernet are scored against your profile on latency, scale ceiling, cost per port, supply-chain breadth and the operational skill each demands. The premium options have to justify themselves in your numbers, not in general.

OUTPUT SCORED FABRIC MATRIX, ONE PAGE
Design the topology

Design the topology

Rail-optimized, non-blocking spine-leaf sized to your node count and growth, with oversubscription stated explicitly rather than buried. Storage and management traffic are planned alongside the back-end fabric, not bolted on later.

OUTPUT TOPOLOGY AND CABLE PLAN
Build and cable

Build and cable

Switches, adapters, optics and cabling procured and installed to plan. Optics selection and clean termination matter more than they should — contamination and marginal transceivers are a leading cause of the intermittent link errors that later look like software problems.

OUTPUT BUILT FABRIC, LABELLED AND MAPPED
Tune congestion control

Tune congestion control

Flow control, congestion notification and adaptive routing are configured and then tested under deliberate burst, not left at defaults. For Ethernet builds this step is the difference between a fabric that matches InfiniBand in practice and one that doesn't.

OUTPUT TUNED CONFIGURATION, DOCUMENTED
Validate under collective load

Validate under collective load

The fabric is burned in and benchmarked with real collective operations at full scale, checking for link errors, flap, throughput against design and consistency across rails. You get numbers at handover, not assurances.

OUTPUT BENCHMARK REPORT AND BURN-IN RESULTS
favcon

What's included

What The Fabric
Unit Covers.

InfiniBand Builds

InfiniBand Builds.

Current-generation switching and adapters with subnet management.

RoCEv2 on Ethernet

RoCEv2 on Ethernet.

Merchant-silicon fabrics engineered for lossless behavior.

AI-Optimized Ethernet

AI-Optimized Ethernet.

Adaptive routing and congestion control for larger deployments.

Staging and Burn-in

Staging and Burn-in.

Rail-optimized non-blocking spine-leaf with stated over subscription.

Optics and Cabling

Optics and Cabling.

Transceiver selection, structured fibre, clean termination and labelling.

In-Network Reduction

In-Network Reduction.

Switch-based collective acceleration where the hardware supports it.

Congestion Tuning

Congestion Tuning.

Flow control and congestion notification tested under burst.

Storage Fabric Integration

Storage Fabric Integration.

RDMA paths for GPUDirect planned with the back-end network.

Burn-in and Benchmarking

Burn-in and Benchmarking.

Collective-level validation before handover.

Ongoing Fabric Monitoring

Ongoing Fabric Monitoring.

Error counters and flap detection under the management unit.

Fabrics Compared

Three Fabrics. Different
Trade-offs, Not Different
Quality.

The row that decides it for most teams isn't latency — it's operational depth and supply chain. Be honest about which you have.

Latency floor
Losslessness
Vendor ecosystem
Cost per port at 400G+
Operational skill needed
Shares tooling with your DC network
Strongest case

InfiniBand

Lowest
Native, credit-based in hardware
Effectively single-source
Highest
HPC networking specialists
No, second management plane
Very large tightly-coupled training

RoCEv2 on
Ethernet

Slightly higher, tunable
Emulated via flow control
Broad, multi-vendor
Lowest
Deep flow-control expertise
Yes
Cost-sensitive builds, Ethernet-standardized teams

AI-optimized
Ethernet

Between the two
Emulated, with adaptive routing
Narrower than commodity Ethernet
Middle
Ethernet skills, plus tuning
Yes
Large Ethernet builds needing scale headroom

Questions

Before You
Commit

We won't answer that before step one, and any vendor who does is guessing. For a large, tightly coupled pre-training cluster the latency floor and deterministic behavior of InfiniBand are worth paying for. For inference serving, fine-tuning, or a team already standardized on Ethernet with Kubernetes tooling, a properly tuned Ethernet fabric usually wins on total cost and operational simplicity. The profile decides it.

Yes. Many mature AI footprints deploy InfiniBand strictly for backend GPU-to-GPU compute communications (all-reduce/all-gather), alongside an AI-optimized RoCEv2 or standard Ethernet fabric for front-end ingestion, storage IO, and cluster orchestration.

Yes. We regularly perform fabric audits on existing clusters, inspecting PFC/ECN buffer headroom, asymmetric routing traps, transceiver error rates, and rail alignment mismatches to recover lost collective throughput.

For current-generation accelerators (such as H100/H200 and Blackwell architectures) scaling past several racks, 800G significantly cuts cable count and switch tiers. For smaller training clusters or inference footprints, 400G often offers lower cost per port with identical workload performance.

Bring us your job profile.

We'll bring
the one-page
matrix.

Scored against your workload and your team's operational depth — not a generic vendor comparison chart.

Talk to an engineer right-arrow-circle

The Rest of the Stack

Six Units. One
Assembled Stack.

Note on figures. Latency, throughput and collective efficiency vary with message-size distribution, cluster topology, rail alignment and tuning parameters. Benchmark figures shown reflect reference architectures under controlled conditions, not an unconditional guarantee across all workloads.