Cornelis raised $205M, and the practical issue is idle GPU time
On September 14, 2026, TechCrunch reported that Cornelis Networks raised $205 million to expand its AI networking fabric business. The company is positioning its fabric as a way to reduce the time AI accelerators spend waiting for data to arrive from other nodes, storage systems or peer accelerators. The strategic backdrop is competition around the infrastructure stack that sits beneath AI training and inference: accelerators, interconnects, switches, storage paths and the software that decides where work runs.
For operators of a quarter rack to a few racks of GPUs, the useful question is not whether a new fabric challenges any one vendor. The useful question is narrower: how much of your paid-for GPU time is lost to data movement, and can you recover enough of it to change the economics of the rack?
GPU underutilisation is often discussed as if it were only a scheduling problem. Sometimes it is. A node is idle because no tenant has a job queued, or because a VM was left running after a notebook session ended. But on real clusters, especially mixed Kubernetes and VM estates, idle spend also appears while a GPU is technically allocated. The tenant sees a running job. The platform sees the device assigned. Finance sees a GPU-hour. Yet the accelerator may be stalled because input pipelines, checkpoint writes, model parallel traffic, storage reads, or east-west network traffic cannot keep up.
That is where fabric matters. Faster or lower-latency networking can reduce stalls for distributed training, retrieval-heavy inference, parameter synchronization, checkpointing and storage access. But fabric alone does not create a governed GPU service. A small operator still needs provisioning, tenancy, GPU partitioning, metering, quota policy, storage placement, power accounting and a way to explain chargeback to tenants.
ClastIQ, a product of Cognition AI and Technology Innovation SPC, is aimed at that operating layer: hardware-agnostic orchestration for GPU and CPU estates owned by the operator, not rented from a hyperscaler. The point is to make the cluster measurable and governable enough that improvements in fabric, storage or scheduling convert into higher effective utilization, not just better lab results.
The bottleneck operators actually hit
A small GPU estate does not usually fail in one dramatic place. It loses utilization in increments.
A data science team reserves four full GPUs for a notebook but only needs bursts of acceleration. A computer vision team runs inference on a full device when a MIG slice or time-sliced share would be enough. A training job spans multiple nodes and spends a measurable share of wall-clock time waiting on gradient exchange. A tenant writes checkpoints to shared storage every few minutes, causing storage congestion that slows other jobs. A VM tenant pins GPUs for a long-running environment, while Kubernetes users queue for capacity. A sovereign workload must stay on-premises or within a national boundary, so the operator cannot simply spill to a public cloud region when the queue grows.
The result is familiar: utilization dashboards look acceptable if they measure allocation, but finance still sees poor revenue per megawatt and users complain about wait times.
Operators should separate three utilization numbers:
| Metric | What it measures | Why it matters |
|---|---|---|
| Allocated GPU utilization | GPUs assigned to pods or VMs | Useful for capacity planning, but can hide stalls |
| Active device utilization | Time the accelerator is doing useful work | Shows data, scheduling and code bottlenecks |
| Billable productive utilization | Tenant-accounted GPU-hours that meet policy | Connects engineering to chargeback and revenue |
Cornelis is speaking to the second line: reducing wasted device time caused by data movement. ClastIQ addresses the system around all three lines. It provisions the bare metal, creates Kubernetes and VM tenancy on the same fleet, partitions accelerators where useful, meters per tenant, enforces quotas, and gives the operator a way to recover cost through transparent chargeback.
Fabric helps most when the rest of the estate is disciplined
A better network fabric can improve cluster behavior, but only if the workload placement and storage paths allow it to matter.
For example, a distributed training job may benefit from lower latency between nodes, but not if the scheduler spreads it across a congested part of the rack or mixes it with noisy neighbors doing checkpoint-heavy writes. A retrieval-augmented inference service may need fast access to vector indexes, but fabric improvements will not help if Ceph pools are undersized, placement groups are unhealthy, or the hot dataset sits behind a slow path. Multi-tenant GPU sharing can raise headline utilization, but without quotas it can also create unpredictable tail latency for paying tenants.
This is why operators should treat fabric as one component in a control plane, not as a standalone cure. The operating model should answer practical questions:
- Which jobs need full GPUs, and which can use MIG slices or time-sliced access?
- Which workloads need node-locality, rack-locality or storage-locality?
- Which tenants are allowed to burst, and who gets preempted first?
- Which network paths are reserved for storage, tenant traffic, management and east-west accelerator traffic?
- Which idle allocations are billed, reclaimed or flagged to the tenant?
- Which workloads must remain in-country, in-facility or in an air-gapped environment?
ClastIQ is designed for these decisions on operator-owned hardware. It uses bare-metal provisioning with MAAS and Juju patterns, supports Kubernetes and VM tenancy on the same physical fleet, and integrates storage and metering so GPU use can be governed instead of guessed. The estate can be NVIDIA, AMD, CPU-heavy, GPU-heavy or mixed. Fabric choice should be made against measured workloads, not vendor preference alone.
A worked example: 32 GPUs and a fabric-related stall
Consider a four-node GPU pod in a quarter-rack deployment, with eight accelerators per node: 32 GPUs total. Assume the operator runs a mix of training, batch inference and interactive notebook workloads.
There are 32 GPUs x 24 hours x 30 days = 23,040 calendar GPU-hours per month.
Assume the all-in internal cost of one calendar GPU-hour is $1.36. That includes amortised accelerator and server capital, power at the facility meter, cooling overhead, support, storage and network allocation. This is not a market price; it is an accounting input for the example.
If only 35% of calendar GPU-hours become billable productive usage, the estate produces:
23,040 x 0.35 = 8,064 productive GPU-hours per month.
At an internal chargeback rate of $4.00 per productive GPU-hour, monthly recovery is:
8,064 x $4.00 = $32,256.
Now suppose the operator improves the data path and orchestration together. Fabric upgrades reduce distributed job stalls. Ceph pools are separated for checkpoint-heavy workloads. Kubernetes placement policies keep multi-node jobs within the best-connected part of the rack. MIG and time-slice policy move small inference and notebook jobs off full devices. Idle VM GPU assignments are metered and reclaimed after a policy window. Productive utilization rises from 35% to 55%.
23,040 x 0.55 = 12,672 productive GPU-hours per month.
At the same $4.00 chargeback rate:
12,672 x $4.00 = $50,688 per month.
The difference is $18,432 per month, or $221,184 per year, on the same 32 GPUs.
Now translate that to revenue per megawatt. If the quarter rack consumes 40 kW of IT power and the facility runs at a 1.35 PUE, the facility draw is 54 kW. At 35% productive utilization, $32,256 per month divided by 0.054 MW equals about $597,333 per MW-month. At 55%, $50,688 divided by 0.054 MW equals about $938,667 per MW-month.
The fabric improvement is not valuable because it is faster in isolation. It is valuable if it moves the cluster from allocated-but-waiting to productive-and-billable, while preserving tenant boundaries and service policy.
What to measure before buying fabric
Before changing the network layer, operators should build a baseline. At small scale, it is possible to inspect enough detail to avoid making a blind purchase.
Start with accelerator metrics. Measure active device utilization, memory utilization, power draw, PCIe or accelerator interconnect counters where available, and time spent waiting in the input pipeline. Then correlate those metrics with Kubernetes pods, VM tenants, storage pools and network interfaces. The correlation is more important than any single metric.
A short diagnostic loop can be as simple as this, adapted to the tools in the estate:
# Map GPU workloads to tenants and nodes
kubectl get pods -A -o wide | grep -E 'gpu|accelerator'
# Check storage health before blaming the accelerator fabric
ceph -s
ceph osd df
# Inspect link errors and dropped packets on candidate data paths
ip -s link show
# Review recent quota pressure and pending GPU workloads
kubectl get resourcequota -A
kubectl get pods -A --field-selector=status.phase=Pending
For NVIDIA estates, vendor tools can expose device-level activity and MIG allocation. For AMD estates, use the equivalent ROCm and platform metrics. For mixed estates, normalize the data into tenant, node, device, job and time dimensions. The fabric conversation should start only after the operator can say which workloads are network-sensitive and how much billable time is lost.
ClastIQ helps by tying the infrastructure view to tenancy. A stalled GPU that belongs to a research tenant, a paid inference tenant and an internal platform job should not be treated the same way. The operator needs to know who consumed the allocation, whether policy allowed it, whether the chargeback record should include it, and whether a quota or placement change would prevent recurrence.
Kubernetes and VM tenancy complicate the fabric question
Many small GPU estates are not pure Kubernetes clusters. They usually need both Kubernetes and VMs.
Kubernetes is suitable for batch jobs, inference services, notebook platforms and pipelines that can be containerised. VMs remain necessary for tenants with custom images, licensed software, older drivers, specific security controls, or development environments that do not fit a container workflow. The same physical GPU fleet often has to serve both.
That mixed model can waste fabric capacity if not orchestrated. A VM may reserve full GPUs and sit mostly idle while Kubernetes jobs queue. A Kubernetes namespace may burst into storage-heavy jobs and affect VM tenants. A multi-node training job may be placed across nodes that are technically available but poorly connected for its traffic pattern.
ClastIQ's role is to make the mixed estate governable. Bare-metal provisioning brings nodes into a known state. Kubernetes and VM tenancy share the same capacity plan rather than becoming separate islands. GPU partitioning lets operators decide when a full device is required and when MIG or time-slice access is enough. Metering connects actual consumption to the tenant. Quotas prevent a single user from turning fabric, storage or power into a shared bottleneck.
Fabric investment then becomes a capacity-planning decision. If the queue is caused by idle VM reservations, buy less fabric and fix policy first. If training jobs are active but waiting on all-reduce or storage reads, fabric and storage paths may be the right next constraint to address. If inference latency is caused by CPU preprocessing or small batch sizes, accelerator networking will not be the first lever.
Storage is part of the fabric story
Data movement is not only node-to-node traffic. It is also storage-to-node and node-to-storage traffic. For small estates, Ceph is often used because it gives the operator a sovereign, on-premises storage layer without depending on a hyperscaler service. But Ceph still needs design discipline.
Checkpoint-heavy training can generate bursts that collide with dataset reads. Object workloads and block volumes may have different latency and throughput needs. VM images, container registries, model weights, training data and tenant outputs should not all be treated as one undifferentiated pool.
An operator-grade design separates storage classes and measures them against workload patterns. Hot datasets may need replication and placement close to the GPU nodes. Checkpoint traffic may need its own pool, schedule or quality-of-service policy. Tenant data may require encryption, retention policy and locality controls. Sovereign workloads may require that replicas never leave a facility, city or national boundary.
In GCC and MENA deployments, this is not an academic point. Operators serving government, education, healthcare, finance or national AI programs often need capacity inside a sovereign environment. Some sites are air-gapped or semi-connected. Importing a hyperscaler operating model is useful, but exporting data to a hyperscaler region may not be allowed. The rack must therefore deliver utilization and governance locally.
ClastIQ's value in that context is not that it replaces fabric. It gives the operator the control plane to make fabric, storage and GPU policy work inside the boundary: per-tenant metering, quotas, chargeback, provisioning, and local human-led support on the operator's own hardware.
Power and revenue per megawatt should drive priorities
GPU operators should evaluate fabric through power economics. A faster data path that raises productive utilization without increasing the power envelope can improve revenue per megawatt. A fabric change that adds power but does not remove stalls may make the facility less efficient.
This is why metering should include power and not only GPU-hours. Tenants that run low-utilisation full-GPU VMs are consuming capacity, power and cooling headroom. Tenants that run well-packed inference slices may produce more revenue per watt. Training jobs that saturate accelerators but create storage congestion may reduce the estate's total output even if their own GPUs look busy.
Policy should reflect those trade-offs. Operators can use quotas for maximum GPU count, maximum simultaneous full-device allocations, maximum burst duration, storage throughput classes and priority tiers. Chargeback can include higher rates for reserved full GPUs, lower rates for shared slices, and penalties or alerts for idle allocations beyond a grace period. These are engineering controls, not accounting paperwork.
ClastIQ provides the framework to implement these controls across a small or mid-size estate. It does not require the operator to standardize on one accelerator vendor or one tenant model. The estate can evolve: start with bare-metal provisioning and Kubernetes, add VM tenancy, introduce MIG or time-slicing, segment Ceph storage, then tune quotas and metering as real usage appears.
How to evaluate Cornelis-like fabric claims
When a fabric vendor says GPUs waste time waiting for data, assume the statement may be true and still require proof on your workloads.
Ask for tests that match your topology: number of nodes, GPUs per node, storage design, tenant mix, model sizes, checkpoint frequency and batch sizes. Include both best-case distributed training and messy multi-tenant periods. Measure wall-clock job completion, active accelerator utilization, storage latency, network retransmits or drops, queue time, and chargeback output.
Avoid testing only a single benchmark that saturates the interconnect. Benchmarks are useful for maximum capability, but operators need sustained estate behavior. The question is whether the fabric reduces idle spend after policy, storage and placement are included.
A practical acceptance test should answer four questions:
- How many additional productive GPU-hours per month does this create?
- Which tenants or workload classes benefit?
- What new power, support and operational complexity does it add?
- How does it change revenue per megawatt and payback period?
If the answers are measurable, the fabric decision becomes an investment case. If they are not, the operator is still guessing.
What to do this week
- Build a 30-day baseline of allocated GPU-hours, active device utilization and billable productive GPU-hours by tenant.
- Identify the top five jobs with the largest gap between allocated time and active accelerator time; classify the cause as data input, storage, network, CPU preprocessing, idle VM reservation or user behavior.
- Separate storage measurements for datasets, checkpoints, VM images and model artifacts; do not average them into one Ceph number.
- Review quota policy for full-GPU reservations, MIG or time-slice access, burst rights and idle reclamation.
- Model one utilization improvement scenario, such as 35% to 55%, and express it in GPU-hours, chargeback recovery and revenue per megawatt.
- If evaluating new fabric, test it with your tenant mix and placement policy, not only a vendor benchmark.
ClastIQ runs this operating model on the operator's own hardware. Request a demo.
