RACK TRAIL
Solutions
VMware MigrationOpenStack at scaleProxmox at scalePrivate AICloud Cost OptimizationManaged Private CloudLarge Scale MigrationBig Data InfrastructureConfidential Computing InfrastructureCephKubernetes WorkloadsCloud Infrastructure for Startups
Why RacktrailGroundworkAbout Talk to an Engineer
13 minutes

Five places new GPU clouds struggle, and what to do about each

Utilization, failing GPUs, tenant isolation, metering and enterprise trust: where new GPU clouds struggle, with public data and what to fix first.

Buying GPUs is the easy part of starting a GPU cloud. The hard part is running them as a service that a bank, a telco or a research lab will trust with production work, and that is where most new providers stall.

A GPU cloud (people now say “neocloud”) rents GPU compute by the hour, the token or the month, usually to AI teams that can’t buy or run their own hardware. In Nepal alone, at least three have launched or opened early access in 2026, ranging from startups renting consumer RTX cards to H200 racks in Kathmandu. ABI Research expects North America’s share of neocloud revenue to fall from 88% in 2026 to 72% by 2030 as sovereign GPU clouds spread to other regions.

At Racktrail we design and operate open infrastructure for cloud providers. Below are the five areas where we see new GPU clouds struggle most, with public data on each and what I’d fix first. None of it needs a hyperscaler’s budget.

Key takeaways

  • A GPU cloud breaks even at roughly 70% utilization, and H100 rental prices have swung from $7–10 an hour to under $2 and back up since 2023.
  • GPUs fail often: Meta’s 16,384-GPU Llama 3 cluster hit an interruption about every three hours, so health checks and automatic draining matter from day one.
  • Containers alone don’t isolate tenants. NVIDIAScape (CVE-2025-23266) showed a three-line container escape affecting 37% of cloud environments.
  • Customers expect an API and an invoice that matches their dashboard, and enterprises expect an SLA, SOC 2 and datacenter-class hardware.

Utilization decides whether the business works

A GPU cloud makes money only while its GPUs are rented, and the breakeven sits around 70% utilization. Moduledge’s model of a 1,024-GPU H100 cluster loses about $330,000 a month at 55% utilization and earns about $340,000 at 85%. Bare-metal gross margins of 55–65% before depreciation leave very little room between those two numbers.

The price you can charge also moves under you. H100s rented for $7–10 an hour in 2023 and fell to $2–4 by late 2025, according to Silicon Data. Then demand turned, and the rate rose about 40%, from $1.70 in October 2025 to $2.35 in March 2026 (Crypto Briefing). A business plan built on one price point will be wrong within a year, in one direction or the other.

Small markets make this harder. A launch fills up with students and researchers who rent a card for a few hours, and per-second billing (which customers love) leaves gaps between jobs that nobody pays for. Depreciation keeps running through those gaps. Most operators assume a 5–6 year useful life, and some analysts argue the real economic life is closer to 2–3 years.

What to fix first

  • Track utilization per GPU every week, split into reserved, on-demand and idle. Most new providers can tell you their revenue but not this number.
  • Sell committed capacity (monthly or annual reservations) alongside on-demand. Reserved contracts pay for the hardware, and on-demand fills what’s left.
  • Slice cards for inference and notebooks. MIG partitions let one H100 or H200 serve up to seven small jobs, which turns an idle full card into paid partial ones.
  • Model your prices at a 30% lower rate than today. If the business still works, a price drop becomes an inconvenience instead of a crisis.

GPUs fail more often than new providers plan for

GPUs are the least reliable part of the server, and a cloud has to catch their failures before customers do. When Meta trained Llama 3 on 16,384 H100s, the cluster hit 419 unexpected interruptions in 54 days, about one every three hours (Data Center Dynamics). Faulty GPUs caused 148 of them and HBM3 memory another 72. CPUs caused two.

419unexpected interruptions in 54 days on Meta’s 16,384-GPU Llama 3 training cluster, roughly one every three hours.

Apply the same per-GPU rate to a single 72-GPU rack and you get roughly one interruption a month. That sounds manageable until it lands on a customer’s three-day fine-tuning run at 2am, with nobody watching. Meta kept effective training time above 90% because automation handled almost every failure. Engineers stepped in manually only three times.

SemiAnalysis rates GPU clouds in its ClusterMAX system, which covers about 90% of the rental market. One of its clearest dividing lines is health checking: smaller providers typically lack automated weekly active checks such as NCCL tests and DCGM diagnostics.

In emerging markets, power adds a second failure source. At a Kathmandu GPU cloud launch in July 2026, the operator said plainly that uptime would be “probably not 99.99”, and pointed to winter power constraints that hit every facility in the country (TechSansar).

What to fix first

  • Burn in every node before a customer touches it, with DCGM diagnostics and NCCL all-reduce tests under sustained load. Early failures show up in the first days.
  • Run passive checks all the time (XID errors, ECC counts, thermals, link flaps) and active checks on a schedule.
  • Drain a failing node automatically and move the tenant to a spare. Keep a small pool of hot spares for exactly this.
  • Tell customers how to checkpoint, and give them a status page. A failure they were warned about costs you far less trust than a silent one.

Sharing GPUs safely is harder than sharing CPUs

Sharing GPUs between tenants is how small clouds push utilization up, and it is also where isolation breaks. GPUs were built for one owner. Their memory isolation is weaker than a CPU’s, and a container shares the host kernel with every other container on the node.

The risk is concrete. In July 2025, Wiz disclosed NVIDIAScape (CVE-2025-23266, CVSS 9.0), a container escape in the NVIDIA Container Toolkit that needed a three-line Dockerfile and reached root on the host. Wiz estimated it affected 37% of cloud environments (The Hacker News). A provider that isolates tenants with containers alone was one malicious customer away from exposing every other customer on that node. ClusterMAX lists container-only isolation as a basic weakness for the same reason.

The sharing method you pick sets how much isolation you get:

MethodIsolation between tenantsFitsWatch out for
Time-slicingNone for memory or faultsOne team’s dev and testNever between paying customers
MIG partitionsHardware paths through memory, up to 7 per GPUInference, notebooks, small jobsDatacenter GPUs only (A100, H100, H200, Blackwell); no MIG on GeForce cards
vGPUPer-VM, with IOMMU protectionEnterprise VMs, mixed workloadsNVIDIA licensing cost
Full card or passthroughSingle tenantTraining, regulated dataLowest utilization if you can’t sell whole cards

Source: Introl, NVIDIA documentation. For how these patterns run on an open control plane, see our guide to running GPUs, FPGAs and DPUs on OpenStack.

What to fix first

  • Put a VM or microVM boundary (for example Kata Containers) around every untrusted tenant. Kubernetes can run on top, but it shouldn’t be the only wall.
  • Patch the NVIDIA stack on a schedule. The NVIDIAScape fix shipped in Container Toolkit 1.17.8 and GPU Operator 25.3.1.
  • Wipe GPU memory and local NVMe between tenants, and default-deny network traffic between them.
  • Keep tamper-evident audit logs of who ran what, where. Enterprise buyers will ask for them in the first security review.

A cloud needs a control plane, and most new ones launch without one

Many new GPU clouds launch as a contact form and a chat group: a customer asks for GPUs, an engineer provisions them by hand, and someone sends an invoice at the end of the month. That works for ten customers. At fifty it eats the engineering team, and the customers who matter most (AI teams used to AWS or RunPod) leave before they get that far. They expect an API, a Terraform provider and a console, and Hosted.ai’s guide for telcos flags ticket-based provisioning as a gap to close.

Metering is the part that quietly breaks. GPU billing has to account for GPU model, memory, MIG slice, region and reservation type. Lago’s guide recommends per-second billing for inference and per-minute for training with a one-minute minimum, and a clear rule for what happens to a job that a failed node killed. When the dashboard and the invoice come from different systems, they disagree, and every disagreement becomes a support ticket and a small loss of trust. The strongest small providers we’ve seen generate both from the same metering ledger.

The scheduler and storage matter as much as the GPUs. Training customers expect Slurm with containers (the pyxis plugin) and topology-aware placement. Inference and notebooks fit Kubernetes. ClusterMAX found that badly tuned shared storage hits the “lots of small files” problem hard enough to make something as basic as import torch painfully slow.

What to fix first

  • Ship an API and a Terraform provider before you build a prettier console. Your best customers automate everything.
  • Build one metering pipeline that feeds the usage dashboard, the invoice and your own utilization reports.
  • Offer Slurm for training and Kubernetes for inference instead of forcing one model on everyone.
  • Benchmark storage with a real PyTorch startup and a dataset of millions of small files before launch, not after the first complaint.

Students fill a launch, but enterprises pay the bills

The customers who keep a GPU cloud alive sign contracts, and they buy on trust signals that most new providers don’t have yet. ABI Research warns that neoclouds without direct enterprise relationships risk staying anonymous back-end capacity. Vultr makes a related point: small providers that compete head-to-head with hyperscalers on hourly price can’t sustain it (Vultr).

The trust gap shows up in four places:

  • Certifications. ClusterMAX notes a long tail of emerging neoclouds without SOC 2 or ISO 27001. A small company typically spends $20,000–$60,000 on SOC 2 Type I and $30,000–$50,000 on Type II (Bright Defense), plus months of evidence collection.
  • SLAs. Enterprises want uptime and response targets in the contract, with credits attached. A site that says “enterprise-grade” without a published SLA gets screened out early.
  • Hardware licensing. Consumer GeForce cards are cheap, but NVIDIA’s driver licence has said since at least 2018 that “the software is not licensed for datacenter deployment”, with an exception only for blockchain (DCD). A procurement team or an NVIDIA partner review will ask about it.
  • Support. A customer whose run dies at 2am wants an engineer who can fix it, quickly and in their own timezone. Local, same-timezone support is one of the few advantages a small provider has over a hyperscaler, so make it visible.

Sovereignty is the best wedge a regional GPU cloud has. Banks, telcos, government agencies and companies training on local-language data often can’t send that data abroad. That market is smaller than “every AI developer”, but it values the two things a local provider offers most easily, in-country data and same-timezone engineers.

What to fix first

  • Publish an SLA, even a modest one you know you can meet. A written 99.5% beats an implied 99.99%.
  • Start SOC 2 readiness early, because Type II needs months of evidence.
  • Run enterprise and regulated workloads on datacenter-class GPUs, and keep consumer cards for clearly labelled dev or education tiers.
  • Pick two or three verticals where data residency matters and build references there before chasing volume.

A one-page checklist for new GPU clouds

If you fix one thing in each area this quarter, make it these:

AreaFirst fixHow you know it worked
UtilizationWeekly utilization report per GPU: reserved, on-demand, idleYou can state your utilization within five minutes of being asked
ReliabilityBurn-in plus automated drain-and-replaceA failed GPU is out of rotation before the customer opens a ticket
Multi-tenancyVM or microVM boundary per tenant; MIG for shared cardsA container escape would reach one tenant’s VM, not the host
PlatformAPI, Terraform provider, one metering ledgerA customer gets GPUs with no human involved, and the invoice matches the dashboard
TrustPublished SLA and a started SOC 2 programmeYou pass a bank’s first security questionnaire without rewriting your answers

None of these needs more GPUs. They need the operational layer that turns a room of hardware into a cloud, and that layer is where most of the work in this business sits.

That operational layer is what we build at Racktrail. We design, deploy and operate open infrastructure (OpenStack, Kubernetes, Ceph, upstream and unmodified) for cloud providers, either fully managed or alongside your own team. Your customers stay yours, the configuration is documented and handed to you, and you can run it without us later.

Frequently asked questions

What utilization does a GPU cloud need to break even?

Around 70% for a typical H100 cluster. Moduledge’s model of a 1,024-GPU cluster loses about $330,000 a month at 55% utilization and earns about $340,000 at 85%. The exact figure depends on your hardware cost, power price and financing.

Can a GPU cloud run on consumer RTX cards?

Technically yes, but NVIDIA’s GeForce driver licence excludes datacenter deployment, and GeForce cards don’t support MIG partitioning. Consumer cards suit clearly labelled education or dev tiers. Enterprise and regulated workloads belong on datacenter GPUs.

Is MIG enough to isolate tenants on a shared GPU?

MIG gives each partition its own hardware path through GPU memory, which is strong isolation on the GPU itself. Tenants still share the host, so pair MIG with a VM or microVM boundary per tenant and keep the NVIDIA container stack patched.

How often do GPUs fail in production?

Meta’s Llama 3 run saw 419 unexpected interruptions in 54 days across 16,384 H100s, about one every three hours, and GPUs or their HBM3 memory caused just over half of them. Scaled to one 72-GPU rack, that is roughly one interruption a month.


Sources

Your VMware renewal is a decision, not an invoice.

Tell us what's running and when your renewal lands. An engineer will map the estate and give you a costed plan to set against your renewal quote.

GPUs where your data is allowed to live.

Tell us the workload, the GPUs you need and the region it has to stay in. An engineer will tell you honestly what to build, rent or skip.

Tell us what needs to run.

Send us the workload, the constraints, and the region it has to live in. An engineer — not a sales rep — will tell you honestly whether we're the right fit.