Buying GPUs is the easy part of starting a GPU cloud. The hard part is running them as a service that a bank, a telco or a research lab will trust with production work, and that is where most new providers stall.
A GPU cloud (people now say “neocloud”) rents GPU compute by the hour, the token or the month, usually to AI teams that can’t buy or run their own hardware. In Nepal alone, at least three have launched or opened early access in 2026, ranging from startups renting consumer RTX cards to H200 racks in Kathmandu. ABI Research expects North America’s share of neocloud revenue to fall from 88% in 2026 to 72% by 2030 as sovereign GPU clouds spread to other regions.
At Racktrail we design and operate open infrastructure for cloud providers. Below are the five areas where we see new GPU clouds struggle most, with public data on each and what I’d fix first. None of it needs a hyperscaler’s budget.
Key takeaways
- A GPU cloud breaks even at roughly 70% utilization, and H100 rental prices have swung from $7–10 an hour to under $2 and back up since 2023.
- GPUs fail often: Meta’s 16,384-GPU Llama 3 cluster hit an interruption about every three hours, so health checks and automatic draining matter from day one.
- Containers alone don’t isolate tenants. NVIDIAScape (CVE-2025-23266) showed a three-line container escape affecting 37% of cloud environments.
- Customers expect an API and an invoice that matches their dashboard, and enterprises expect an SLA, SOC 2 and datacenter-class hardware.
Utilization decides whether the business works
A GPU cloud makes money only while its GPUs are rented, and the breakeven sits around 70% utilization. Moduledge’s model of a 1,024-GPU H100 cluster loses about $330,000 a month at 55% utilization and earns about $340,000 at 85%. Bare-metal gross margins of 55–65% before depreciation leave very little room between those two numbers.
The price you can charge also moves under you. H100s rented for $7–10 an hour in 2023 and fell to $2–4 by late 2025, according to Silicon Data. Then demand turned, and the rate rose about 40%, from $1.70 in October 2025 to $2.35 in March 2026 (Crypto Briefing). A business plan built on one price point will be wrong within a year, in one direction or the other.
Small markets make this harder. A launch fills up with students and researchers who rent a card for a few hours, and per-second billing (which customers love) leaves gaps between jobs that nobody pays for. Depreciation keeps running through those gaps. Most operators assume a 5–6 year useful life, and some analysts argue the real economic life is closer to 2–3 years.
What to fix first
- Track utilization per GPU every week, split into reserved, on-demand and idle. Most new providers can tell you their revenue but not this number.
- Sell committed capacity (monthly or annual reservations) alongside on-demand. Reserved contracts pay for the hardware, and on-demand fills what’s left.
- Slice cards for inference and notebooks. MIG partitions let one H100 or H200 serve up to seven small jobs, which turns an idle full card into paid partial ones.
- Model your prices at a 30% lower rate than today. If the business still works, a price drop becomes an inconvenience instead of a crisis.
GPUs fail more often than new providers plan for
GPUs are the least reliable part of the server, and a cloud has to catch their failures before customers do. When Meta trained Llama 3 on 16,384 H100s, the cluster hit 419 unexpected interruptions in 54 days, about one every three hours (Data Center Dynamics). Faulty GPUs caused 148 of them and HBM3 memory another 72. CPUs caused two.
Apply the same per-GPU rate to a single 72-GPU rack and you get roughly one interruption a month. That sounds manageable until it lands on a customer’s three-day fine-tuning run at 2am, with nobody watching. Meta kept effective training time above 90% because automation handled almost every failure. Engineers stepped in manually only three times.
SemiAnalysis rates GPU clouds in its ClusterMAX system, which covers about 90% of the rental market. One of its clearest dividing lines is health checking: smaller providers typically lack automated weekly active checks such as NCCL tests and DCGM diagnostics.
In emerging markets, power adds a second failure source. At a Kathmandu GPU cloud launch in July 2026, the operator said plainly that uptime would be “probably not 99.99”, and pointed to winter power constraints that hit every facility in the country (TechSansar).
What to fix first
- Burn in every node before a customer touches it, with DCGM diagnostics and NCCL all-reduce tests under sustained load. Early failures show up in the first days.
- Run passive checks all the time (XID errors, ECC counts, thermals, link flaps) and active checks on a schedule.
- Drain a failing node automatically and move the tenant to a spare. Keep a small pool of hot spares for exactly this.
- Tell customers how to checkpoint, and give them a status page. A failure they were warned about costs you far less trust than a silent one.
Sharing GPUs safely is harder than sharing CPUs
Sharing GPUs between tenants is how small clouds push utilization up, and it is also where isolation breaks. GPUs were built for one owner. Their memory isolation is weaker than a CPU’s, and a container shares the host kernel with every other container on the node.
The risk is concrete. In July 2025, Wiz disclosed NVIDIAScape (CVE-2025-23266, CVSS 9.0), a container escape in the NVIDIA Container Toolkit that needed a three-line Dockerfile and reached root on the host. Wiz estimated it affected 37% of cloud environments (The Hacker News). A provider that isolates tenants with containers alone was one malicious customer away from exposing every other customer on that node. ClusterMAX lists container-only isolation as a basic weakness for the same reason.
The sharing method you pick sets how much isolation you get:
| Method | Isolation between tenants | Fits | Watch out for |
|---|---|---|---|
| Time-slicing | None for memory or faults | One team’s dev and test | Never between paying customers |
| MIG partitions | Hardware paths through memory, up to 7 per GPU | Inference, notebooks, small jobs | Datacenter GPUs only (A100, H100, H200, Blackwell); no MIG on GeForce cards |
| vGPU | Per-VM, with IOMMU protection | Enterprise VMs, mixed workloads | NVIDIA licensing cost |
| Full card or passthrough | Single tenant | Training, regulated data | Lowest utilization if you can’t sell whole cards |
Source: Introl, NVIDIA documentation. For how these patterns run on an open control plane, see our guide to running GPUs, FPGAs and DPUs on OpenStack.
What to fix first
- Put a VM or microVM boundary (for example Kata Containers) around every untrusted tenant. Kubernetes can run on top, but it shouldn’t be the only wall.
- Patch the NVIDIA stack on a schedule. The NVIDIAScape fix shipped in Container Toolkit 1.17.8 and GPU Operator 25.3.1.
- Wipe GPU memory and local NVMe between tenants, and default-deny network traffic between them.
- Keep tamper-evident audit logs of who ran what, where. Enterprise buyers will ask for them in the first security review.
A cloud needs a control plane, and most new ones launch without one
Many new GPU clouds launch as a contact form and a chat group: a customer asks for GPUs, an engineer provisions them by hand, and someone sends an invoice at the end of the month. That works for ten customers. At fifty it eats the engineering team, and the customers who matter most (AI teams used to AWS or RunPod) leave before they get that far. They expect an API, a Terraform provider and a console, and Hosted.ai’s guide for telcos flags ticket-based provisioning as a gap to close.
Metering is the part that quietly breaks. GPU billing has to account for GPU model, memory, MIG slice, region and reservation type. Lago’s guide recommends per-second billing for inference and per-minute for training with a one-minute minimum, and a clear rule for what happens to a job that a failed node killed. When the dashboard and the invoice come from different systems, they disagree, and every disagreement becomes a support ticket and a small loss of trust. The strongest small providers we’ve seen generate both from the same metering ledger.
The scheduler and storage matter as much as the GPUs. Training customers expect Slurm with containers (the pyxis plugin) and topology-aware placement. Inference and notebooks fit Kubernetes. ClusterMAX found that badly tuned shared storage hits the “lots of small files” problem hard enough to make something as basic as import torch painfully slow.
What to fix first
- Ship an API and a Terraform provider before you build a prettier console. Your best customers automate everything.
- Build one metering pipeline that feeds the usage dashboard, the invoice and your own utilization reports.
- Offer Slurm for training and Kubernetes for inference instead of forcing one model on everyone.
- Benchmark storage with a real PyTorch startup and a dataset of millions of small files before launch, not after the first complaint.
Students fill a launch, but enterprises pay the bills
The customers who keep a GPU cloud alive sign contracts, and they buy on trust signals that most new providers don’t have yet. ABI Research warns that neoclouds without direct enterprise relationships risk staying anonymous back-end capacity. Vultr makes a related point: small providers that compete head-to-head with hyperscalers on hourly price can’t sustain it (Vultr).
The trust gap shows up in four places:
- Certifications. ClusterMAX notes a long tail of emerging neoclouds without SOC 2 or ISO 27001. A small company typically spends $20,000–$60,000 on SOC 2 Type I and $30,000–$50,000 on Type II (Bright Defense), plus months of evidence collection.
- SLAs. Enterprises want uptime and response targets in the contract, with credits attached. A site that says “enterprise-grade” without a published SLA gets screened out early.
- Hardware licensing. Consumer GeForce cards are cheap, but NVIDIA’s driver licence has said since at least 2018 that “the software is not licensed for datacenter deployment”, with an exception only for blockchain (DCD). A procurement team or an NVIDIA partner review will ask about it.
- Support. A customer whose run dies at 2am wants an engineer who can fix it, quickly and in their own timezone. Local, same-timezone support is one of the few advantages a small provider has over a hyperscaler, so make it visible.
Sovereignty is the best wedge a regional GPU cloud has. Banks, telcos, government agencies and companies training on local-language data often can’t send that data abroad. That market is smaller than “every AI developer”, but it values the two things a local provider offers most easily, in-country data and same-timezone engineers.
What to fix first
- Publish an SLA, even a modest one you know you can meet. A written 99.5% beats an implied 99.99%.
- Start SOC 2 readiness early, because Type II needs months of evidence.
- Run enterprise and regulated workloads on datacenter-class GPUs, and keep consumer cards for clearly labelled dev or education tiers.
- Pick two or three verticals where data residency matters and build references there before chasing volume.
A one-page checklist for new GPU clouds
If you fix one thing in each area this quarter, make it these:
| Area | First fix | How you know it worked |
|---|---|---|
| Utilization | Weekly utilization report per GPU: reserved, on-demand, idle | You can state your utilization within five minutes of being asked |
| Reliability | Burn-in plus automated drain-and-replace | A failed GPU is out of rotation before the customer opens a ticket |
| Multi-tenancy | VM or microVM boundary per tenant; MIG for shared cards | A container escape would reach one tenant’s VM, not the host |
| Platform | API, Terraform provider, one metering ledger | A customer gets GPUs with no human involved, and the invoice matches the dashboard |
| Trust | Published SLA and a started SOC 2 programme | You pass a bank’s first security questionnaire without rewriting your answers |
None of these needs more GPUs. They need the operational layer that turns a room of hardware into a cloud, and that layer is where most of the work in this business sits.
That operational layer is what we build at Racktrail. We design, deploy and operate open infrastructure (OpenStack, Kubernetes, Ceph, upstream and unmodified) for cloud providers, either fully managed or alongside your own team. Your customers stay yours, the configuration is documented and handed to you, and you can run it without us later.
Frequently asked questions
What utilization does a GPU cloud need to break even?
Around 70% for a typical H100 cluster. Moduledge’s model of a 1,024-GPU cluster loses about $330,000 a month at 55% utilization and earns about $340,000 at 85%. The exact figure depends on your hardware cost, power price and financing.
Can a GPU cloud run on consumer RTX cards?
Technically yes, but NVIDIA’s GeForce driver licence excludes datacenter deployment, and GeForce cards don’t support MIG partitioning. Consumer cards suit clearly labelled education or dev tiers. Enterprise and regulated workloads belong on datacenter GPUs.
Is MIG enough to isolate tenants on a shared GPU?
MIG gives each partition its own hardware path through GPU memory, which is strong isolation on the GPU itself. Tenants still share the host, so pair MIG with a VM or microVM boundary per tenant and keep the NVIDIA container stack patched.
How often do GPUs fail in production?
Meta’s Llama 3 run saw 419 unexpected interruptions in 54 days across 16,384 H100s, about one every three hours, and GPUs or their HBM3 memory caused just over half of them. Scaled to one 72-GPU rack, that is roughly one interruption a month.
Sources
- Moduledge, Neocloud: how the GPU cloud business works
- Silicon Data, H100 rental price over time
- Crypto Briefing, Nvidia H100 GPU rental costs surge (August 2026)
- Data Center Dynamics, Meta report details hundreds of GPU and HBM3 related interruptions to Llama 3 training run
- SemiAnalysis, The GPU cloud ClusterMAX rating system
- TechSansar, YetiCloud.AI launches as Nepal’s first GPU cloud (July 2026)
- The Hacker News, Critical NVIDIA Container Toolkit flaw (July 2025)
- Introl, Multi-tenant GPU security: isolation strategies
- Hosted.ai, How CSPs and telcos launch a GPU cloud
- Lago, GPU compute billing
- ABI Research, Neocloud market trends
- Vultr, Will your GPU provider survive the great neocloud consolidation of 2026?
- Bright Defense, SOC 2 certification cost
- Data Center Dynamics, Nvidia updates GeForce EULA to prohibit data center use