Accelerators are the most expensive thing in a modern rack, and they are also the easiest to leave idle. A simulation team books GPUs for a CFD campaign that runs for three weeks and then goes quiet. A design group buys workstations with professional graphics cards that sit idle overnight. A data-science team wants a few GPUs for inference and gets told to wait for the next hardware budget.
The hardware is rarely the problem. The problem is the layer above it: who decides which workload gets which device, how a GPU is split or handed over whole, how tenants are kept apart, and how usage is counted and charged back. That layer is often either a proprietary virtualisation suite or a set of scripts that one engineer understands.
OpenStack, the open-source cloud platform maintained under the OpenInfra Foundation, has become a serious answer to that question. It schedules GPUs, FPGAs, SmartNICs and NVMe devices as first-class cloud resources, and it runs in production at supercomputing centres, national research clouds and commercial GPU clouds. This article covers how it does that, who already relies on it, and what you should know before you try.
Key takeaways
- OpenStack gives you four ways to hand a GPU to a workload: whole-GPU passthrough, time-sliced vGPU, MIG-backed vGPU and bare metal through Ironic.
- The same control plane also schedules FPGAs, SmartNICs and DPUs, fast NVMe and confidential VMs.
- It already runs accelerators in production at Cambridge’s Dawn, Jetstream2, CERN and OVHcloud.
- The hard parts are vGPU licensing, static partitioning, upgrades and finding experienced operators.
Four ways to hand a GPU to a workload
Every accelerator strategy starts with one decision: how much of the physical device a workload gets, and how directly. OpenStack supports four patterns, and one cloud can offer several of them side by side.
1. Whole-GPU passthrough. The hypervisor hands an entire physical GPU to one virtual machine using PCI passthrough (VFIO). The VM sees the real device and runs the vendor’s standard driver, with near-native performance. The cost is density: one GPU, one VM. OpenStack’s compute service (Nova) matches these devices to flavors in its scheduler, optionally tracking them in the Placement service, so flavors can ask for “one A100” or “two L40S” and land on a host that has them free.
2. Time-sliced vGPU. A licensed vGPU driver on the host splits one GPU into several virtual GPUs. Each gets a dedicated slice of memory and takes turns on the compute engines. This is the natural fit for virtual workstations running CAD and visualisation, where many users need a real GPU but rarely all at once. Nova exposes each vGPU as a mediated device (or, with newer NVIDIA drivers, as an SR-IOV virtual function) and schedules it like any other resource.
3. MIG-backed vGPU. On NVIDIA GPUs that support Multi-Instance GPU (MIG), such as the A100, A30, H100, H200 and B200, the card is partitioned in hardware, so each slice gets its own compute engines as well as its own memory. An A100 can be cut into as many as seven instances. Performance is more predictable than time-slicing, and noisy neighbours matter less.
4. Bare metal through Ironic. OpenStack’s bare-metal service provisions an entire physical server, with all its GPUs and interconnect, as if it were a VM: same API, same images, same networks. Nothing sits between the workload and the silicon. That matters for multi-GPU jobs that depend on NVLink or InfiniBand bandwidth, for large solver runs, and for anything that needs to read real hardware identity.
| Passthrough | Time-sliced vGPU | MIG-backed vGPU | Bare metal (Ironic) | |
|---|---|---|---|---|
| Unit given to a workload | Whole GPU | Fraction (shared engines) | Fraction (dedicated engines) | Whole server |
| Performance | Near-native | Varies with neighbours | Predictable per slice | Native |
| Multi-GPU / NVLink | Possible, host-dependent | Limited | No | Full |
| Density | Low | High | Medium | Lowest |
| Guest licensing (NVIDIA) | Standard driver | vGPU licence | vGPU licence | Standard driver |
| Live migration | Limited, since 2025.1 (supported devices only) | Yes, since 2024.1 | Yes, since 2024.1 | No |
| Typical engineering fit | FEA/CFD solvers, rendering | CAD and visualisation desktops | Inference, teaching, light simulation | Large solver runs, distributed training |
The OpenStack accelerator toolkit
No single OpenStack project “does GPUs.” The capability comes from several services working together, which is also why it is flexible.
- Nova and Placement are the core. Placement keeps an inventory of every schedulable resource (CPUs, memory, PCI devices, vGPUs), each with a resource class and descriptive traits. Operators then define flavors that request, say, one vGPU of a given profile. Since the 2023.1 release, generic PCI devices can be scheduled through Placement as well, as an opt-in feature.
- Ironic enrols physical servers through standard management protocols such as IPMI and Redfish and provisions them on demand. It plugs into the same identity, image and networking services as virtual machines.
- Cyborg is OpenStack’s accelerator-management project. It discovers devices such as FPGAs, GPUs, SmartNICs, SSDs and AI chips on each host and reports them to the scheduler. It also handles device programming, such as loading an FPGA bitstream, and attaching the result to an instance.
- Neutron provides tenant networking, including SR-IOV for high-throughput NICs and support for DPUs that run the virtual switch on the card instead of the host.
- Storage and data services round it out: Ceph behind the block and image services, Manila for shared file systems, and Swift or S3-compatible object storage for datasets and results.
- Kubernetes on top. Sites such as Jetstream2 and CERN offer Kubernetes clusters on their OpenStack clouds, giving container users GPUs without giving up the cloud’s isolation and quota model.
Beyond GPUs
GPUs get the headlines, but engineering and research workloads use a wider set of hardware, and the same control plane covers it.
FPGAs. Cyborg treats an FPGA, a region of one, or a loaded function as an allocatable accelerator. This suits signal processing, low-latency inference and hardware prototyping. The Chameleon research testbed, for example, offers FPGA nodes alongside GPU nodes through the same OpenStack interface.
SmartNICs and DPUs. Neutron can bind ports to off-path DPUs, where the virtual switch and network agent run on the card’s own processor. Networking work then stops competing with simulation jobs for host CPU cores.
Fast local NVMe. NVMe devices can be scheduled through Placement like GPUs. That is useful for solvers and post-processing jobs that are limited by scratch-disk speed.
Confidential computing. For manufacturers protecting design IP, or suppliers running jobs on a partner’s hardware, confidential VMs encrypt guest memory so that even the host operator cannot read it. Nova already supports AMD SEV and SEV-ES. The 2026.2 “Hibiscus” release, scheduled for 30 September 2026, extends this to AMD SEV-SNP and Intel TDX. Nova enables attestation but does not perform it, so a verification service is still needed on top.
Who already runs GPUs on OpenStack
This is not a lab curiosity. A sample of public deployments:
| Deployment | Accelerator pattern | What it shows |
|---|---|---|
| Dawn, University of Cambridge (UK) | 1,024 Intel Data Center GPU Max accelerators on 256 Dell servers, run on StackHPC’s Scientific OpenStack | Supercomputing delivered as a cloud, for both AI and simulation |
| Jetstream2, Indiana University / NSF (US) | A100s split by vGPU from a quarter-GPU up, full-GPU passthrough, plus H100 and L40S nodes | Fractional and whole GPUs in one national research cloud |
| CERN (Switzerland) | PCI passthrough and vGPU in its private OpenStack cloud | GPUs offered as a regular resource in a large research organisation’s private cloud |
| Chameleon Cloud (US) | Bare metal via Ironic, with GPU and FPGA nodes | Hardware-level reconfigurability for systems research |
| OVHcloud (Europe) | Public GPU instances (H100, L40S, L4, A10) on KVM in OpenStack regions | A commercial public GPU cloud built on OpenStack |
| G-Research (UK) | Whole-GPU passthrough, with GPU, NVMe and general-purpose VMs co-scheduled through Placement | A private-sector operator’s reasoned trade-off |
| Quansight open-gpu-server | A single GPU server running Kolla-Ansible OpenStack with passthrough flavors | The approach works at one-server scale too |
Spotlight: Dawn. Dawn is the clearest evidence that OpenStack belongs in accelerated computing. Built by Dell, Intel and the University of Cambridge with UK firm StackHPC, it combines 512 Xeon processors and 1,024 Intel GPUs across 256 liquid-cooled nodes. Researchers reach it through a single cloud control plane, which also opens up the rest of the university’s research computing estate. It was built for both AI and simulation, including a digital twin of a fusion power plant and climate modelling. It is not alone: StackHPC reports that OpenStack underpins several TOP500 systems, five of them its own clients.
Spotlight: Jetstream2. Jetstream2 shows how to serve very different GPU users from one pool. Its 90 A100 nodes are carved with NVIDIA vGPU, so a student can take a quarter of a GPU while a research group takes a whole one. Full-GPU flavors use passthrough with a standard driver, partial flavors use the vGPU driver, and newer H100 and L40S nodes sit alongside. Its public documentation is also candid about limits. Suspending GPU instances is unsafe on its libvirt version, and some CUDA features such as unified memory work only on full-GPU flavors. That candour is part of what makes it a good model.
A reference architecture for a simulation and AI cloud
For an engineering organisation, the four patterns map naturally onto a single cloud with three accelerator pools behind one API.
- A workstation pool of vGPU-enabled hosts serves CAD, pre-processing and visualisation desktops. Interactive users share cards, and each desktop keeps a real GPU.
- A solver pool of passthrough hosts runs FEA, CFD and rendering jobs that need a full GPU for hours at a time.
- A bare-metal pool managed by Ironic serves large multi-GPU runs and distributed training over InfiniBand or RoCE, with no hypervisor in the path.
All three share one identity service, one set of projects and quotas, one image catalogue and Ceph-backed storage. A model can move from a designer’s desktop to a solver run to a training job without being copied between systems. Kubernetes clusters can be created on top for teams that prefer containers. The telemetry and rating services (Ceilometer and CloudKitty) turn usage into chargeback reports per team or per project, so each team can see what its GPUs actually cost.
Hard truths about GPUs on OpenStack
An honest guide has to cover what is still difficult.
- Partitioning is static. vGPU and MIG profiles are set per physical card by the operator. Users cannot resize a slice on demand, so profiles need planning against real demand.
- Licensing is separate. NVIDIA vGPU requires guest licences and a licence server. Passthrough and bare metal avoid that cost and the extra moving part.
- Lifecycle operations need recent components. vGPU live migration only arrived in 2024.1 and needs recent libvirt and QEMU. The NVIDIA driver does not support more than one vGPU per instance on a single GPU.
- Placement details matter. GPUs are bound to NUMA nodes. By default, when Nova knows the topology, it places a VM on the same NUMA node as its GPU. A flavor that doesn’t account for this can fail to schedule or leave performance on the table.
- Drivers, kernels and upgrades must move together. Host GPU drivers, the kernel, libvirt and the OpenStack release form one tested set. Upgrading any piece without rehearsal is where production GPU clouds get hurt.
- The scarce resource is people. Experienced OpenStack operators are scarce. Plan for the operations team, not only the build.
None of these are reasons to avoid OpenStack. They are the difference between a proof of concept and a service people depend on.
OpenStack, Kubernetes-only or VMware?
Kubernetes-only is excellent when every workload is containerised and one team owns the cluster. GPU sharing through MIG and time-slicing is mature there, and there is no guest licensing.
VMware remains familiar for Windows workstation estates, but its post-acquisition licensing changes have made the cost of staying a live question.
OpenStack earns its place when you need several of these at once: VMs and bare metal, Windows desktops and Linux solvers, multiple teams with hard isolation and quotas, and chargeback. In practice the choice is often not either-or, because Kubernetes runs comfortably on top of OpenStack.
The bottom line
Accelerated computing is too expensive to run on idle capacity and guesswork. The organisations above show what an open control plane can do. It shares GPUs across very different users, keeps tenants isolated, and turns a pile of costly hardware into a service with a single front door. The software is free and the patterns are proven. What decides success is careful design of the pools and disciplined operation afterwards.