GPU Scheduling on Kubernetes: Requests, Sharing and Keeping GPUs Busy
GPU nodes are usually the most expensive nodes in a cluster, and Kubernetes treats them very differently from CPU and memory. A GPU is handed out whole, it is not overcommitted, and once a pod holds it, nobody else can use it, even if the pod does nothing. Most GPU waste comes from that model, not from the scheduler being "dumb".
This post covers how GPU scheduling works, where idle GPU time comes from, and what to do about it.
How Kubernetes sees a GPU
Kubernetes does not know about GPUs by itself. A vendor device plugin (NVIDIA, AMD, Intel) runs on each GPU node and advertises an extended resource such as nvidia.com/gpu. The node also needs the vendor driver and container runtime support. For NVIDIA, the GPU Operator installs the driver, container toolkit, device plugin and DCGM exporter as one package.
Extended resources follow stricter rules than CPU:
- You set GPUs in
limits. If you omitrequests, the request defaults to the limit. - If you set both, they must be equal. A different value is rejected by the API server.
- You cannot set a GPU request without a limit.
- Whole units only.
0.5is not valid. - No overcommit, and by default one GPU belongs to one container.
apiVersion: v1
kind: Pod
metadata:
name: train-my-model
spec:
restartPolicy: Never
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
containers:
- name: trainer
image: registry.example.com/ml/trainer:1.4.2
resources:
limits:
nvidia.com/gpu: 1
memory: 32Gi
requests:
cpu: "8"
memory: 32Gi
Keep other pods off GPU nodes
Taint GPU nodes so that only pods that tolerate the taint land there:
kubectl taint nodes gpu-node-1 nvidia.com/gpu=present:NoSchedule
Without a taint, ordinary pods fill the CPU and memory on GPU nodes, and then a GPU job cannot be scheduled even though the GPU is free. The ExtendedResourceToleration admission plugin can add the matching toleration automatically to pods that request nvidia.com/gpu. Some managed providers taint GPU node pools for you; check your provider's docs.
If you have several GPU models, label nodes (or use Node Feature Discovery and the GPU Operator's labels) and select the model with nodeSelector or node affinity. A small inference service does not need the largest card in the cluster.
Where idle GPU time comes from
- Interactive workloads. Notebooks and dev pods hold a GPU all day and use it for minutes.
- Inference services sized for peak traffic, holding a full GPU at low load.
- Data loading bottlenecks. The GPU waits on CPU, disk or network, so it is allocated and mostly idle.
- Fragmentation. Pods requesting one GPU are spread over many nodes, so a job that needs all GPUs of one node cannot start, and the cluster autoscaler adds another node.
- Runaway jobs that never finish.
Measure first
Allocation is not usage. Install DCGM exporter (the GPU Operator does this) and compare what is allocated with what is used:
# average GPU utilization per pod over the last hour
avg by (namespace, pod) (avg_over_time(DCGM_FI_DEV_GPU_UTIL[1h]))
# framebuffer memory used, MiB
max by (namespace, pod) (DCGM_FI_DEV_FB_USED)
The pod and namespace label names depend on your exporter and scrape configuration. DCGM_FI_DEV_GPU_UTIL only shows the share of time a kernel was running, not how much of the GPU it used. For a better picture, look at profiling metrics such as DCGM_FI_PROF_SM_ACTIVE if your GPUs support them.
Pick a threshold that makes sense for your workloads and report pods that stay below it for days. Those are the candidates for sharing or a smaller GPU.
Share GPUs where it is safe
- Time-slicing (NVIDIA device plugin) advertises each physical GPU as several replicas. Pods interleave on the same GPU with no memory or fault isolation. Good for notebooks and light inference you trust.
- MIG (Multi-Instance GPU, on supported NVIDIA data center GPUs) splits a GPU into hardware-isolated instances with their own memory. Good for multiple inference services.
- MPS lets processes share a GPU concurrently with some limits. Check the device plugin docs for its current support level.
Time-slicing config for the NVIDIA device plugin:
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
namespace: gpu-operator
data:
any: |-
version: v1
flags:
migStrategy: none
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4
Point the operator at it through devicePlugin.config in the ClusterPolicy or Helm values. Training jobs that need the full card should not run on time-sliced nodes. Keep separate node pools for shared and dedicated GPUs.
Pack GPU workloads
To avoid fragmentation, have the scheduler prefer nodes that are already busy. The NodeResourcesFit plugin supports the MostAllocated strategy with per-resource weights:
apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
profiles:
- schedulerName: gpu-binpack
pluginConfig:
- name: NodeResourcesFit
args:
scoringStrategy:
type: MostAllocated
resources:
- name: cpu
weight: 1
- name: memory
weight: 1
- name: nvidia.com/gpu
weight: 5
On managed clusters you usually cannot change the default scheduler's config, so run this as a second scheduler and set schedulerName: gpu-binpack on GPU pods. Packing also lets the cluster autoscaler remove empty GPU nodes.
Queue batch jobs
The default scheduler places pods one at a time and has no idea of team quotas or of a job that needs eight pods to start together. For training workloads, add a queueing layer:
- Kueue (a Kubernetes SIG project) holds Jobs suspended until their team's quota has room, and supports borrowing unused quota between teams.
- Volcano is a batch scheduler with gang scheduling, so a distributed job starts all its pods or none.
Set activeDeadlineSeconds on Jobs so a stuck run does not hold GPUs forever.
Dynamic Resource Allocation
DRA is GA since Kubernetes 1.34. Instead of a counted resource, pods reference a ResourceClaim and a DeviceClass, which lets drivers describe devices with attributes (model, memory, partitions) and lets the scheduler select on them. It needs a DRA driver from your GPU vendor. If you are designing a new GPU platform, check whether your vendor's driver is ready before building on the device plugin model.
Checklist
- Device plugin or GPU Operator installed, drivers match the node image.
- GPUs set in
limits; requests omitted or equal. - GPU nodes tainted; pods tolerate and select the GPU model they need.
- DCGM metrics in Prometheus; allocation compared with real usage.
- Time-slicing or MIG for light workloads, dedicated GPUs for training.
- Bin packing for GPU pods, queueing for batch jobs, deadlines on Jobs.
