Dedicated multi-node GPU clusters with high-speed interconnect, reserved for you alone. Sized to your job, priced for months rather than minutes.
An on-demand instance is one box you rent by the second. A cluster is a block of interconnected nodes held for you across the whole term — so throughput stays flat and nothing gets reclaimed mid-run.
From eight GPUs to several hundred, on one high-speed fabric. Scale the block up between phases instead of re-architecting.
Single-tenant hardware, your own network segment, direct SSH. Nothing shared, nothing preempted.
Capacity already exists on the network. Most clusters go from signed spec to first job inside a day, not a procurement cycle.
A committed term buys a rate well under the spot market. One line item, no per-hour surprises.
You're training or serving continuously, and several teams are queueing for the same cards. Reserved capacity costs less than paying spot around the clock.
Distributed training, simulation or big-data work where nodes talk to each other constantly. Interconnect topology decides your throughput, not raw FLOPS.
A production deadline, a customer-facing endpoint, or a run that loses days if a node disappears. You need the capacity guaranteed, in writing.
Pick the GPU generation, the node count, the CPU-to-GPU ratio, the interconnect, and the storage tier. Right-size between phases rather than paying for headroom you aren't using.
Utilisation and cost tracked per node and per team, in real time. Know which run burned the budget while it is still running, not at the end of the month.
PyTorch, TensorFlow, vLLM, Slurm, Kubernetes — anything that runs in a container runs here. The same CLI and SDK drive the cluster as drive a single instance.
A named contact who knows your topology, plus written SLA terms on availability and response. Not a ticket queue you're shouting into.
Not every workload needs a dedicated cluster. Plenty of teams are better served renting single instances by the second — and we would rather tell you that up front.
Occasional training or inference that a single instance handles. Bootstrapped or side projects with sporadic compute needs.
Short, ad-hoc bursts while you figure out what the model needs. Committing to a term before you know the shape of the job costs more, not less.
If you're happy taking interruptible capacity and restarting from a checkpoint, a dedicated cluster is overkill.
One OpenAI-compatible endpoint for 300+ models, with routing and failover.
API Gateway →Search GPUs, deploy containers and script the whole workflow from your shell.
OpenLink CLI →A typed client for compute and models. Build agents that scale their own GPUs.
Python SDK →Tell us the shape of the job — node count, interconnect, term — and we'll come back with a spec and a number.