Accelerated compute, procured and governed like infrastructure — not bought like hardware.
GPU capacity has become a material line on the technology budget, and it is being bought the way cloud was bought in 2015 — urgently, in isolation, and without the commercial discipline applied to everything else. We work both sides of that: standing supply relationships for NVIDIA H200, B200 and B300 capacity, and the workload analysis that decides how much of it you should be holding. Capacity is placed against a modelled utilisation curve, under the same FinOps controls we put around cloud spend.
The problem
Most GPU spend is committed before anyone models the workload.
The market conditions make this easy to get wrong. H200 is now effectively a commodity — tracked across more than twenty-five providers, with genuine price competition. Blackwell Ultra is the opposite: scarce, and moving almost entirely through reserved contracts rather than public on-demand rates. Buying both the same way is how organisations end up overpaying for one and unable to get the other.
Committed too early
A three-year reservation signed before the training run is characterised, then paid for at 30% utilisation.
Wrong silicon
Frontier-tier GPUs bought for inference that a previous generation would serve at a fraction of the cost per token.
No unit economics
Spend reported as a single monthly figure, with no cost per training run, per model or per business owner.
Single-source exposure
One provider, one region, one contract — and no position to negotiate from when the renewal arrives.
What we do
Four things, and we stay accountable for all of them.
Workload characterisation
Before any commitment. We profile the actual jobs — model size, context length, batch behaviour, memory ceiling, interconnect sensitivity — and establish whether you are memory-bound, compute-bound or network-bound. That determines the silicon. Everything else follows from it.
Supply and placement
We hold relationships across hyperscalers and specialist GPU providers, so a requirement can be placed against capacity already available to us rather than worked up from a public price list. Reserved terms typically run from one to thirty-six months, and the position against on-demand is negotiated, not published.
Landing zone and operations
Capacity is not a platform. We build the surrounding environment — identity, networking, storage throughput, scheduling and queueing, observability — so the cluster is usable by your team on day one and does not sit idle while someone works out how to submit a job.
FinOps for accelerated compute
Utilisation tracking against the commitment, cost per training run and per million tokens, chargeback to the team that caused the spend, and a standing review before every renewal. This is the part most providers have no incentive to give you.
The fleet
Matching silicon to the workload, not to the datasheet.
A mixed fleet almost always beats a single-SKU strategy: current-generation capacity for production serving, frontier-tier reserved for the training that genuinely needs it. Specifications below are NVIDIA's published figures.
| GPU | Memory | Bandwidth | Best fit | Availability |
|---|---|---|---|---|
| NVIDIA H200 | 141 GB HBM3e | 4.8 TB/s | Production inference, long-context serving, fine-tuning. Roughly 1.4× H100 on training and up to 1.8× on inference for memory-bound work. | Broad supply |
| NVIDIA B200 | 192 GB HBM3e | ~8 TB/s | Blackwell-generation training and high-throughput inference where FP4/FP8 paths are in use. | Contracted |
| NVIDIA B300 | 288 GB HBM3e | 8 TB/s | Blackwell Ultra. Reasoning-model training and inference where the memory ceiling per GPU is the binding constraint. | Reserved only |
| GB300 NVL72 | ~20 TB per rack | 130 TB/s NVLink | 72 Blackwell Ultra GPUs and 36 Grace CPUs as a single NVLink domain — around 1.1 exaFLOPS FP4. Frontier-scale training. Draws roughly 120 kW per rack, which is a facilities conversation before it is a compute one. | Allocation |
Availability reflects what can realistically be placed through current supply arrangements rather than a published price list. Supply, region and terms are confirmed against the live position at the point of engagement.
Commercial models
The contract shape is where the money is won or lost.
Most organisations need more than one of these at once. The skill is in the proportion — enough committed capacity to earn the discount, enough elastic capacity to absorb the spikes without paying for them all year.
On-demand
Experimentation and burst
- No commitment, highest unit rate
- Right for evaluation, spikes and proof-of-concept work
- Wrong as a steady state — this is where budgets quietly disappear
Reserved capacity
One to thirty-six months
- Materially below on-demand — the gap is negotiable and worth negotiating
- The only realistic route to Blackwell Ultra at present
- Sized against a modelled utilisation curve, not an optimistic one
- Reviewed before renewal, with a documented position
Dedicated cluster
Isolated and configured to you
- Single-tenant, with control over topology and interconnect
- For sustained frontier-scale training and strict data isolation
- Facilities, power and network are part of the decision
Australian context
Where the compute physically sits is a board-level question.
Training data does not stop being regulated because it has been turned into weights. For organisations under the Privacy Act, APRA CPS 234, or government data-classification obligations, the residency of the GPU matters as much as its specification — and the cheapest capacity is frequently in the wrong jurisdiction. We treat residency, sub-processor disclosure and exit terms as selection criteria, not as paperwork to be resolved afterwards.
How an engagement runs
Four phases. You can stop after any of them.
Each phase produces something you own and could hand to another party. We do not hold the artefacts hostage to the next stage.
1. Characterise
Profile the workloads and establish the binding constraint. Output: a sizing model with the assumptions written down and challengeable.
2. Place
Match the requirement against what is currently available to us and normalise the options onto comparable terms. Output: a written position — what can be secured, on what commitment, at what unit cost — that your procurement can defend.
3. Land
Stand up the environment around the capacity — access, networking, storage, scheduling, observability. Output: a cluster your team can actually submit work to.
4. Govern
Utilisation against commitment, unit economics, chargeback, and a renewal position prepared before the renewal date. Output: a standing review your finance team can read.
Questions
What clients ask before committing.
Do you provide the capacity, or advise on it?
Both, and we tell you which one applies before any work starts. We hold supply relationships on one side and run the workload analysis on the other, so in most engagements the capacity can be arranged through us rather than leaving you to approach the market alone. Where contracting directly is the better outcome for you, we will say so and structure it that way. Either way the commercial structure, including how we are paid, is on the table at the outset.
We only need inference. Is Blackwell Ultra worth it?
Usually not. Frontier-tier silicon earns its premium on training and on reasoning workloads where per-GPU memory is the binding constraint. A large share of production inference runs more economically on H200, and the cost per million tokens is the number that should decide it — not the generation on the datasheet.
How long does capacity take to secure?
Current-generation capacity in a common region can move quickly, and we can usually tell you what is available against your requirement in the first conversation. Blackwell Ultra and NVL72-class allocations are a different matter and are planned in quarters, not weeks. You will know which situation you are in before you commit to a timeline.
Can you work with capacity we have already contracted?
Yes, and this is a common starting point. If a commitment is already signed and under-utilised, the work is recovering value from it — improving utilisation, restructuring how it is shared across teams, and building the position for renegotiation at renewal.
Do you handle the facilities side?
For hosted capacity, the provider does. For dedicated or on-premises deployments it becomes central: an NVL72-class rack draws on the order of 120 kW and has cooling requirements that most existing floor space cannot meet. We bring that into the assessment early, because discovering it late is expensive.
Start with the workload, not the purchase order.
A scoping conversation is thirty minutes and does not require you to have a budget approved. If the answer is that you should not be buying GPU capacity yet, we will say so.
Book a scoping call

