GPU instances are now billed with per-second granularity, with a one-minute minimum. For teams running many short jobs, that change removes most of the waste that hourly billing introduced. Hyperparameter sweeps, batch inference and rendering jobs rarely run for a clean hour, and until now they were billed as though they did. The gap is small per job and large per month.

Nothing else about the instance changes. You still choose the GPU class, the image and the volume when you launch, and the instance behaves exactly as it did before. The only difference is on the invoice, where the line for compute now reflects the time the instance was actually running. Only the invoice arithmetic changed.

What this changes in practice

A job that runs for four minutes is billed for four minutes, not sixty. On a sweep of two hundred short runs, that difference is usually larger than the cost of the instance itself. The saving is not a discount and it is not a promotion. It is the removal of a rounding rule that suited billing systems and penalised users.

  • Per-second billing, one-minute minimum
  • Applies to on-demand GPU instances
  • Monthly reservations unchanged
  • No change to how you reserve capacity

How the meter works

The meter starts when the instance reaches a running state and stops when it is terminated, not when shutdown is requested. Time is recorded per instance, so a parallel sweep is billed for each instance separately. The one-minute minimum exists because provisioning and teardown still consume real capacity on the host. Teardown is not instant, and the host holds the allocation until it completes.

For a long-running job, per-second billing changes very little, which is why monthly reservations are untouched. For exploratory work, it changes the arithmetic completely. A notebook session that runs for eleven minutes costs eleven minutes, and a job that finishes early stops the meter immediately. That last part matters more than it sounds, because abandoned experiments are common.

The rounding rule also distorted behaviour in a way that was easy to miss. Teams batched small jobs into longer sessions to avoid paying for a full hour, which made failures more expensive and debugging slower. Billing by the second means the cheapest way to run a job is also the simplest way to run it.

We rebuilt the metering pipeline before switching this on and ran it in shadow against real usage for a period. During that time we compared per-second figures against the hourly figures customers would have paid, and treated every disagreement as a defect rather than a rounding difference. The change applies automatically, with no setting to enable and no new agreement to sign. Shadow mode keeps billing defects off customer invoices.

Checkpoints still persist

Attach a persistent NVMe volume and your checkpoints survive between jobs, so a short run does not mean starting over. A sweep can write state after every configuration, stop, and resume from the best result without retraining from the beginning. Volumes are billed separately from compute and can be detached when you are finished with them.

What stays the same

Monthly reservations are unchanged, and the way you reserve capacity has not moved. Reserved capacity remains the right answer for steady training workloads, and it still carries the same terms. On-demand instances remain the right answer for everything that is short, unpredictable or exploratory, which is exactly the work this change is aimed at. Both appear on the same usage view.

  • Same GPU classes and images
  • Same private networking between nodes
  • Same persistent NVMe volumes
  • Same support during MYT business hours

Where to see it

Usage appears in the portal with per-second detail, so a sweep can be reconciled against the invoice without estimating. Historical invoices are not restated and nothing already billed is recalculated. If you want help sizing a workload against the new model, our team is available during local business hours. Usage detail exports in a format finance teams can read.

This is the first of several billing changes we are making to the GPU platform this year. Each one is aimed at the same problem, which is aligning what you pay with what you consume. We will publish each change here before it takes effect, and we will not restate historical invoices when we do.

You should pay for the compute you used, not for the hour it happened to fall inside.