11 May 2026
How to read a GPU-hour line on a Malaysian cloud invoice
The GPU-hour row is rarely one number. Here is how we unpack it when a Kuala Lumpur team brings last month’s bill to Jalan Thambi Dollah.
Most invoices we see in Kuala Lumpur still print a single GPU-hour total. That total usually mixes on-demand hours, a reserved block that was paid whether or not it ran, and a few hours that belong to a failed job the team already forgot. Before anyone argues about the amount, the line has to be split.
Ask the billing owner for the usage extract that sits behind the invoice, not only the PDF summary. On the extract, sort by instance type and by whether the hour was reserved. Reserved hours that show zero utilisation are not a mystery; they are a decision that was made when traffic was higher, and they stay on the bill until someone cancels the reservation.
Failed training jobs deserve their own column. A job that dies after forty minutes still bills those forty minutes. If the team reruns the job twice, the invoice shows three attempts and one useful result. Finance rarely hears that story unless someone writes it down next to the line.
We also mark hours that ran in a second region “for latency” after a campaign. Those hours often survive the campaign. If nobody owns the replica, it keeps drawing GPU time through the next quarter. The invoice will not label it as leftover campaign capacity; you have to match dates to the marketing calendar.
When the extract and the invoice disagree by more than a small rounding difference, stop and ask the vendor for a usage file covering the same dates. Do not average the two numbers. The findings memo we write always states which file we trusted and why.