Working session around a table with notebooks and a laptop in a bright meeting room

A sitting with invoices, not a product login. Jalan Thambi Dollah rooms, or video.

Kuala Lumpur · AI application bills

We sit with the invoices your AI feature leaves behind.

Compute Pulse Grid is a small practice that reads compute cost and usage for AI applications already in production. GPU-hour lines, token charges, reserved blocks that outlived a campaign — we mark them, then write a findings memo you can take to finance.

Request a cost review

01

Invoices

The PDF summary is a starting page. We ask for the usage extract behind it so reserved hours, failed jobs, and second-region replicas are visible as separate stories.

02

Usage windows

A noisy week after a model swap or a Hari Raya campaign is not the same as a quiet run rate. We name the dates before we name a number.

03

Allocation

When three products share an endpoint, the bill is one line. We write a method the teams can recognise, including what the method cannot see.

Flagship sitting

A compute cost review for one named application

Over three weeks we read ninety days of invoices and a sample of usage logs for a single AI application and one billing account. Two working sessions correct what we have misread. You receive a findings memo, an idle-capacity note, and a first-pass allocation table.

We do not rewrite your routing, negotiate vendor credits, or leave you a dashboard. The work ends on paper, with a follow-up call.

From RM 12,400. Extra accounts and extra applications are quoted after intake.

Read the review in full

Person comparing printed billing pages with notes on a laptop

Also on the books

Other readings we take

A full review is not always the right sitting. Some teams need one noisy month explained. Some need a shared endpoint split across product owners. Some already have a memo and want the same shape of note each period.

Close view of a calculator, printed figures, and a pen on a desk

Short reading

Usage window briefing

A focused sitting with one billing period — a campaign week, a model swap, or a month that jumped — so you can explain that window to finance.

Three colleagues talking through papers around a small meeting table

Shared endpoints

Cost allocation mapping

A mapping of one shared model endpoint or GPU pool onto the product teams that actually call it, so the bill is no longer a single unexplained number.

Two people reviewing a document together at a bright office table

Retainer

Recurring spend briefing

A monthly or quarterly written briefing on the same application, so finance sees the same shape of note each period instead of a scramble after every invoice.

From a sitting

A replica that outlived a Johor pilot

Farah sat with three months of GPU-hour lines and found a replica we had left up after a pilot in Johor. The memo gave finance a date and an instance name, which ended a month of vague arguing. I still wish the kickoff had spelled out how long our security team would take to export the usage file — that wait ate a week we had not planned.

Ahmad Rahman, Product lead, document answering feature · Compute cost review

More client stories

How a review starts

Send the application name and three months of invoices

We reply within two working days with whether the pack is enough to start. Incomplete logs are common; we will say so rather than fill gaps. Sessions are in our rooms at Wisma Ho Po or by video.

See the engagement steps

Bring a named billing owner. We do not ask for production keys at intake. If a later export is needed, the grant is written first.

Address for sittings: 65-67 5Th Floor Wisma Ho Po Ckt Thambi Dollah,Kuala Lumpur,Wilayah Persekutuan,55100,Malaysia

Notes

From recent readings

Short pieces on GPU-hour lines, token gaps, leftover campaign capacity, and the questions finance asks after a feature ships.

Handwritten notes beside a laptop during a working session

2 April 2026

Token counts and the bill are not the same story

A product team can watch tokens rise slowly while the invoice jumps. The gap is usually context length, retries, or a second model that never appeared on the dashboard the team actually looks at.