We run about twenty small services on one 4-vCPU / 8 GB box. One of them transcribes meeting audio with whisper.cpp, on that same box, with no GPU. This is what we actually measured, and what we got wrong on the first pass.
The constraint is CPU seconds, not API spend
Most hosted transcription bills you per minute of audio because it costs them per minute of inference. We compiled whisper.cpp on the server, so the marginal cash cost of one more transcript is zero. What we spend instead is CPU time — and CPU time is shared with every other service on the machine.
That changes the shape of the problem. There is no bill that grows smoothly with usage. There is a wall, and past that wall everyone's queue gets slower at once.
The numbers
Benchmarked on a 60-second file, 2 threads:
| model | realtime factor |
|---|---|
base | 0.434x |
small | 0.838x |
So small processes roughly 1.19 minutes of audio per wall-clock minute per slot. With two concurrent slots and two threads each, all four cores are busy.
Multiply that out and a month of wall-clock time gives about 87,700 audio-minutes in theory. We do not sell 87,700. We sell 30,000.
Where the other 57,700 went
Two haircuts, and one of them we now think was wrong.
Contention (×0.85). The other services on the box are not idle. They are node processes serving HTTP, running cron, talking to Postgres. Threads do not get clean cores.
Utilisation (×0.40). This was the guess. The reasoning was that usage is not evenly spread — people upload during working hours, not at 4am — so the ceiling that matters is the peak, not the average.
The 0.40 is where we were sloppy. It is not measured. It is a number that felt safe.
Measuring instead of guessing
We now read the actual CPU history from sar rather than trusting the constant. The thing that matters is not average idle but idle during the busy stretches — we take the 10th percentile.
Today's reading on our box: median idle 80.2%, but p10 idle 68.4%. That means even when things are busy, about 2.74 of 4 cores are free. Feeding that into the same arithmetic gives roughly 141,000 audio-minutes per month — about 4.7 times what we currently sell.
We did not raise the limit.
Why not raise it
Because we have zero transcription jobs in production. The realtime factor above is a single benchmark on a single file. The gap between "one 60-second file on an idle box" and "someone uploads a 90-minute board meeting while three other people do the same" is exactly the gap that ruins queues.
So the tooling now does something asymmetric on purpose:
- If measurement says capacity is lower than configured, it lowers it automatically. Overselling is the failure that hurts users.
- If measurement says higher, it refuses to raise, and prints what evidence would justify raising: twenty real transcriptions on the
smallmodel, one month of actual-versus-committed usage, three paying subscribers for a load pattern.
An automation that can only tighten is not much of an automation. But an automation that loosens on thin data is worse than none.
The capacity guard is a hard stop
When committed throughput reaches 80% of capacity, new subscriptions stop being sold. Not a warning, not a slower queue — the plan comes off the shelf.
This was the most argued-about decision in the design. A soft degrade keeps revenue flowing. But a queue that quietly gets slower is a service that is broken in a way nobody can see, and every existing subscriber pays for the new ones. A "sold out" sign is honest. We would rather explain why you cannot buy today than explain why your transcript took nine hours.
What we would tell you to check first
If you are considering the same architecture:
- Measure idle at the percentile that hurts, not the mean. The average is comfortable and useless.
- Write down which of your constants are measured and which are guesses. Ours were mixed together in one number, and it took a rewrite to separate them.
- Decide the failure direction before you need it. Ours is: refuse to sell rather than degrade. That decision is much harder to make honestly once there is revenue attached to it.
The service is at voice.star365.site if you want to see what the constraints produced. Free tier is 30 minutes a month; paid plans start at 9 USDT. Payment is USDT on BSC only — there is no card processing, which we know filters out most people.