← Blog/Billed Before the First Request

Billed Before the First Request

Billed Before the First Request
SEP 29, 2026

Stand up the smallest production-shaped deployment you could defend in a code review. One managed Kubernetes cluster. One load balancer in front of it. One NAT gateway so the nodes can reach the internet. One public IPv4 address. One secret in a managed store. One block storage volume.

Now serve nothing. No traffic, no users, no requests. Read the invoice.

The control plane is ten cents per cluster-hour on two of the three largest providers, whether the cluster runs three nodes or three hundred. That is roughly seventy-three dollars a month. The NAT gateway is four and a half cents an hour (about thirty-two dollars a month), plus four and a half cents for every gigabyte it processes. And it begins billing the moment it is created, with no free allowance. The load balancer costs a little over two cents an hour simply to exist, sixteen to eighteen dollars a month, before the capacity units that measure any actual work. The public IPv4 address is half a cent an hour, three dollars sixty a month, charged whether it is attached to anything or not. The secret is forty cents a month. The volume bills on provisioned capacity rather than used capacity, from creation until deletion; in Frankfurt, just under ten euro cents per gigabyte-month.

Call it a hundred and twenty-five dollars a month. Before compute. Before storage anyone reads. Before the first request.

Then add a staging environment, because you are professional, and a development environment, because you are sane. Every charge above is per resource, and every environment needs its own. The floor does not know how small you are. It triples.

There is a further charge worth sitting with. On one of those providers, when the Kubernetes version you are running leaves its fourteen-month standard-support window, the same control plane goes to sixty cents an hour. Six times the price, around four hundred and thirty-eight dollars a month, for an identical cluster. Nothing about the capacity changed. The charge is for being late. The operator most likely to be late is the one without a platform team to keep them current.

What the floors actually buy

There is a second complaint about deployment that has nothing to do with money, and it cuts the other way. Software written to be handed to other people cannot take the shortcuts a deployment you control can. TLS termination cannot be assumed; it has to be documented and handed off to whatever reverse proxy the operator already runs. A dependency on one database extension means everybody installing the software has to modify a shared database before it will start. Serving readers without JavaScript and serving an interactive client are two implementations of the same feature, and both have to be maintained. The path of least resistance out of all of that is to ship one container with the entire stack inside it and support nothing else.

That path fails for a reason worth stating precisely: there is no single simple way to deploy because there is no single deployment. Every operator arrives with an opinionated setup already running: their own proxy, their own certificate handling, their own access control. An application carrying its whole stack has to fight that setup rather than fit into it. The same objection comes back from the operator's side as a matter of taste: nobody wants software that arrives with the kitchen attached. Fixed, indivisible units are exactly what both sides are complaining about, which is why the arithmetic above deserves stating in its own terms rather than being waved past.

The floors buy interoperability. They buy a control plane that is genuinely running, genuinely replicated, and patched by someone who is not you. One provider is explicit about the trade: the free tier of its managed Kubernetes carries no financially backed availability commitment and is recommended below ten nodes, while the paid tier buys a 99.95 percent commitment on the API server and additional control-plane replicas. The address charge tracks a real scarcity: the entire IPv4 space is allocated, and the provider that introduced the charge cited a rise of more than three hundred percent in acquisition cost over five years. These are cost-recovery charges attached to things that cost money to provide. And the alternative that gets recommended in every one of these threads, one server patched by hand, carries costs that appear on no invoice at all, until the month a stale plugin turns your side project into somebody's botnet.

So the complaint is not that fixed costs exist. Fixed costs are honest. The complaint is about what they are attached to.

Who the floor lands on

Offset printing had this shape for most of a century. The plates cost what they cost before a single sheet runs, so setting up a run of fifty and a run of five thousand cost nearly the same, and the per-copy price of the short run was mostly setup. Nobody chose that out of contempt for small publishers. It was what the machine required before it could produce anything at all. What eventually fixed it was not a discount for small orders. It was a press that needed no plates, and once the setup was gone, a single copy became a sensible thing to buy and print on demand became an industry.

The awkward property of a fixed charge is not its size. It is that its weight runs inversely to what you do with it. A hundred and twenty-five dollars a month is a rounding error against a real production workload and the entire budget of a side project that makes forty euros. Identical charge, opposite meaning. At the top of the range the floor disappears into the variable cost of serving real traffic. At the bottom it is most of the invoice, and what it pays for is not the traffic but the eligibility to receive any.

The ratio also gets worse the more careful you are. A project that right-sizes its compute, caches well and keeps its images small is shrinking the part of the bill that responds to effort and leaving the part that does not, so optimization runs into a floor of its own, below which there is nothing left to win. The reward for being small and efficient is a bill that is almost entirely composed of things you cannot influence.

The obvious defence is that at least the capacity is being used. The published measurements do not support that, though they need reading carefully.

The most-cited dataset comes from a vendor that sells Kubernetes cost optimization, which is worth holding in mind: its 2026 report puts average CPU utilization across the clusters it measures at 8 percent and memory at 20 percent, both down from prior years, with GPU utilization around 5 percent. An observability vendor's container research found that more than 65 percent of Kubernetes workloads use less than half the CPU and memory they request, and that fewer than one percent of organizations run the vertical pod autoscaler that would correct it. A broader industry survey of 753 respondents put estimated cloud waste at 29 percent in 2026, the first increase in five years, though that figure is what practitioners estimate about their own spend rather than anything measured.

One thing none of those datasets shows is that small deployments are the worst offenders. The largest of them excludes clusters below fifty CPUs outright, on the grounds that they are too small to form a reliable sample. Within what it does measure, the trend runs the other way: around 13 percent CPU utilization in the fifty-and-up band, rising to 44 percent in the very largest clusters, which are under one percent of the sample. Bigger teams manage capacity more closely, because they have people whose job that is.

So this is not the argument that small clusters waste more. The measurements that would settle it leave them out. It is an argument about what they are billed, which is a different and simpler claim, and one anybody can check against a public price list this afternoon.

The unit is the problem

Look again at what each floor is charging for, and a pattern shows up that has nothing to do with pricing strategy.

A cluster is the smallest thing that can have a control plane. A load balancer is the smallest thing that can terminate a connection. An address is the smallest thing that can be routed to. A volume is the smallest thing that can hold a filesystem. None of them divides. You cannot buy a tenth of a control plane, so the smallest project buys a whole one, and the bill is not really a price list. It is a description of the architecture's granularity, converted into money.

Other industries have made that transition, and they made it by changing the machine rather than the price list. Printing did it when the press stopped needing plates. Manufacturing did it wherever the mould that had to be cut before the first part gave way to machines that will produce a batch of one. Compute has not made that transition. It is sold in units sized for the largest customer, and the small operator's bill is what that sizing looks like from underneath.

This is the problem we work on at LILY. The platform is designed so a small project does not have to start by buying and operating a whole production-shaped stack. There is no cluster to stand up, no control plane to keep current, and no orchestrator for the user to patch. The early beta is live, it runs in European data centres, and there is a free tier.

Which leaves the question the original thread never asked, and could not have, because it was arguing about reverse proxies.

What should small cost?

Not "less." That is a discount, and a discount still assumes the unit. The answer worth having is that the number should be a function of what you actually ran, and nothing else: no floor, no seat, no charge for merely existing. Nobody can say yet what that number is, because almost nothing on the market is divisible enough to find out.

You build. We carry.

Start for free.    Read the docs.    Talk to us about migrating a workload.

Step into the world after the cloud.
Start for free, integrate in minutes, and scale when you need to.

LILY

What is coming after the cloud

LILY Labs GmbH, Maudacherstraße 45, 67065 Ludwigshafen am Rhein
Amtsgericht Ludwigshafen am Rhein, HRB 70791

© 2026 LILY Labs GmbH. All rights reserved.