The price list moved in both directions this month. For a small business the first budgeting decision isn’t which model. It’s how you pay.
Google doubles the price of its default small model on January 1. It says so on the pricing page, in plain text, next to this year’s rate: one price “through December 31, 2026,” twice that “starting January 1, 2027.” If someone on your team ran a pilot on it this fall and wrote the pilot’s bill into next year’s budget, the number is wrong by half before the year starts. I know how that mistake happens because I made a version of it in June, watching my research agent’s bill hit $30 a day on a line I had budgeted as a rounding error.
Here is the short answer to the question in the title. You do not budget AI by picking a model and multiplying, because the price of a model is not a number anymore; it moves with how you use it and with the calendar. What you can budget is the way you pay. There are three, each has a different risk, and choosing among them is the decision a small business actually owns.
There are only three ways to pay for AI: flat, metered, or owned
Flat. A subscription per person per month. ChatGPT, Claude, Gemini, Copilot, most of the AI features inside tools you already have. Predictable, capped, and the cap is the catch: heavy days run into usage limits, and the plan does not scale to an agent that works while you sleep.
Metered. You pay per unit of work, through the API, which is what any custom agent, automation, or “AI feature we built ourselves” runs on. This is where the price list lives, and the price list is the problem.
Owned. You run a model on your own hardware. The price is the machine and the electricity, and you set it. For most small businesses this sounds like a developer’s hobby, and for the routine, repetitive work it is quietly the cheapest option on the list.
The budget question for 2027 is not “how much for AI.” It is “which of these three, for which work,” and that decision is what makes the number predictable.
The metered price is a function, not a rate, so it can’t be budgeted by model
This is the technical part, kept short, because your finance person will ask.
On a metered plan you are billed per token, roughly per word in and per word out. The rate per token looks like a fixed price. It is not, for four reasons, all taken from the vendors’ own pricing pages this week:
- The rate has dates on it. Google’s January 1 doubling is printed. Intro pricing has an end date, and this one is on the page.
- The same rate can cost you more per task. Anthropic’s pricing page notes that its newer models use a tokenizer that “produces approximately 30% more tokens for the same text.” The rate did not move. The bill for the same job did.
- Repeated work gets discounted, but only if it’s shaped that way. The same page shows cached input, the instructions and reference files an agent re-reads on every run, at a fraction of the fresh price, and on its top-tier models that fraction just fell to a quarter of what it was. Great if your agent repeats itself; irrelevant if it doesn’t. You would have to know which.
- Tools bill on a second meter. Web search on Claude is $10 per thousand searches on top of the tokens; search grounding on Gemini is free for the first 5,000 requests a month, then $14 per thousand. An agent that searches on every run has a cost line nobody wrote down.
One vendor, DeepSeek, charges by the hour of the day: half price outside its peak window, which is set in UTC on Chinese working days. That is not a complaint; it is a clear illustration of the point. The metered price is a function. You can budget a function only if you know its inputs, and most small businesses do not, because nobody told them the inputs existed.
Moving the judgment work to a flat plan ended my metered overages
In June I wrote that every agent has an operating cost and you should price a single run before you build one. The firsthand part was my research agent, Metis, burning through more than $300 in three weeks on a metered plan, on a path to about $900 a month. Moving the routine, repetitive work to a model on my own machine took that line to roughly $25 a month.
Here is the part I did not write in June, because it happened after. The metered overages stopped entirely. Not because the price list got kinder, but because I moved the judgment-heavy work, the part that needs the best model, onto a flat plan: the $200 tier of Claude’s Max plan. So the bill for a solo consultancy is now one flat subscription plus the cost of the machine that does the repetitive work, and no surprises since. The metered price list, the one that doubles in January and charges by the UTC hour, is a list I no longer have to read.
That is the whole recommendation for a small business, and it is the reverse of how the industry sells it. Flat for the work that needs judgment. Owned for the work that repeats. Metered only for a workload that has proven, with a priced run, that it earns its meter.
The 2027 AI budget line is three rows, not a model name
For each AI thing you pay for or plan to, the budget carries three things: which of the three ways you are paying, what the price depends on (nothing, for flat; the machine, for owned; the shape of the run and the calendar, for metered), and the date, if any, on which it changes. Three rows, not a model name.
Then one check before the number goes in the deck. For anything metered, apply every dated change you know about and ask June’s three questions again: what does one run cost, is the run worth more than it costs, and if not, what changes. If a doubling on January 1 turns a run that was worth it into one that is not, the conversation is not “how much.” It is “which way do we pay for this,” and that is a far better conversation to have in October than in February.
Uber spent its entire 2026 AI budget in four months and now caps each employee at $1,500 a month per coding tool. Nobody there misread a price. They budgeted a number where they needed a decision about the meter.
Three questions I get asked about this
We already pay for ChatGPT or Copilot seats. Is that the flat tier? Yes, and it is the right home for the work that needs judgment. The catch is the usage cap, not the price: a seat does not scale to an agent that runs while you sleep, and heavy days hit the limit. Budget it as a fixed line and watch the limit, not the bill.
Is a model on our own machine good enough for real work? For repetitive, well-specified work, yes; that is where mine went, and it is the cheapest tier on the list. For anything that needs judgment, no. That split is the whole recommendation, and it is why the two tiers are budgeted separately rather than as one “AI” line.
How do we find the dated price changes? On the vendor’s pricing page, in the date qualifiers and footnotes, not the headline rate. If a price says “through” a date, that date goes in the budget. If a page mentions a new tokenizer, a peak window, or a per-call fee for tools, those go in too, because each one moves the bill without moving the rate.
Sources and Further Reading
- Gemini API pricing — Google AI for Developers, accessed September 26, 2026
- Claude API pricing — Anthropic, accessed September 26, 2026
- Models & Pricing — DeepSeek, accessed September 26, 2026
- Uber caps employee AI spending after blowing through budget in 4 months — TechCrunch, June 2, 2026
- Your AI Agent Has an Operating Cost. Price It Before You Build It. — Pallas Advisory, June 30, 2026



