← Back to writing

Doing More With Less Is a Trap: Why Cheaper AI Will Not Save Government

10 min read
Doing More With Less Is a Trap: Why Cheaper AI Will Not Save Government

Your agency's AI bill will keep climbing even as the price per token falls, because "do more with less" is the wrong promise and cheaper inference does not bank a saving. It buys more inference, more sprawl, and a bigger bill (Jevons, 1865). The assurance we wrap around it measures the wrong thing and then gets gamed until the dashboard is green and the risk is not managed (Strathern, 1997). And because every platform your agency already owns now ships its own model, government AI is fragmenting to mirror the shape of the org chart, one captive feature at a time (Conway, 1968). The public value never turns up, because none of this compresses the thing that actually matters: the time it takes to turn raw information into a decision a person can defend and a citizen can contest.


The win is not cheaper AI. It is a shorter, assured, governed path from a question to a defensible decision, built on tools your staff already hold. Stafford Beer put the test plainly: the purpose of a system is what it does (Beer, 2002). If your AI programme reliably produces more spend, greener dashboards, and no faster decisions, then that, and not the business case, is its purpose.

If your AI gets ten times cheaper and your Agency does ten times more with it, what exactly did you save?

Introduction

It is budget week. Someone walks into the executive meeting with good news. The new model costs a fraction of last year's, the price per token has fallen through the floor, and the efficiency dividend is banked on a slide before a single decision has been made faster or better. Heads nod. The word productivity gets said three times. The meeting moves on.

But the good news on that slide is not the story worth telling. The uncomfortable part is that the cheaper the AI gets, the more of these meetings end with a bigger bill and a slower agency, and almost nobody in the room can say why. The why is old, and it is not a technology problem. It is three laws that were true long before any of us started building these systems, and they do not care about your business case.

The Efficiency Promise, and Why It Inverts

In 1865 the economist William Stanley Jevons noticed something awkward about coal. As steam engines got more efficient and the cost of burning coal fell, Britain did not use less coal. It used far more, because cheaper coal made a hundred new uses of coal worth it (Jevons, 1865). Efficiency did not shrink demand. It unleashed it. The modern version is called the rebound effect, and it is why efficient lighting did not lower your power bill and a faster commute did not hand you back your evening.

Inference is coal. When the cost per task drops, the honest response inside an agency is never to do the same amount of work for less money. It is to find ten new things worth doing that were not worth doing at the old price. Summarise every case, not just the hard ones. Draft every letter. Score every transaction. Triage every alert. Each is defensible on its own. Together they are Jevons in a business suit. Usage climbs, the meter runs on every token, and the total bill goes up while the per-unit price goes down.

I will be careful not to overclaim it. The size of the rebound varies by task: some new uses genuinely pay their way, others are pure waste, and the effect is a tendency, not an iron law that guarantees your bill balloons. That variation is exactly the argument for the control I keep coming back to. If you cannot say in advance which new uses are worth it, you need a measure that decides it after the fact, and that measure is cost per decision, not cost per token.

A service-delivery lead running a visa or welfare backlog has the strongest version of the objection to everything above, and it deserves answering directly rather than waved off as a rebound anecdote. For them, "more" is the queue of real people who have been waiting months for a decision because processing every case individually was never affordable at the old price, and cheaper inference is the first time in years that clearing the backlog properly is actually within reach. Calling that Jevons in a business suit sounds like an argument against helping the people at the back of the queue.

That objection is right, and I concede it in full. More usage is sometimes exactly the point, and a hospital that uses cheaper inference to clear a diagnostic queue faster is finally delivering the service the queue always needed, not committing rebound waste. Where the distinction looks like more versus less, it is really whether the extra usage compresses time to a defensible decision for someone who was genuinely waiting, or whether it is discretionary volume that exists because it is now cheap to generate: summarising cases nobody asked for, drafting variants nobody will read, scoring transactions nobody will act on differently. The backlog-clearing case passes that test cleanly, because someone at the other end gets an answer sooner. Most of the "do more" that shows up on a budget slide does not pass it, because nobody can point to who benefits from the twentieth summary of a case that was already fine. Jevons is an argument for being honest about which "more" you are buying, not an argument against doing more.

Notice the sleight of hand in the language. "Do more with less" assumes the "more" is free. In a metered, usage-priced world, the "more" is the expensive part.

Drag the price down. Total spend climbs while the pitch line falls. That gap is Jevons.

When the Metric Becomes the Target

So leaders reach for control. They stand up governance: an AI register, an assurance process, a set of controls to tick, a continuous authority to operate (one day). Good. Except the moment you manage AI by a metric, you meet Charles Goodhart (1984). As Marilyn Strathern sharpened it, when a measure becomes a target, it ceases to be a good measure (Strathern, 1997).

Count models registered, and teams register the harmless ones and route the real work around the register. Measure controls passed, and the controls get passed on paper while the model drifts underneath. Certify a system once and call it assured, and you have certified a photograph of something that changes the day after you sign it. That last one is the trap I see most. A continuous authority to operate is meant to be a live state, not a certificate on a wall, and I have written before about why (Hall, 2024). Managed as a certificate, it games itself: the dashboard goes green, the evidence is a snapshot, and the actual behaviour of the model in production belongs to no one's metric.

A governance system that reliably produces green dashboards and shelfware assurance has, as its real purpose, the production of green dashboards and shelfware. Not safety. What a system does is what it is for, whatever the policy says it is for.

Push the metric and it climbs. The outcome it stood for peaks, then slips.

Your Org Chart Is Now Your AI Architecture

Here is the third law, and it is the quiet one. In 1968, Melvin Conway observed that any organisation designing a system produces a design whose structure copies the organisation's own communication structure (Conway, 1968). Build software in four teams that do not talk, and you ship four subsystems bolted together at the seams.

Now watch how government actually buys AI. It is not buying "an AI system" it can point at and govern. Every platform it already owns has quietly shipped a model. The workflow tool has an assistant. The customer system has an assistant. The document suite, the case-management system, the security tooling, each ships a captive model, each procured by a different team on a different contract, with its own logging and no shared evaluation. This is Conway's Law running at speed: fragmented buying produces a fragmented, ungoverned estate of dozens of little models, and the aggregate slips through every framework that assumes there is one system to inspect. No architect drew this. Procurement drew it, one captive feature at a time, and the org chart signed off.

The fix is boring and it works. Route the reasoning-grade tasks through one governed model layer you own, a thin layer that sits in front of whatever models you choose, with a single place for logging, provenance, evaluation, and swapping models when the frontier or the politics moves. Let each platform's built-in AI keep the small stuff that never leaves its own data. It costs integration effort. It buys back a system you can actually govern, and a shape that matches the accountability you are on the hook for, rather than the accidental shape of who bought what.

Left, the estate the org chart drew. Right, the single seam accountability needs. Same platforms, very different governability.

The So What?

So if cheaper AI does not save you, what does? The loop does, specifically the only number that earns its keep in public-sector AI: time-to-decision, how long it takes to turn raw information into a decision a person can defend, lawfully, and a citizen can contest when it is wrong. Compress that honestly and you have public value. Everything else is motion.

A few moves, and none of them start with buying a cheaper model.

  • Measure cost per decision, not cost per token. A cheap model that triples the review queue is not cheap. Model the whole loop: volume, tokens per decision, model price, human-review rate, reviewer cost. If the economics only work at pilot scale, the capability is already broken.

  • Make time-to-decision the target, then watch it for gaming. If the AI does not shorten the path to a defensible decision, it is a cost with a nice demo. And the moment you can hit the target without the decision getting faster or safer, you have built a Goodhart machine.

  • Govern one model layer, not forty embedded ones. Bring consequential reasoning under a thin model layer you own. Hand the small, coupled tasks to the platform's built-in AI, and hand the assessments that will be scrutinised to somewhere you can log, evaluate, and swap. Be precise about what "own" means: you own the governance seam, the logging, evaluation and routing, not the model weights. And this is not a single point of failure, because the layer's job is to make models interchangeable behind it, so no one model is load-bearing. "Swap" is doing quiet work in that sentence, mind: changing the model underneath means re-evaluating against your own tasks before you trust it, so the seam is what makes the swap safe, not free.

  • Treat assurance as a live state, not a certificate. Evaluate against production behaviour continuously, not once against a snapshot. Robodebt is the standing proof of what an opaque, uncontestable automated decision does at scale, and the lesson generalises straight onto today's models (Royal Commission into the Robodebt Scheme, 2023).

  • Uplift the staff you have, with tools they already hold. The fastest safe gain is rarely a new frontier model. It is giving an analyst a governed, assured version of a tool already in their hand, inside the loop they already run, so the human stays in command and the decision gets quicker without getting sloppier. Norbert Wiener warned decades ago against building systems that turn people into servants of the machine (Wiener, 1954). A human dropped into the loop only to rubber-stamp is exactly that, and it is not oversight.

This article is getting long and you get the point. There is a proper conversation to be had about routing, caching, and model tiering as the actual cost controls, and that is one for my architecture friends another day.

Interactive: make inference cheaper, push on the metric, and watch the bill and the gap do the rest.
Interactive: route consequential tasks through one governed model layer, and swap models without rework.

Conclusion

The efficiency story is seductive because it sounds like thrift and feels like progress. But thrift measured on the per-unit price of inference is a mirage. Jevons says the cheaper it gets, the more you will tend to use, and unless that extra use is winning back something real, like the backlog cleared above, your bill climbs regardless of the per-unit price. Goodhart says the moment you manage it by a metric, the metric gets gamed and the risk goes unmanaged. Conway says that if you let procurement draw your architecture, you will spend your life governing a shape you never chose. None of this is a reason to avoid AI. All of it is a reason to stop buying the promise and start buying the outcome, which is a shorter, assured, contestable path to a decision a person can stand behind.

So before you bank next year's efficiency dividend, ask the one question that survives contact with these three laws. When your AI is ten times cheaper and doing ten times more, will your agency decide faster and defend those decisions better, or will you have simply bought a bigger, greener, more fragmented way to do exactly what you did before, at a price that quietly went up?

Thanks for reading.

The views expressed in this article are my own and do not represent those of my employer, or any of my clients.


References

  • Beer, S. (2002). What is cybernetics? Kybernetes, 31(2), 209–219. https://doi.org/10.1108/03684920210417283

  • Conway, M. E. (1968). How do committees invent? Datamation, 14(4), 28–31.

  • Goodhart, C. A. E. (1984). Monetary theory and practice: The UK experience. Macmillan.

  • Hall, B. (2024). Embedding FinOps: Achieving a continuous authority to operate (cATO). LinkedIn. Embedding FinOps: Achieving a Continuous Authority to Operate (cATO)

  • Jevons, W. S. (1865). The coal question: An inquiry concerning the progress of the nation, and the probable exhaustion of our coal-mines. Macmillan.

  • Royal Commission into the Robodebt Scheme. (2023). Report. Commonwealth of Australia. https://robodebt.royalcommission.gov.au

  • Strathern, M. (1997). "Improving ratings": Audit in the British university system. European Review, 5(3), 305–321.

  • Wiener, N. (1954). The human use of human beings: Cybernetics and society (2nd ed.). Houghton Mifflin.

9 min read1 like

Exempt by Design: The AI Governance Gap in the National Intelligence Community

The TL;DR is that we assume the most consequential users of AI must be the most tightly governed, and in Australia it runs the other way. The mandatory framework for government AI, the Digital Transformation Agency's policy now at v2.0 with a mandatory use-case register, exempts the Defence portfolio and the National Intelligence Community (NIC) (Digital Transformation Agency, 2025). The National AI Plan shelved the proposed mandatory high-risk guardrails and stood up an AI Safety Institute instead (Department of Industry, Science and Resources, 2025). What is left for the community is point-in-time authorisation under the Protective Security Policy Framework and an oversight office that has appointed a single Chief AI Officer (Inspector-General of Intelligence and Security, 2025).

AIGovernanceNational SecuritySovereign AIOversight
7 min read1 like

Starved by Design: The Sovereign AI Paradox

The TL;DR is that Sovereign AI is being sold as a capability upgrade for our most sensitive work, when the opposite is closer to the truth. The more classified the network, the worse the Artificial Intelligence running on it tends to be, and that is baked into the physics of how these systems are built, not a funding gap we can spend our way out of. As Australia invests in Sovereign and next-generation capability through programs like the Australian Signals Directorate's REDSPICE and the National AI Plan (Australian Signals Directorate, 2022; Department of Industry, Science and Resources, 2025), we owe ourselves an honest conversation about what a sealed enclave can and cannot do.

Sovereign AICLOUD ActGovernance
15 min read2 likes

The Undeclared Dependency: Every Technical Supply Chain Has a Strait of Hormuz

The TL;DR is that AI cannot secure a supply chain you have not described. Australia's fuel stocks, your cloud builds, and the defence industrial base are all exposed through the same weakness, dependencies nobody wrote down, and the unglamorous act of declaring the graph and then verifying it continuously is the actual security control. The clever model you point at the problem comes a distant second.

Critical InfrastructureSupply ChainAINational SecurityArchitecture