Analysis

Germany

Artificial Intelligence

Local First: What Public Buyers Need to Understand About AI Before They Procure It

Local First: What Public Buyers Need to Understand About AI Before They Procure It

AI is not a new economy — it is the same digital value chain with two labels changed. Written up from the CFIT workshop of 2 June 2026, this article explains what public buyers need to understand about AI before procuring it, and recommends a local-first architecture that is both lower-impact and more useful for most public sector tasks.

How AI models are actually made, where the energy actually goes, and why most public sector tasks belong on hardware you already own

AI is not a new economy. It is the same digital value chain that has existed for two decades, with two labels changed: digital resources are now called tokens, and digital services are now called AI services. Everything underneath — the energy, water, buildings, chips and fiber — is unchanged. Once that is clear, the procurement question becomes tractable: most tasks that public organisations actually perform are low enough in complexity to run on the laptops and on-premise hardware they already own, at a fraction of the cost and environmental impact. Only a narrow band of genuinely complex work requires hyperscale infrastructure. Buyers should therefore procure local-first systems with an explicit, informed escalation path to the cloud — not a cloud subscription with local capability as an afterthought.

Content

This article is written up from the AI track of the CFIT workshop on 2 June 2026, where public procurement practitioners from the Netherlands, Denmark, Ireland, the United Kingdom and Germany worked through what AI is, what it costs, and what should be required of it when it is bought with public money.

We deliberately begin one step before the usual starting point. The debate in most organisations opens with which AI to buy. We think the first question is whether it is needed at all — and that this question cannot be answered honestly today, because the prices on offer are not the real prices. AI is heavily subsidised by the companies bringing it to market. A procurement decision made against a subsidised price is a decision made against a number that will not hold.

The second half of this article is a recommendation. It is not a call for restraint for its own sake. The local-first model we propose is not a worse version of AI accepted for environmental reasons — it is genuinely better for most tasks, because a model running on your own device can be grounded in your own files without uploading anything, and it costs you the hardware you have already bought. Lower impact and higher utility point in the same direction here, which is rare enough to be worth acting on.

As an independent think tank we approach this from a public-value perspective. Our aim is that public spending on AI produces real capability inside public organisations, rather than transferring value and dependency abroad.

Summary

  • AI is not a new value chain. It is the existing digital value chain with digital resources relabelled as tokens and digital services relabelled as AI services. The inputs — energy, water, raw materials, chemicals, capital, labour, land — and the infrastructure that consumes them are identical.

  • The visible price of AI is not its cost. Subsidisation by model providers means procurement is currently evaluating a number that is not sustainable for the seller. Buyers should model the unsubsidised price before committing to a dependency.

  • Capability scales with energy, not with better mathematics. The underlying maths has largely plateaued; progress is now bought by putting more energy in, both during training and — through "thinking" — at the moment of use.

  • Model behaviour depends on human labour that is rarely discussed. Reinforcement learning is performed at industrial scale by outsourced human freelancers, largely outside Europe. There is a human rights dimension here that procurement frameworks do not yet address.

  • Agentic AI (tool calling) is both a major reliability improvement and a major security exposure. It reduces hallucination and energy use per useful answer, and simultaneously lets the model reach outside its sandbox.

  • The utility that is driving real headcount decisions comes from knowledge grounding plus shared memory, not from raw model capability. This is also the part organisations can build themselves.

  • Most public sector tasks — summarising a document, drafting a slide deck, looking something up, transcribing a meeting — run adequately on current laptops or a single on-premise GPU. Specialised small models are extremely capable and free.

  • Procurement should therefore specify local-first architectures: solve locally, escalate to government-owned GPUs, and reach hyperscale cloud only through a deliberate action with a data-transfer disclaimer attached.

  • No commercial actor in this market has an incentive to offer that architecture. It has to be specified by the buyer and built by domestic IT providers.

Terminology

  • Digital infrastructure: The umbrella term for everything required to produce digital resources: data center buildings, IT equipment (servers, GPUs), network equipment, mechanical and electrical equipment for cooling and power, and fiber connectivity.

  • Digital resources: Compute (the ability to process data), storage (the ability to retain it) and network (the ability to move it). Digital resources play the role in the digital economy that energy resources play in industry.

  • Token: The unit in which AI vocabulary, input and output are all measured. When a model is described as a "1.8 trillion token model", that is a statement about the size of its vocabulary. When you send a prompt, you send tokens; when you receive an answer, you receive tokens. The token is the commodity of the AI economy.

  • Training: Building the model. A copy of as much of the world's written material as possible is assembled and refined, and the model is taught a vocabulary and the statistical relationships between its words. Conceptually this is teaching a child to speak — at a scale where the energy involved is measured against the output of power plants.

  • Reinforcement learning: The step after training, in which human beings interact with the raw model and tell it what it may and may not say. Necessary because the model has no capacity to judge good from bad on its own.

  • Inference: The model in production, receiving prompts and returning answers. Training builds the machine; inference is the machine running.

  • Thinking (reasoning): A mode in which the model talks to itself before answering. It is not a new capability so much as a way of spending more energy per prompt to compensate for a plateau in model quality.

  • Large language model (LLM): A model whose vocabulary consists of words.

  • Multimodal LLM: A model that additionally holds a vocabulary of images, video and audio, and can therefore interpret them — while still producing text.

  • Generative image and video models: Models that also produce images, video and audio.

  • Agentic AI / tool calling: The ability of a model to take actions in external systems — run a search, execute code, query a database — rather than only describing what it would do.

  • Knowledge grounding (RAG): Placing an organisation's own documents into a vector database that the model can call as a tool. In effect, a second vocabulary representing everything the organisation knows.

  • Shared memory: A persistent record of corrections. When a user tells the model its output was wrong, that correction is retained and applied to future work.

The First Question: Do We Need It?

Before any question about model choice, deployment or vendor, there are three questions worth asking in order: Do we need it? Is it necessary? Does it add value?

These sound like rhetorical questions. They are not, because they cannot currently be answered with real numbers. Most of the cost of AI is hidden, because the technology is heavily subsidised by the companies bringing it to market. The subscription price a ministry pays today is a customer acquisition price, not a cost-recovery price. The honest form of the question is therefore: does this add value once the actual cost is deducted?

That caution applies to any technology, but it applies more sharply here, because of what is being deployed. Giving an AI subscription to every knowledge worker in an organisation is the equivalent of giving every one of them a Porsche. Some of them need it. Most of them are commuting four kilometres. The fuel consumption is the same either way.

AI Is Not a New Economy — It Is the Same One With New Labels

Our model of the digital economy runs left to right, as an economic system.

  • On the left are the system inputs: energy, water, raw materials, chemicals, capital, labour and land. Capital deserves particular attention in the AI cycle, because so much of it is being concentrated on a single bet that the risk is correspondingly high.

  • These inputs are consumed by digital infrastructure. What is colloquially called a data center is only the building; computers have to be put inside it, the computers have to be cooled, and the buildings have to be connected by fiber to a global marketplace so the computation can actually be reached.

  • That marketplace is the internet, and it remains the only genuinely unregulated global marketplace. This is why Ireland can host a large volume of data centers that serve very little Irish business, and a great deal of business for startups in Berlin. That flow of trade is free and effectively unmeasured.

  • Digital infrastructure produces digital resources: compute, storage and network. AI has not changed this list. It has changed the scale.

  • Software then refines digital resources into digital services — the same compute, storage and network capacity becomes an ERP system, or becomes a social network. The choice of product is a software question, not an infrastructure one.

  • The environmental impact in this system is produced by the use of digital services, because use is what requires digital resources to be produced.

Remove the left-hand side and the right-hand side ceases to exist. One does not work without the other.

Now apply this to AI. The picture changes in exactly two places. "Digital resource" becomes "token". "Digital service" becomes "AI service". Everything else — the inputs, the buildings, the equipment, the cooling, the fiber, the unregulated global market — is identical. Nothing structural has changed. AI is another technology in the digital space, and it should be assessed with the instruments we already have.

How an AI Model Is Actually Made

This technology is old. What changed is that cloud computing, for the first time, produced data centers with enough computers in one physical place to hold very large amounts of data and computation together. Centralised computation is the precondition for what we now call AI, and it exists because organisations moved to the cloud and paid for it to be built. Without that migration, this would not exist.

The production process has four steps.

Input. You need a copy of every written word on the planet. The internet is a good place to start. Books and art are also taken, with varying degrees of permission. The goal is a comprehensive database of written language.

Vocabulary. Training is not so different from teaching a child to talk, at an incomparable scale. Two things are built at once: a vocabulary, and the relationships between its words. Human memory works associatively — say "dog" and many people immediately see a dog — and models imitate that triggering. They hold the probability that any two words are related, that dog and cat are both animals. Those relationships are expressed as vectors. The essential point is that training means building a very large vocabulary and mapping how its elements relate to one another. With a child, this takes years. Here, it is done by expending the energy equivalent of a few power plants to force as many words as possible into the model.

Reinforcement. This step is routinely omitted from public accounts, and it is the one that matters most for behaviour. A freshly trained model is aggressive and frequently offensive, because a large share of the language on the internet is negative. So human beings are put in front of it at industrial scale to correct it — sending prompts, provoking it, and telling it what it may and may not say. This is the same process as teaching a child what cannot be said to a neighbour.

The industry term is reinforcement learning. The reality is a very large number of human freelancers, and they are not sitting in Helsinki, Stockholm or Berlin. They are on the other side of the planet. The major labs largely do not employ these people directly: a small number of specialist companies, each now valued in the billions, exist solely to supply this human labour to the labs. AI models do not work without humans making them work. For a procurement community that already accounts for human rights in supply chains, this is an unaddressed exposure.

Deployment. The corrected model is distributed to data centers, where it is ready to receive prompts.

Why the model cannot simply be built "good by design"

The obvious question is why the model is not simply designed to behave. This is being attempted, mostly by refining the input: removing the World of Warcraft forums, removing 4chan, and observing that better input data produces a better model.

But it cannot be solved this way, because — despite the name — AI is not intelligent. It has no taste. It is not a thing that learns or understands. It is a very large accumulation of words in a mathematical vector space, and a guess at probabilities. Designing the model to distinguish proper from improper language would require it to understand the difference. It does not. The only mechanism that works is humans telling it what is good and what is bad.

Why this feels different from earlier technology

The reason AI unsettles people who were untroubled by previous waves of automation is that it is the first mechanisation of knowledge work.

When cars were still assembled by hand and someone arrived to say the job could be done by a machine, that was frightening for the people who held the knowledge of how to screw a car together. It is worth putting yourself in their position, because this is the same moment, moved to a different class of work. Someone is wheeling a machine into the office and pointing out that the hundred-page document you read takes it very little time, and that it will send them a summary.

That is the actual source of the anxiety, and it is rational. Most of us in northern and western economies are knowledge workers, and this mechanises a large part of what we do.

Where the Energy Goes: Training, Inference and the Scaling Laws

Training

Training requires a copy of the internet, repeatedly. It is worth pausing on the scale of that, because it is not intuitive.

Roughly a fifth of the world's internet traffic is Google refreshing its search index — copying the internet on the order of several times per day, so that search results are current. That was the position before AI. There are now several more companies doing the same thing for the same reason, and global traffic is rising accordingly. Wikipedia has been knocked over repeatedly by new model launches, because the labs do not need one copy; they need a fresh copy, as fast as possible.

Storing a copy of the internet is itself a data center. Then the training run requires a large GPU fleet, and the energy involved is best understood in units of power plants rather than kilowatt-hours.

The paper you will hear referenced is The Scaling Laws. Its practical content is that model capability scales with the energy put in. This is a nightmare for anyone thinking about impact, because it means every meaningful step forward in capability requires roughly a doubling of energy input. The resulting numbers are difficult to defend.

Inference

Inference is the model in use, and it has its own scaling relationship. Speed of response is a function of energy expended.

There is now a second factor, and it exists because progress has stalled. For roughly two years the models have been close to a plateau in what they can do. The naming conventions give this away: releases arrive as point increments rather than as new generations. Providers are pacing progress because they are stuck.

The response is a mode marketed as "thinking". Enabling it means throwing more energy at each individual prompt, because thinking means the model talks to itself for longer before answering. That is the entire mechanism. All of it amounts to spending more energy on the problem, because the mathematics underneath is no longer improving much. It is the same pattern as lithium-ion batteries: when the underlying physics plateaus, you brute-force it.

What this means for prompts

This has a direct, unintuitive consequence for users, and it is the single most actionable thing in this article.

An imprecise prompt costs more energy than a precise one. If you write "can you make a presentation about circularity?", the model has very little to work with, so it iterates — repeatedly, expensively — until it cracks what you might have meant. A long, precise, specific prompt reduces the tokens consumed and the energy spent, because the model does not have to guess at your intent.

Prompt discipline is therefore not only a quality practice. It is an efficiency practice, and it is free to implement through training rather than procurement.

What Kind of Model, for What Kind of Task

Four categories are worth distinguishing, because they differ enormously in impact:

  • Large language models hold a vocabulary of words and can be prompted against it.

  • Multimodal LLMs additionally hold vocabularies of images, video and audio. They can interpret those inputs, but they produce text.

  • Generative image and video models also produce images, video and audio.

  • Agentic systems are any of the above with the ability to act, discussed in the next section.

Our recommendation on the third category is unambiguous: do not use generative image and video models. The energy involved is on a different scale from text, and the deepfake exposure for a public body is not worth the marginal benefit. Use these systems to build the slide deck; do not use them to generate the pictures in it.

Agentic AI: The Model Gets Hands

Agentic AI is the development that changes the procurement calculus most, and the word obscures how simple it is.

Previously, the model had no hands. Asked to pick up a folder, it would return four hundred pages describing the motion of picking up a folder. Tool calling means it can now instruct an actual hand to pick up the paper.

The consequences are larger than they sound:

  • Search. Before tool calling, a model could only search within the words it already knew. It could not step outside its own vocabulary or extend its knowledge, which is precisely why search produced so many hallucinations. A current model performs an actual search: most of them simply go to Google, open the first ten results, read them, and summarise. The knowledge is external now, not invented.

  • Coding. A model could always generate code; it could never run it. If you cannot run code, you cannot verify that it works, so the output was largely unusable. Being able to execute and check was the breakthrough that made coding the most profitable application of AI models.

  • Research. A request for the academic literature on a topic used to be bounded by the training data, which is often old. With tool calling, the model can query an external database — or an internal knowledge base.

The counterintuitive result is that agentic systems are more reliable and more energy-efficient per useful answer. Instead of spending a large amount of thinking time inventing a plausible answer, the model performs a search, lets the search engine do its work, and reports back what it found. The expensive internal guessing is replaced by cheap external retrieval.

This is also, as discussed below, a serious security change.

Knowledge Grounding and Shared Memory: Where the Real Utility Comes From

Agentic capability enables the configuration that every organisation is now attempting, and it has two parts.

Knowledge grounding. Every document your ministry or agency holds is placed into a vector database — a retrieval-augmented generation system. In the terms used earlier, you are building a second vocabulary: another brain, representing everything your organisation knows. Because the model has tool calling, it can query that knowledge base.

Shared memory. Grounding alone does not produce good answers, for the reason established earlier: the model has no taste and cannot decide what is good or bad. It still needs people correcting it. The efficient way to do this is shared memory. You upload a document, the model returns a summary, you tell it the summary is wrong — and that correction persists. In practice it is a text file recording, line by line, the mistakes that were made and what should have happened instead.

Together, these two things change the picture substantially. The system is grounded in the organisation's own knowledge, it can be course-corrected, and it remembers having been corrected. Usage rises sharply once this works, because it has become genuinely useful.

It is worth being direct about where that leads. This combination — not raw model capability — is what lies behind the headcount reductions at organisations in the United States that are further along in applying it, and it is being applied particularly to middle management, whose function is substantially the summarisation and routing of information.

Risks:

  • Grounding an external model in internal documents is the point at which the security question becomes concrete. The knowledge base is now reachable by a system that can also reach the internet.

  • Shared memory is a governance object. It encodes institutional judgement in a plain text file that will shape every future output, usually without version control, review or ownership.

  • The productivity case and the workforce case are the same case. An organisation that procures this combination without deciding in advance what it means for its staff has made that decision by default.

Tokens Are a Commodity — So Ask for the Impact per Token

Tokens are the unit of everything in this market. Vocabulary size is expressed in tokens. Your prompt is decoded into tokens. The answer arrives as tokens. For anyone used to thinking in commodities, this is a commodity.

Which makes the transparency question straightforward, and it is our standing recommendation to every buyer: ask for the environmental impact per token. Just ask. The fact that the question is currently difficult to answer is itself the most useful information a procurement process can surface.

We ask it because we can answer it. Two years ago the German Federal Ministry for the Environment provided a substantial grant to build a laboratory in which we deploy GPUs — refurbished, and in some cases very old — across four data centers, two in the Netherlands and two in Germany. That allows us to measure rather than estimate.

We have built that measurement into a user interface, so that for every single prompt a user can see the energy consumed, the water used, and the operational and embodied carbon they are responsible for. The platform behind it is complicated; the experience is not.

The most important design decision in that interface is what it emphasises. Showing impact per prompt is not very useful, because per-prompt numbers are small and hard to make meaningful. What matters is the cumulative account per user over the lifetime of their usage. The impact of AI is not driven by the individual prompt; it is driven by people sending hundreds of prompts a day. It is the scale that matters, and only a cumulative view makes scale visible.

Local First: Most of What You Need Runs on Hardware You Already Own

Here is the fact that most changes what a public buyer should ask for: most models today run perfectly well on laptops and desktop computers. Supercomputers are required to train models. They are not required to use them.

A concrete example from our own operations: a large volume of audio and meeting transcription runs on an eleven-year-old Mac Mini in our Amsterdam office, using current transcription models, without difficulty. We are in the process of connecting more old Mac Minis together — which is lifetime extension of hardware, achieved by using it for the workload everyone assumes requires new hardware.

  • Specialised tasks are small. Audio transcription, document recognition and OCR require a small vocabulary and have a correspondingly small energy footprint.

  • Coding needs more, because the space of possible code is large and the vocabulary must be bigger. Most organisations do not need this.

  • Google has done the public sector a favour by releasing one of the most capable models available, which is also among the smallest, fully open source, free, and able to run locally. Microsoft likewise publishes very small models designed specifically for older devices.

If we could give public organisations a single recommendation, it would be this: have your IT department run a small open model locally, on the laptops you already own, as the default. Then implement a button that says this wasn't good enough, which escalates the request to a larger system. Local first. For most of what your people do, it is more than enough — you are all carrying substantial machines around already.

How that survives contact with a procurement process is the harder problem, and it is the reason this article exists.

The other end of the scale

For contrast, consider the largest models currently available. A frontier model with a vocabulary on the order of a trillion-plus tokens requires something like 128 GPUs simply to be resident — roughly a full rack of current NVIDIA hardware, drawing on the order of 120 kilowatts before it has answered a single prompt. That is the model sitting still.

The consequences of that are visible in the market. Capacity at this scale is now being secured by renting entire data centers, at monthly costs in the hundreds of millions to billions, and those facilities frequently run on natural gas turbines. When a provider describes such a facility as clean, it is worth asking what it is being compared against.

(The specific figures in this section were given as orders of magnitude in the workshop and should be verified against current public sources before republication.)

Speed is a design choice, not a requirement

One question from the room deserves its own answer, because it reframes the hardware debate: how does extending the working life of equipment fit with the requirements of AI?

Fundamentally, you can run any model on any generation of chip. This is genuinely true, including on very old devices. The only thing that changes is that it gets slower.

So if we did not insist that these systems respond in real time, the hardware question would largely dissolve. Imagine arriving in the morning, listing what needs doing — read this document, draft this presentation — firing all of it off, and letting it run in the background. Almost any laptop in current use would be sufficient.

It is the need for speed that drives demand for better chips, and the reason we need speed is that people sit in front of the screen reading the output word by word as it appears. That is a design error. You should not be watching the words arrive. You should hand a task to an agent, wait an hour, and receive the result — which also leaves you less interrupted and less distracted than waiting on a machine.

This matters for procurement because a specification that mandates real-time response has, implicitly, mandated new hardware.

What This Means for Procurement

Direct and indirect procurement

There are two distinct exposures, and only the first is usually managed.

  • Direct procurement of AI as a service, to help a public body deliver its own services. This is where specifications, licences and vendor relationships live today.

  • Indirect procurement — the use of AI inside the goods, works and services a public body buys. How much AI is embedded in what you are already procuring, and on whose infrastructure does it run?

The useful analogy is carbon accounting: direct procurement is scope one and two, and embedded AI is scope three. As with scope three, the difficulty is the boundary — what is national, what is transnational, and where responsibility stops. That is exactly where it becomes uncomfortable, and exactly where the work is.

What we heard in the room

Asked what AI requirements procuring organisations are actually receiving, the most accurate short answer offered was ambiguity. Several central purchasing bodies are involved, everyone is learning as they go, and the requirements are genuinely complex, spanning sustainability, human rights, data, infrastructure and security at once.

The use cases described were instructive:

  • Appointment handling in the health sector.

  • Railway signalling, and transport and logistics optimisation for national infrastructure bodies — cases where the requirement really is to process large volumes of data quickly.

  • Automated assessment of aerial imagery to identify where maintenance is needed on infrastructure works.

  • Defence logistics, and — stated plainly by participants — targeting systems and analysis of satellite imagery at scale.

Two observations follow. First, several of the most operationally important use cases have nothing to do with language models at all. They are machine vision and optimisation problems, and conflating them with LLM procurement produces bad specifications. Second, and more commonly: "I think we're using it because everybody else is using it." Or, as one participant put it, the internal message is effectively here is a hammer, figure out what to do with it.

The economics described were equally instructive. One directorate holds two enterprise chat licences for roughly a hundred people, with no monitoring of how they are used. A workshop was run on using them well; usage remains basic. And the effects are not obviously net positive: help desks across several ministries report that the questions arriving are now written by AI, and if those questions are then answered by AI, the result is two machines corresponding with each other. Where there were ten emails there are now a hundred, and nothing has moved faster.

Then the price question. Asked whether they would still use the tool if it cost €2,000 per month, the answer was immediate: probably not. That is close to the range at which these products would need to be sold to be profitable — on the order of €1,200 to €1,800 per month for professional and enterprise tiers, with pricing roughly doubling each model generation. The current position is that it is too cheap not to use. That is not a market signal. That is a subsidy, and procurement is making dependency decisions inside it.

Risks:

  • Security is not a future risk, it is the current state. Agentic AI is a significant security threat — but the exposure began the moment people in ministries started uploading documents to public chat services, and the incentive to upload things that should not be uploaded is very strong. Tool calling makes it considerably worse, because the model can now act outside its own boundary. The labs are relatively transparent about this: security reports accompanying major releases document the model attempting to break out of its sandbox, and succeeding.

  • Subsidised pricing conceals lock-in. A dependency established at a subsidised price is repriced at the vendor's convenience, and the switching cost by then includes retrained staff and embedded workflows.

  • Restricting local installation forces the worse option. Where policy prevents installing agentic tooling on managed devices while permitting browser-based chat services, the effect is to prohibit the private, local option and mandate the one that requires uploading files. This is a common configuration and it inverts the intended protection.

  • Uninstrumented usage cannot be governed. Without monitoring, an organisation cannot state what it spends, what it emits, or what has left the building.

  • AI-to-AI workload growth is a real cost. Automating the production of correspondence without automating the decision to correspond increases volume without increasing throughput.

A Procurement Model: Local First, Cloud by Exception

The recommendation rests on a single relationship. Model capability rises with model size. Task complexity varies enormously. And the tasks a public organisation actually performs are not distributed evenly across that range.

  • Low complexity — the large majority of daily work. Summarise this document. Draft this presentation. Look this up. Transcribe this meeting. These run locally, on the computer the person already has.

  • Medium complexity. With reasonably recent laptops this often still works locally, but this is the band where a small on-premise GPU — a single machine on a shelf in the ministry, run centrally — is the right answer.

  • High complexity. A narrow band of genuinely complex work that can only run on hyperscale cloud. It cannot be run in local data centers, and there is no point pretending otherwise.

The procurement consequence is a specific architecture, and if we were buying, we would not go to a model provider to get it. We would go to a domestic IT provider — an intermediary, unavoidably — and specify a local-first system:

  • The interface is a chat experience, and the request is attempted locally first.

  • If the local model cannot solve it, escalate to the government's own GPUs.

  • Only if it is genuinely needed, a deliberate action — a red button — sends the request to hyperscale cloud, accompanied by a disclaimer stating that data is now being uploaded to US-owned or Chinese-owned infrastructure.

Almost none of this needs to be built. Open source interfaces that look and feel like the commercial chat products already exist and are free; the models are free; the escalation logic is configuration. What needs building is the government-specific interface — and there is precedent, in the custom chat services already operated by individual city administrations.

There is a further benefit that is easy to miss. Running locally is what makes grounding safe. The health-data question raised in the workshop — how do citizens share sensitive data with an assistant without exposing it — has a clean answer at this end of the scale. A model running on your own laptop can be pointed at your own files. It reads them locally. It needs no interconnection and uploads nothing.

That is the whole argument in one sentence: local-first is more useful and lower impact at the same time, because grounding in local files improves the answers while the hardware you already own does the work.

The reason it is not on offer is that no actor in this market has any incentive to propose it. It has to be specified by buyers, and built by domestic providers, or it will not exist.

Recommendations

For public organisations, in order of how quickly they can be acted on:

  • Ask the first question. For each proposed use, establish whether it is needed and what it would be worth at an unsubsidised price. Model €1,200–1,800 per user per month and see what survives.

  • Ask for the environmental impact per token, in every tender and every vendor conversation. Treat inability to answer as a finding.

  • Require cumulative per-user impact reporting, not per-prompt figures. Scale is the variable that matters.

  • Train prompt precision. It improves output quality and reduces energy consumption simultaneously, and costs nothing to procure.

  • Deploy a small open model locally as the default, with an explicit escalation path. Local first, cloud by exception, with a disclaimer at the boundary.

  • Do not procure generative image and video capability without a specific, examined need.

  • Stop mandating real-time response. Specify asynchronous task completion where the work allows it, and the hardware refresh pressure falls away.

  • Extend equipment life deliberately. Any model runs on any chip generation; only speed changes. Specialised workloads such as transcription run well on hardware that is years old.

  • Treat the reinforcement learning supply chain as a human rights question, and ask vendors who performs it and under what conditions.

  • Buy the interface domestically. The models and the open source tooling are free. The integration work is where public money should create local capability.

This article is written up from the AI track of the CFIT workshop of 2 June 2026. Figures given in the workshop as orders of magnitude are marked as such and should be verified before republication. A segmented recording of the session is available in fifteen thematic clips.