Einblick in die hochmoderne Fertigung der CiS electronic GmbH in Krefeld: KI-gestützte Prozesse steigern Effizienz, Qualität und Innovation in der Kabelkonfektion. © CiS electronic GmbH

Research

AI Agents

Compute is the constraint. The endpoint is the tier nobody is pricing

Article by

Dr. Anoj Winston Gladius

·

In late June 2026 it was reported that Google had told Meta it could not have as much access to Gemini models as it wanted, because Google did not have the capacity to serve it. Two of the five most valuable companies in existence, both with effectively unlimited cash, and the one that designs its own silicon could not supply the other. In the same period Google raised its 2026 capital expenditure guidance to between 180 and 190 billion dollars, said it expected that figure to increase significantly in 2027, moved to raise close to 85 billion dollars in equity explicitly to fund AI compute infrastructure, and signed a lease worth roughly 920 million dollars a month for data centre capacity from xAI. Inside the company, DeepMind researchers were queuing for TPUs that had been sold to external customers — Anthropic having contracted for up to a million Ironwood chips, five gigawatts of capacity across five years, in an arrangement worth up to forty billion dollars, with Meta signing separately. Demis Hassabis acknowledged the constraint in public and located it precisely: a few suppliers of a few key components, which is the courteous description of a high-bandwidth memory bottleneck sitting across three vendors and gating every advanced accelerator on the planet. Researchers, he noted, need a great many chips to experiment on new ideas at a meaningful scale. None of this is the behaviour of a company that is hedging. It is the behaviour of a company that has hit a physical wall — and a wall that binds on Google binds on everyone behind it. For a European enterprise the interesting question is not who wins the capacity race. It is what you do about the fact that the winners will ration you, that the meter is being redesigned to charge for compute-time rather than tokens, and that the best inference silicon on earth cannot lawfully enter your building.

This is the eighteenth piece in a series I have been writing for neuland.ai. [¹] Earlier pieces have argued that the model is a component rather than a moat, that the routing table matters more than any single model choice, and most recently that sovereignty of location is about to commoditise while sovereignty of control is not for sale. [²] This piece is about the layer underneath all of that, which the previous seventeen have treated as a given: the compute itself. It has stopped being a given.

What the two hedged giants actually did

There is a tempting story in which Google and Apple are sitting out the AI race while OpenAI, Anthropic, xAI and the Chinese labs fight it out. It is wrong about Google in a way that is easy to check, and it is wrong about Apple in a more interesting way.

Google is not sitting out. Google is running in the race while simultaneously supplying most of the other runners. It trained its own frontier model entirely on its own silicon, shipped an enterprise agent platform to general availability, and sold vast blocks of accelerator capacity to two of its most direct competitors — with the consequence that its own research organisation now waits behind paying customers, and at least one long-tenured researcher has left for a startup with better access to chips. [³] That is a company so committed to the race that it is cannibalising its own experimental capacity to serve the market, and renting from a rival to cover the gap.

Apple looks more like abstention and is not. Apple assessed its own models as twelve to eighteen months behind the frontier on conversational quality, reasoning and multi-step task completion, and in January 2026 licensed a custom Gemini variant for roughly a billion dollars a year rather than shipping something inferior and waiting. [⁴] But look at the architecture that arrived at WWDC in June rather than the headline. A small on-device model of roughly three billion parameters handles latency-sensitive and privacy-sensitive work. A controlled cloud environment handles the middle. The heaviest reasoning routes to the licensed frontier tier. Apple's own next-generation models and its own server silicon are in development, with the intention of substituting the licensed model out from 2027. [⁵]

And the detail that matters most for this argument is one that received almost no coverage outside developer commentary: Apple's foundation model framework is no longer restricted to Apple's own models. Any model conforming to a new public protocol can be used, which means an application can swap the licensed frontier model for a competitor's, or for a local one, without rewriting its features. The developer assessment was blunt — provider abstraction has become table stakes, and coupling your own codebase to a single provider now looks like technical debt. [⁶]

So the two companies with the most to lose from getting this wrong both arrived at the same architecture: rent frontier capability where it is genuinely needed, retain control of the inference layer, keep small models running locally for work that is private or must be fast, abstract the provider so substitution stays cheap, and build the substrate underneath in the meantime.

That is not abstention. It is the most sophisticated hedge available, executed by organisations that can afford any option they like. It is also, almost line for line, the architecture this series has been arguing for since February — which I note not as vindication but because it is the strongest available evidence that the architecture is correct rather than merely convenient for us.

Three consequences that land on the enterprise

The meter is being redesigned, and not in your favour. Google moved its consumer AI tiers to compute-based usage limits, where allocation is consumed according to the actual processing required rather than the number of requests made, and removed the bundled monthly credits from the paid tier. Its enterprise platform introduced separate billing for agent runtime. [⁷] The direction is unmistakable and it is arriving at precisely the wrong moment for buyers: metering is shifting from tokens toward compute-time exactly as agents begin running for hours rather than seconds. A pricing model indexed to elapsed compute, applied to workloads whose defining characteristic is that they run for a long time, is not a rounding error on the budget.

Rationing runs downhill. The most instructive fact of the quarter is not the size of anyone's capital expenditure. It is that a hyperscaler under capacity pressure limited the model access of one of the five most valuable companies in the world. If that is what happens to Meta, no enterprise should build a critical workflow on an assumption of guaranteed allocation. This is the same argument the flexibility piece made after a model was withdrawn by export-control directive, arriving from a completely different direction: your access to frontier capability is contingent on someone else's constraint, and contingency has to be designed for rather than hoped about.

The best silicon is structurally unavailable to you. This is the one I would put in front of a board. Google's accelerators are reachable only through Google Cloud. There is no on-premises path, no licensing route, no air-gapped option. The forthcoming inference-optimised generation claims something like an eighty percent improvement in performance per dollar for low-latency serving of large mixture-of-experts models — and any organisation with a data residency requirement, an air-gapped network, or a policy against public cloud cannot use it at all, regardless of how good it is. [⁸] That exclusion is permanent and it is architectural. It does not soften as the vendor relationship matures.

Put those three together and the conclusion is uncomfortable for anyone whose AI strategy consists of a contract with a hyperscaler. The capability is real, the price is rising, the metering is being restructured, the allocation is discretionary, and the best of it is off-limits to precisely the enterprises this series is written for.

The compute you already own

Here is the part that almost nobody is pricing.

An enterprise with ten thousand employees owns roughly ten thousand devices containing neural processing units, plus a comparable number of phones. Apple's current silicon has redesigned neural accelerators in every core with something like forty percent more AI throughput than the preceding generation, and equivalent hardware now ships in essentially every business laptop sold. That fleet was purchased on a refresh cycle that has nothing to do with AI, it sits inside the corporate perimeter by construction, and for inference purposes it runs at close to zero utilisation.

It is also the one compute layer that cannot be rationed by a supplier, cannot be withdrawn by an export-control directive, cannot be repriced mid-contract, and does not require a data residency negotiation — because it is already inside the building and already paid for.

For Europe specifically this matters more than it does elsewhere. Europe is not going to win a capacity race against six hundred billion dollars a year of hyperscaler capital expenditure; that argument was over before it started, and pretending otherwise has been a persistent European weakness. But the endpoint fleet across European enterprises is very large, entirely inside European jurisdiction, owned outright rather than leased, and almost completely idle. The previous piece in this series argued that what survives commoditisation is what capital cannot buy. Compute capacity is the clearest example of something capital can buy, which is exactly why competing on it is futile — and exactly why the compute Europe has already bought is worth taking seriously.

I want to be careful here, because the naive version of this argument is wrong in several ways, and the honest version is more useful.

The enterprise corpus does not live on the endpoint. Retrieval still has to cross to a server where the index and the permission model live, which means local compute serves the reasoning steps rather than the grounding. Distributing and updating models across thousands of heterogeneous devices is a genuine operations problem, not a footnote. The variance between a five-year-old laptop and current silicon is enormous, so any routing decision has to account for device class. Local execution does not confer privacy by itself — the harness still has to enforce entitlement, and an unconstrained local agent is exactly as dangerous as an unconstrained remote one. Small models are capable, not frontier-capable, so workload triage is essential rather than optional. And users notice when the fan spins up.

Which resolves into a claim that is modest in form and significant in consequence: the endpoint is a tier in the routing table, not a replacement for the cloud. The routing decision that this series has been describing for a year — which model, at what cost, under which policy, in which jurisdiction — acquires one more dimension. Where does this run.

Why bring-your-own-model stopped being interesting

Model agnosticism was a differentiator for about eighteen months. It is now a protocol in a consumer operating system. When the platform vendor ships provider abstraction as a public interface so that any conforming model can be substituted without rewriting the application, the argument is over: bringing your own model is a checkbox, and any vendor still presenting it as a strategic advantage is describing table stakes.

The interesting question has moved one layer up. Not which model runs but whose loop runs it.

By the end of 2026 a large share of enterprise applications will ship with task-specific agents already embedded, up from a very small fraction a year earlier. [⁹] Those enterprises have built harnesses. They have agent loops, tool definitions, evaluation suites, and opinions about all of it. What they do not have — as the previous piece in this series argued at length — is permission-aware retrieval that trims to the requesting identity before the model reasons, a maintained model of their own business to ground against, an append-only record that answers what a given identity could have seen at a given moment, or a routing table that survives a provider being rationed or withdrawn.

Telling those customers to rebuild on someone else's agent framework is a bad offer. Bring your own harness is the better one: keep the loop you built, and plug it into the governed substrate underneath. The platform stops trying to be the agent and becomes the thing the agent runs on.

Where neuland.ai stands

The neuland.ai HUB has been model-agnostic by construction from the beginning, for the structural reason the earlier pieces set out: we have nothing to sell at the model layer, so the routing decision can be made on the workload's merits rather than ours. Deployment options span sovereign cloud, private cloud, on-premises and genuinely air-gapped operation, which is possible only because we can run entirely on weights the customer holds — and which is precisely the capability that cloud-only accelerators structurally cannot offer.

Two extensions follow from the argument above and are where our engineering attention currently sits.

The first is bring your own harness: the customer's existing agent loop, tool definitions and execution logic connected to the platform's permission-aware retrieval, ontology, audit trail and model routing, rather than replaced by ours. Alongside it, durable streaming for long-running work, because an agent that runs for hours needs to survive interruption, resume from where it stopped, and remain inspectable throughout — properties that have to be in the substrate rather than in the agent.

The second is the endpoint as a routing tier. If compute is the binding constraint and cloud compute is rationed and repriced, then routing appropriate work to hardware the customer already owns is a cost and sovereignty optimisation at the same time. The honest framing is that this is a direction rather than a finished capability: the reasoning tier moves more easily than the retrieval tier, the device-class problem is real, and the model distribution question is an operations discipline we would rather solve properly than announce early. But the architecture already treats where inference happens as a policy decision rather than a product constraint, which is the part that is difficult to retrofit. [¹⁰]

The research effort behind this remains systems and infrastructure work rather than model research. Nobody at neuland.ai is going to design an accelerator. What we can do is make sure that when the compute constraint tightens further — and every public statement from the people who own the constraint suggests it will — the customer's architecture treats that as a routing change rather than an outage or a renegotiation.

Personal take

The compute squeeze is the most under-discussed strategic fact in enterprise AI, and I think it is under-discussed because it is unflattering to everyone. It is unflattering to the labs, whose economics depend on capacity they do not control. It is unflattering to the hyperscalers, who are rationing customers while announcing record investment. It is unflattering to European policy, which has spent three years discussing sovereignty in terms of data location while the binding constraint turned out to be high-bandwidth memory supply across three vendors. And it is unflattering to enterprises, whose AI budgets were modelled on a token price that is being replaced by a compute-time price.

What is genuinely encouraging is that the two best-informed and best-capitalised companies in the industry independently reached the same conclusion about what to do: do not bet the architecture on one model, keep control of the inference layer, run locally what can run locally, and abstract the provider so that substitution stays cheap. Neither of them is doing this out of principle. They are doing it because it is the correct engineering response to a constraint they can see more clearly than anyone.

European enterprises should draw the same conclusion and act on it earlier, because they have less slack. And they should notice that the one compute layer they own outright — the fleet of machines already sitting on their employees' desks, inside their own perimeter, under their own jurisdiction, immune to rationing and to export control — is currently doing nothing at all for them.

A brief note on the regulatory backdrop, since it continues to develop. The EU AI Act became broadly applicable on 2 August 2026, with GPAI enforcement powers under Chapter V binding from that date; the Digital Omnibus agreement of 7 May 2026 postponed the high-risk Annex III obligations to 2 December 2027 and Annex I obligations to 2 August 2028. [¹¹] Worth noting in the context of this piece: an architecture that can move inference between cloud, on-premises and endpoint under policy control is considerably easier to keep compliant across a divergent regulatory landscape than one whose inference location is a property of a vendor contract.

Compute is the constraint. The endpoint is the tier nobody is pricing. That is the work in front of us, and it is the work we have been doing.


¹ Series articles at neuland.ai/en/resources/insights.

² See earlier pieces in this series on the fast-follower workhorse thesis, on flexibility as an architectural property following the withdrawal of a frontier model by export-control directive, on the open-weight frontier, and on the commoditisation of sovereignty of location.

³ Reported June 2026: Google limited Meta's access to Gemini models owing to compute capacity constraints. Google 2026 capital expenditure guidance of 180 to 190 billion US dollars, with expectation of significant increase in 2027; equity raise of approximately 84.75 billion US dollars to fund AI compute infrastructure; data centre capacity lease from xAI at approximately 920 million US dollars per month. Anthropic TPU commitment: up to one million Ironwood chips, five gigawatts of capacity over five years, valued up to 40 billion US dollars, with additional supply lined up through 2027; separate TPU agreement with Meta earlier in 2026. Internal compute queueing affecting DeepMind research timelines reported mid-2026, including departures of senior long-tenured researchers to better-resourced startups. Demis Hassabis public acknowledgement of the constraint, attributing it to a small number of suppliers of key components — the high-bandwidth memory bottleneck across Samsung, Micron and SK Hynix.

⁴ Apple–Google agreement announced 12 January 2026 for a custom Gemini variant, estimated at approximately one billion US dollars annually. Internal assessment reportedly placed Apple's own large language models twelve to eighteen months behind leading alternatives on conversational quality, reasoning and multi-step task completion. For context, Apple pays a substantially larger annual sum for default search placement.

⁵ Apple WWDC, 8 June 2026: rebuilt assistant announced on new Apple Foundation Models with a licensed frontier cloud tier for the most demanding reasoning. Architecture retains an on-device model of approximately three billion parameters for latency-sensitive and privacy-sensitive tasks, with a controlled cloud compute environment as the intermediate tier. Own next-generation multimodal architecture and trillion-parameter proprietary models under development with the intention of displacing the licensed model from 2027; own AI server silicon in development. Analyst characterisation of the partnership as easing short-term pressure rather than representing a long-term strategic shift, with on-device AI demand expected to grow more meaningfully from 2027.

⁶ Apple's foundation model framework opened beyond Apple's own models at WWDC 2026 via a new public protocol, permitting substitution of licensed frontier, competitor, or on-device models without application rewrites. Developer commentary characterised provider abstraction as having become table stakes, and single-provider coupling as emerging technical debt. Current-generation Apple silicon features redesigned neural accelerators in every core with approximately 40 percent higher AI throughput than the preceding generation.

⁷ Compute-based usage limits introduced on Google's consumer AI tiers in May 2026, with allocation consumed according to actual computational processing required rather than request count, and removal of previously bundled monthly AI credits from the paid tier. Separate agent runtime billing introduced on the enterprise platform.

⁸ Google TPU access is available exclusively through Google Cloud Platform, with no on-premises deployment path. Organisations with data sovereignty requirements, air-gapped networks or policies precluding public cloud cannot use the hardware regardless of its technical characteristics. The forthcoming inference-optimised generation is stated to offer approximately 80 percent improvement in performance per dollar for low-latency inference on large mixture-of-experts models.

⁹ Gartner projection: approximately 40 percent of enterprise applications will ship with task-specific AI agents by the end of 2026, up from under 5 percent a year earlier, with the 2026 agentic AI Hype Cycle describing a market moving from experimentation into production readiness as governance and orchestration layers arrive.

¹⁰ neuland.ai HUB: model-agnostic routing per workload across frontier proprietary and self-hosted open-weight models, decided on capability, cost, residency and policy; permission-aware retrieval enforced ahead of reasoning; customer-owned ontology and knowledge layer; append-only audit across retrieval and access decisions; deployment options spanning sovereign cloud, private cloud, on-premises and air-gapped operation. Bring-your-own-harness support and durable streaming for long-running agentic work are current engineering priorities; endpoint-tier inference routing is a stated architectural direction rather than a shipped capability. neuland.ai AG retains responsibility for content quality and clean delivery of results across all customer engagements.

¹¹ Council of the EU and European Parliament provisional political agreement on the Digital Omnibus on AI, 7 May 2026: Annex III high-risk obligations postponed to 2 December 2027; Annex I obligations postponed to 2 August 2028; Article 50(2) watermarking obligations moved to 2 December 2026. GPAI enforcement powers under Chapter V binding from 2 August 2026.


Image generated using the neuland.ai HUB.