Einblick in die hochmoderne Fertigung der CiS electronic GmbH in Krefeld: KI-gestützte Prozesse steigern Effizienz, Qualität und Innovation in der Kabelkonfektion. © CiS electronic GmbH

Research

AI Agents

Distillation is the strategic destination

Article by

Dr. Anoj Winston Gladius

·

The question is who is distilling into whom, and on whose substrate.

On 23 July 2026, Michael Kratsios, who heads the White House Office of Science and Technology Policy, publicly accused Moonshot AI of using large-scale distillation against Anthropic's Claude Fable 5 to train Kimi K3, alleging that the company had built an internal platform to run the extraction while switching between access methods to avoid detection. Anthropic had earlier alleged that approximately 25,000 fraudulent accounts ran 28.8 million exchanges through Claude between 22 April and 5 June 2026, targeting coding, reasoning and cybersecurity capability specifically. In February 2026 the company had made a similar allegation against DeepSeek, Moonshot and MiniMax collectively. Kratsios presented no direct evidence that Fable 5 outputs had been incorporated into K3, and AI researchers publicly disagreed about whether the timeline made it technically feasible. Others noted the awkwardness of the accusation arriving from a company that had recently settled claims over training on pirated books for 1.5 billion US dollars, and pointed to developer reports of Claude Opus 4.8 identifying itself as Alibaba's Qwen during language tests. The Oxford China Policy Lab has documented that Chinese developers routinely buy proxied Claude access through resale channels at roughly a tenth of list price. Whether the specific allegation is accurate is a question for regulators, courts and the labs involved. What matters for European enterprise architecture is different, and almost nobody is saying it: the technique at the centre of this controversy is the same technique that will define enterprise AI architecture over the next twenty-four months. The controversy is about who is distilling into whom, and on whose substrate. Those are the right questions. They just have a completely different answer at the enterprise layer than they do in a fight between frontier labs.

This is the fourteenth piece in a series I have been writing for neuland.ai . [¹] The thread running through every piece is the same: in enterprise AI, the value, the risk and the moat sit in the layer above and around the model, not in the model itself. The previous piece argued that Kimi K3's release on 16 July 2026 was the empirical resolution of the commoditisation thesis — that the open-weight frontier and the closed-weight frontier now overlap on most enterprise-relevant benchmarks. [²] This piece is about what enterprises do with that condition, and it turns on a technique that a political controversy has just made highly visible for the wrong reasons.

What distillation actually is

Distillation is a well-established machine learning technique with a fifteen-year history and no inherent controversy attached to it. In its general form, a smaller model is trained on the outputs of a larger one, learning to approximate the larger model's behaviour at a fraction of the parameter count, inference cost and deployment footprint. Every major lab uses it internally. It is how frontier models get turned into the cheaper, faster variants that appear alongside them in the same product family. It is also how a general-purpose model gets turned into something narrower and better at a specific class of task.

It is worth being precise about what the technique can and cannot deliver, because this constrains the argument in both directions. Supervised fine-tuning on another model's outputs transfers style, formatting conventions, response structure and — as the identity-confusion reports illustrate — sometimes the source model's own sense of what it is. What supervised fine-tuning does not straightforwardly transfer is frontier-level reasoning capability. Nathan Lambert's technical assessment of the current controversy, reported by TechCrunch, makes this point directly: achieving genuinely frontier-adjacent capability through distillation requires reinforcement learning techniques at substantial scale, with an agent of the larger model grading the smaller model's outputs across very large numbers of runs. That is a different infrastructure problem from generating a few tens of millions of API exchanges, and it is one reason researchers have been sceptical about whether the alleged mechanism could produce the observed result. [³]

The distinction matters for the enterprise argument, because it tells you what distillation is actually good for. It is not a shortcut to owning a frontier model. It is an extremely effective method for producing a smaller, cheaper, faster model that is very good at a specific bounded domain — provided you have the domain data and the grounding to define what "good at this domain" means.

The lab-to-lab version, and why it is contested

The version of distillation currently generating political heat has a specific shape. One lab uses another lab's commercial API, at industrial volume, through accounts and access paths that circumvent the terms under which that API is offered, for the purpose of building a competing product. Anthropic's allegation describes exactly this, with specific figures: roughly 25,000 accounts, 28.8 million exchanges, a six-week window from 22 April to 5 June 2026, and capability targets in coding, reasoning and cybersecurity. Kratsios's public accusation on 23 July extended the framing to state actor involvement and hardware acquisition routes, and named Kimi K3 as the output. [⁴]

The counter-arguments have been made publicly and are worth acknowledging rather than dismissing. No direct evidence linking Fable 5 outputs to K3's training corpus has been presented. Researchers have questioned whether the alleged volume is sufficient for the alleged result. Commentators have noted that a company that recently settled copyright claims over its own training data for 1.5 billion dollars occupies an uncomfortable position when arguing that unauthorised training on others' outputs is illegitimate. Developer reports of Claude Opus 4.8 self-identifying as Qwen suggest that whatever the direction of influence in the current market, it is unlikely to run only one way. [⁵]

I am not going to adjudicate this. It is a genuine dispute with real commercial stakes, live regulatory attention, and evidentiary questions I have no privileged access to. What I want to extract from it is the observation underneath: the reason this fight is happening at all is that distillation works well enough to be worth fighting over. Two of the most sophisticated technical organisations in the world, plus the White House, are treating the ability to compress capability out of a large model into a smaller one as strategically consequential. They are right that it is. They are fighting about it at the wrong layer of the stack for European enterprises to learn anything useful from the fight itself.

The version that matters for European enterprises

Move the same technique down a layer and the entire picture changes.

An enterprise does not need to distill a frontier model into a general-purpose competitor. It has no commercial interest in doing so, no capacity to do so, and nothing to gain from trying. What an enterprise needs is something narrower and considerably more valuable to it: a model that is very good at the specific class of reasoning its own business requires, running at a cost and latency profile that makes high-volume deployment economically sensible, on infrastructure it controls, grounded in its own operational reality.

That is a distillation problem. And it is a distillation problem with none of the properties that make the lab-to-lab version contested. The source models are ones the enterprise is licensed to use, accessed under the terms they are offered under, through the enterprise's own routing table. The training signal comes from the enterprise's own data, its own documents, its own historical decisions, its own domain experts' judgements about what a correct answer looks like. The output is a customer-specific small language model that the customer owns outright, deploys on the substrate it chose, and can retrain or discard as its requirements change. No terms are violated. No IP dispute arises. No jurisdictional exposure is created that did not already exist.

The economics are the part that makes this strategically dominant rather than merely interesting. A frontier model called through a commercial API on every document in a high-volume enterprise workflow is expensive at a per-token level and unpredictable at a budget level — which is why Uber exhausted an annual AI budget in four months and why per-action metering has become the standard monetisation pattern across the major enterprise software vendors.

The frontier labs are aware of this and are responding to it in their own pricing. Anthropic released Claude Opus 5 on 24 July 2026 at half the token price of Fable 5, positioning it explicitly as the model to reach for by default in everyday production work rather than the one that chases benchmark records, and reporting efficiency gains that one early-access customer measured as a twenty-six percent reduction in average token consumption at equivalent output quality. The company's head of product management for research characterised enterprise buyers as looking for value, in a market where they are increasingly unwilling to experiment with expensive models absent a clear picture of the return. [⁶] That is a genuine concession to the cost dynamics described above, and it materially narrows the gap between frontier pricing and open-weight pricing. It does not change the underlying arithmetic for a customer running a genuinely high-volume workload, because a lower per-token price on a general-purpose model is still a per-token price on a general-purpose model — metered by a third party, on a rate card the customer does not set, for a model that has never seen the customer's business. A domain-tuned small model running on the customer's own GPU capacity costs the GPU minutes it consumes. For the workloads that constitute the bulk of enterprise AI traffic — classification, extraction, structured summarisation, routing, compliance checking, first-pass drafting, retrieval grounding — a well-distilled domain model does not merely match frontier performance at lower cost. It frequently exceeds it, because it has been trained on the specific vocabulary, document structures, exception patterns and decision rules of that particular business, which no general-purpose frontier model has ever seen.

Why the ontology is the difference between a small model and a domain model

This is where the argument connects back to the piece earlier in this series on the ontology race. [⁷]

Distillation without a knowledge substrate produces a smaller general-purpose model. Useful, cheaper, faster, and still fundamentally generic. What turns a distilled model into a domain reasoning engine is the ontology: the formal representation of the entities, relationships, decision rules and permitted actions that constitute how a specific business actually operates. Customers, suppliers, contracts, assets, products, approval thresholds, exception paths, regulatory obligations, industry-specific terminology that means something different inside this company than it does in general usage.

With the ontology in place, several things become possible that are not possible without it. The training signal for distillation can be constructed against ground truth rather than against generic quality heuristics — the ontology defines what a correct entity extraction is, what a valid relationship inference is, what an acceptable compliance determination looks like in this specific regulated context. The evaluation of the distilled model can be rigorous rather than impressionistic, because there is a formal definition of correctness to evaluate against. And the resulting model's outputs remain auditable and explainable, because they are grounded in a representation that a human domain expert can inspect and reason about.

The practical consequence is that the ontology and the domain-specific model are not two separate initiatives. They are two halves of one architecture. An enterprise that has invested in a customer-owned ontology has, as a direct consequence, the substrate required to produce customer-specific models that are genuinely better than general-purpose alternatives on its own workloads. An enterprise that has not has, at best, access to a cheaper generic model. This is why the ontology consolidation race described earlier in this series matters as much as it does, and it is why the question of whose ontology it is turns out to determine whose models the enterprise can build.

The industry pattern generalises across the verticals where this work is most valuable. A legal services organisation with a formal representation of its contract taxonomy, clause library and jurisdictional rules can produce a model that reasons about its contracts more reliably than any frontier model. A financial institution with a formal representation of its product structures, regulatory obligations and risk categories can produce a model that handles its compliance workflows with an accuracy profile a general model cannot reach. A manufacturer with a formal representation of its component hierarchies, quality standards and supplier relationships can produce a model that reads its technical documentation the way its own engineers do. In each case the ontology is the asset and the distilled model is the capability the asset makes possible.

Where neuland.ai stands

The neuland.ai HUB is architected around this pattern rather than adapted to it after the fact. The routing table gives the customer access to frontier proprietary and open-weight models across providers and jurisdictions, with per-workload evaluation determining what runs where. The ontology layer generates and maintains the company ontology and per-connector ontologies from the customer's own data, materialised as a knowledge graph with skills, rules and vector indexes, and continuously kept in sync as the underlying data changes. The orchestrator routes each task to the appropriate ontology and the appropriate model per domain and application, deterministic or agentic as the workload requires. The governance, observability and audit layers operate uniformly across all of it. Ingestion runs at petabyte scale on customer-controlled infrastructure. Deployment is sovereign by default on STACKIT, with on-premises, private cloud and hybrid options available at the customer's discretion. [⁸]

The distillation and domain-adaptation work sits naturally on top of that foundation, and it is an active research focus. Our research team is working on adaptation and distillation techniques for producing customer-specific and industry-specific small language models grounded against customer ontologies, with the evaluation discipline that regulated deployment requires — rigorous measurement against real customer workloads rather than public benchmarks, explicit tracking of where the domain model outperforms the general model and where it does not, and honest reporting of the boundary between them. The delivery side of the business is what identifies which workloads in a specific customer environment justify the investment, because that determination requires understanding the customer's actual processes rather than reading a capability matrix. This is the direction of travel described in an earlier piece in this series: the platform capability exists because the engagements required it, not the other way around. [⁹]

Personal take

The distillation controversy is going to run for months. It involves national security framing, export control implications, IP law questions that have no settled answers, and commercial interests large enough to guarantee that all parties will litigate the narrative aggressively. European enterprises have no useful role in that fight and nothing to gain from taking a position in it.

What European enterprises should take from it is the underlying signal. The most sophisticated technical organisations in the world are treating the ability to compress capability out of large models into smaller, cheaper, more specialised ones as strategically decisive. They are correct. The technique works. What they are fighting about is a specific application of it — one lab replicating another lab's general-purpose frontier model without licence — that is both contested and, for an enterprise, entirely beside the point.

The application that matters for a European enterprise is the one where the enterprise distills from models it is licensed to use, into models it owns, grounded on an ontology it controls, deployed on substrate it chose, tuned for workloads only it has. That version raises no IP question, creates no jurisdictional exposure it did not already have, and produces something no frontier lab can sell it: a model that knows its business. Over the next twenty-four months, the enterprises that have built the ontology and the platform discipline to do this will be operating at a cost and capability profile that enterprises still calling frontier APIs on every document simply will not be able to match.

Distillation is the strategic destination. The controversy is about who is distilling into whom, and on whose substrate. For a European enterprise the answers should be: from the models on your routing table, into your own domain models, on your own infrastructure, grounded on your own ontology. That is the version nobody will accuse you of, and it is the version that produces durable advantage.

A brief note on the regulatory backdrop, since it continues to develop. The 7 May 2026 EU Digital Omnibus agreement postponed the high-risk Annex III obligations from 2 August 2026 to 2 December 2027, and Annex I obligations to 2 August 2028. [¹⁰] GPAI enforcement powers under Chapter V remain on the original 2 August 2026 schedule. The strategic implication is unchanged from the previous pieces in this series: the architecture decisions of Q3 and Q4 2026 are the ones that determine whether the European enterprise AI stack survives the 2027 procurement cycles intact.

Distillation is the strategic destination. The question is whose substrate you arrive on. That is the work in front of us, and it is the work we have been doing.

¹ Series articles at Insights from neuland.ai: AI in practice, trends & deep dives . Previous pieces have addressed: control planes and execution surfaces; model drift, multi-LLM strategy and observability; model topology and hyperscaler independence; compliance as a system property; agent security and the lethal trifecta; MCP protocol governance; the fast-follower workhorse thesis with sovereign deployment; the new enterprise data silos created by SAP / Microsoft / ServiceNow / Salesforce AI gateways; flexibility as the architecture in light of the Fable 5 recall; the ontology race started by SAP / Microsoft / Google / Palantir; the case against the visual workflow builder paradigm; the direction of travel from which an enterprise AI services partner is built; and the open-weight frontier following Kimi K3.

² For the commoditisation argument and Kimi K3's benchmark position, see the previous piece in this series. K3 released 16 July 2026: 2.8 trillion parameters, sparse Mixture-of-Experts, 1 million token context, Modified MIT licence, with full weights scheduled for 27 July 2026.

³ Nathan Lambert's technical assessment of the distillation allegations, as reported in TechCrunch, 23 July 2026: supervised fine-tuning transfers surface characteristics including response style and model self-identification, but achieving frontier-adjacent capability through distillation requires reinforcement learning approaches operating at substantially larger infrastructure scale than API exchange volume alone would indicate.

⁴ Michael Kratsios, Director of the White House Office of Science and Technology Policy, public statement on X, 23 July 2026, alleging large-scale distillation by Moonshot AI against Claude Fable 5 in the development of Kimi K3, including allegations regarding access-method rotation to avoid detection and hardware acquisition routes. Anthropic's underlying allegation: approximately 25,000 fraudulent accounts generating 28.8 million exchanges between 22 April and 5 June 2026, with stated capability targets in coding, reasoning and cybersecurity. Earlier allegation, February 2026, named DeepSeek, Moonshot and MiniMax collectively with approximately 24,000 accounts and 16 million exchanges. Reporting: CyberScoop, The New Stack, South China Morning Post, Tom's Hardware, 22–24 July 2026. Alibaba has denied the related allegations concerning Qwen.

⁵ On the absence of direct evidence and researcher disagreement regarding technical feasibility, see South China Morning Post and TechCrunch coverage, 23 July 2026. On Anthropic's 1.5 billion US dollar settlement of claims relating to training data, and on developer reports of Claude Opus 4.8 self-identifying as Qwen, see The Register, 22 July 2026, and TipRanks, 22 July 2026. On proxied commercial API resale channels operating in the Chinese developer market at approximately one-tenth of list price, see Oxford China Policy Lab research as reported in The Next Web, 21 June 2026. These are reported positions in a live commercial dispute; this piece takes no position on the underlying factual questions.

⁶ Anthropic, "Introducing Claude Opus 5," 24 July 2026. Pricing: 5 USD per million input tokens and 25 USD per million output tokens, unchanged from Opus 4.8 and half the 10 USD / 50 USD of Fable 5. Anthropic positions the model as coming close to Fable 5's frontier intelligence at half the price, designed for everyday use rather than record benchmark scores, and reports new state-of-the-art results on Frontier-Bench (43.3 percent, against Fable 5 at 33.7 percent and Opus 4.8 at 18.7 percent) and GDPval-AA v2 (1,861, ahead of both Fable 5 and GPT-5.6 Sol), while remaining behind Mythos 5 on cybersecurity tasks. A tunable effort setting and a faster mode (approximately 2.5x default output speed at twice the base rate) are included. Efficiency measurement cited: Niko Grupen, head of applied research at Harvey, reported output quality matching Opus 4.8 at maximum reasoning with a 26 percent reduction in average token usage. Enterprise value framing: Dianne Penn, Anthropic head of product management for research, in comments to CNBC, 24 July 2026, on enterprise buyers seeking value amid reluctance to commit to expensive models without clear return. On comparative per-task economics, Artificial Analysis weighted average cost per Intelligence Index task at the time of the Opus 5 launch: 2.75 USD for Fable 5, 2.03 USD for Opus 5 at maximum effort, 1.04 USD for GPT-5.6 Sol at maximum, and 0.95 USD for Kimi K3. Reporting: The Register, CNBC, Quartz, Techzine, 24–25 July 2026. Anthropic has published no information on training-lineage relationships between models in the Opus and Mythos-class families; the documented shared-model relationship is between Fable 5 and Mythos 5, which differ in safeguard configuration rather than in underlying model.

⁷ For the ontology consolidation argument and the SAP Knowledge Graph / Microsoft Fabric IQ / Google Enterprise Knowledge Graph / Palantir Foundry Ontology positioning, see the earlier piece in this series.

neuland.ai HUB platform architecture. Core elements: orchestrator routing per ontology and SLM per domain / application (deterministic or agentic); agentic retrieval and reasoning (agentic RAG, recursive language models, agentic threads, deep agents, multi-hop reasoning, KAG, query planning, semantic re-ranking); ontology layer covering company ontology and connector ontologies, auto-generated from customer data and continuously updated, materialised as knowledge graph, skills, rules and vector index; governance and compliance (roles and rights, guardrails, audit trail, EU AI Act / DORA / BaFin / DSGVO / BRAO, explainability via AtMan and Aleph Alpha); observability and logging across all connectors and agentic processes; petabyte-scale ingestion pipelines on Kubernetes GPU clusters; sovereign deployment on STACKIT with on-premises, private cloud and hybrid options; model-agnostic model layer with per-workload evaluation harness driving the routing table. neuland.ai AG retains responsibility for content quality and clean delivery of results across all customer engagements.

⁹ For the direction-of-travel argument regarding how the platform capability emerged from customer engagements rather than being designed in advance of them, see the earlier piece in this series on the services layer.

¹⁰ Council of the EU and European Parliament provisional political agreement on the Digital Omnibus on AI, 7 May 2026. Annex III high-risk obligations postponed from 2 August 2026 to 2 December 2027 (16-month delay); Annex I obligations postponed to 2 August 2028 (12-month delay); Article 50(2) watermarking moved to 2 December 2026. GPAI enforcement powers under Chapter V remain on the original 2 August 2026 schedule.


Image generated using the neuland.ai HUB.