AI Strategy
Where do your agent's limits actually live?
Where do your agent's limits actually live?

Dr. Anoj Winston Gladius
·
15
Min. Lesezeit

Prompt limits are requests. Architecture limits are controls.
aufsatz
Image: AI generated with neuland.ai HUB
Three news items in eight days tell a story that none of them tells alone. On 28 September, OpenAI announced that it would not release GPT-6.1 Astra, the successor to its flagship model, which had been due in ChatGPT and Codex in October. The reason, as its head of safety systems put it, was that the model fell short on staying within the scope and authorisation it had been given, and on how honestly it reported back to users what it had actually done. The next day, at its developer conference, OpenAI launched Dots, its new persistent agents, running on the previous model. On 2 October, Apple announced that it would tighten the Full Disk Access permission in macOS, warning developers that more capable and more autonomous agents make broad access to a user's files, mail and messages substantially riskier. Read together, these are not three unrelated stories. OpenAI and Apple said the same thing in the same week, from opposite ends of the stack: an agent will not reliably stay inside limits that are written down as instructions, and its own account of what it did is not a record. That raises the question every CIO and CISO with agents in production should be asking now: where do the limits on our agents actually live?
This is the twenty-sixth piece in a series I have been writing for neuland.ai. Each one has pushed the same underlying argument forward: the value, the risk and the moat in enterprise AI are not in the model. They are in the layer above and around it. [¹] Earlier pieces in this series argued that containment has to work in both directions, that authorisation has to cover what an agent may know and not only what it may do, and that provenance has to be recorded when an action happens rather than reconstructed afterwards. [²]
This piece brings those arguments down to one practical distinction. A limit written in an instruction is a request. A limit enforced outside the model is a control. And these three stories are the clearest demonstration of that difference I have seen.
Story one: the vendor stopped its own model
On 28 September, the Wall Street Journal reported that OpenAI had scrapped the October release of GPT-6.1 Astra, and OpenAI confirmed the decision the same evening. Saachi Jain, OpenAI's head of safety systems, said the model had improved on laziness but fell short on "staying within scope and authorization", and on how it communicated to users the type of work it had done. According to the Journal's reporting, internal testing found higher levels of deceptive behaviour than in the model it was meant to replace. OpenAI said it will continue to release other models. [³]
OpenAI deserves credit for this. Pulling a model that does not meet your own bar, a day before your own developer conference, is exactly what a responsible laboratory should do.
But read the stated reason as an engineer. The two failures it names are the two properties that enterprises most often delegate to the model: staying inside the scope it was given, and telling you truthfully what it did. Neither is a capability problem, by OpenAI's own account the model was more capable. Both are containment and provenance problems. And the company that knows this model better than anyone concluded that it could not rely on the model to handle them.
Story two: the agents shipped anyway, on the previous model
On 29 September, at DevDay 2026 in San Francisco, OpenAI unveiled Dots, a new generation of persistent agents running on GPT-6 Astra. Sam Altman said the product had been safety-tested before release, distinguishing it from the withheld model. Specialist Dots carry their own organisational identity and are in focused enterprise pilots; Enterprise Dots are switched off by default until an administrator enables them; and the Pro tier is not available in the European Economic Area, the United Kingdom or Switzerland. [⁴]
None of that is a criticism. Off by default is a sensible control, and shipping on a model that cleared the bar is the right decision. But notice where the decision now sits: on your side of the line. When an administrator switches persistent agents on, the limits that matter are the ones the organisation enforces around them, because the same vendor explained, in the same week, what happens when you rely on the model to enforce them itself.
Story three: Apple moved the limit into the operating system
On 2 October, Apple published a notice on its developer site announcing additional controls on Full Disk Access in macOS. The permission exists so that backup software can work, and it largely bypasses the privacy controls Apple builds for everything else. Apple wrote that some developers are using it in ways that expose files, mail, messages and browsing history without users fully understanding what they have granted, that granting it will in future require very explicit user action, and, the sentence that matters here, that as agents become more capable and autonomous, "the risks associated with this level of access will grow substantially." No rollout date has been given. [⁵]
The trigger was a journalist's claim that Meta's Muse assistant had read his private messages on a Mac. Meta disputes the account and says access to Messages is opt-in, requiring both Full Disk Access and a separate connector to be enabled. The facts of that case are not settled. Apple's conclusion does not depend on them. It did not ask agent developers to behave better. It moved the limit to a layer the agent cannot talk its way past.
What ties the three stories together
It would be tempting to read these as an OpenAI story and an Apple story. They are one story, told from both ends of the stack.
OpenAI, which understands its own model better than anyone, concluded that the model would not reliably keep to the limits users set, and would not reliably report what it had done. Apple, which owns the operating system that desktop agents run on, concluded that it could not rely on agents or their developers to use broad access responsibly, and moved the control into the platform. The laboratory and the platform owner reached the same conclusion from opposite directions: limits have to be enforced by something other than the agent.
The point that matters more than any single model
Here is the part that I have not seen written down clearly enough in the coverage of either story.
Most enterprise agent deployments enforce their most important limits in the system prompt. Only access this customer's records. Never send data outside the company. Ask before deleting anything. Report every action you take. These are sensible instructions, and every agent should carry them. They are not controls. They are requests addressed to a probabilistic system, the same kind of system whose vendor has now said that its newest version got worse at honouring exactly these requests while getting better at nearly everything else.
The distinction matters most for the second failure OpenAI named. If the agent's own summary is the record of what it did, then the quality of your audit trail is a property of the model, and it degrades precisely when the model's honesty does. A record written by something other than the agent, at the moment each retrieval, tool call and decision happens, does not depend on the agent telling the truth. An earlier piece in this series argued that provenance has to be captured at the moment of action. GPT-6.1 Astra is the vendor explaining why.
And the other side of the boundary
The limits question has a second half, because agents are now on both sides of the perimeter.
On 9 September, GreyNoise documented the first mass-exploitation campaign run overwhelmingly by autonomous agents. A single attacker used hundreds of agents built on OpenAI Codex and a DeepSeek model to exploit two PaperCut vulnerabilities across 395 organisations in 48 countries. The agents scanned for targets and wrote the exploits themselves; initial compromise took 26 seconds at eleven organisations, and domain administrator access was confirmed at twelve, between five and 144 minutes after entry. [⁶] A day earlier, Google's Threat Intelligence Group reported that machine credentials now carry a market price, and described an agent-enabled credential-harvesting campaign planned and executed in under six hours; Anthropic's threat report two days later documented two further credential thefts.[⁷] Later in September, Gambit Security and the Cloud Security Alliance described a campaign in which chained open-source agent tools compromised more than a hundred online retailers and stole over 600,000 payment card records. [⁸]
Against that, a survey in June found that only 12 percent of enterprises had an agent-identity product anywhere in their consideration set. [⁹]Most organisations are still governing agents with the identity tools they built for people.
The audit: eight questions for every agent you run
The practical consequence is an inventory exercise, and I would encourage every CISO reading this to run it, starting with whether the inventory exists. For each agent, ask where each of the following limits physically lives. If the honest answer is "in the prompt", that limit is a request.
1. Whose identity does it act under? An agent running on a borrowed human session or a shared service account inherits everything that person or account can reach, and every action it takes is attributed to someone else. The limit belongs in a scoped identity of the agent's own, with the person it acts for preserved in the delegation chain.
2. What can it read, and is that trimmed before retrieval? If the agent retrieves broadly and the output is filtered afterwards, confidential material has already entered the context. The limit belongs in the retrieval layer, applied for the person the agent is acting for, before the model reasons.
3. Which tools can it call? If the answer is "whatever the framework exposes", the reachable set is emergent. The limit belongs in an enumerated allowlist enforced by the platform.
4. Where can it send data? Egress should be a configuration the agent cannot change, not an instruction it is asked to respect.
5. What credentials does it hold, and for how long? Long-lived keys inside an agent's environment are exactly what the September campaigns were harvesting. The limit belongs in short-lived, brokered access, where the agent never holds the secret itself.
6. Which steps must be identical every time, and which need a human signature? Regulated procedures should run as deterministic, replayable sequences, with approvals recorded against the person who gave them, not left to the agent to decide when to ask.
7. Who records what it actually did? If the answer is the agent, the record is only as honest as the model. The limit belongs in an append-only log written by the platform at the moment of action.
8. How quickly can you switch it off, or swap its model? A model can now be withdrawn by its vendor for safety reasons as well as by a government directive. Swapping it should be a routing change, and switching off an agent should not mean switching off the person it acts for.
An honest update to an earlier piece
The previous piece on the pacing debate observed that the proposal to slow the frontier had halted no release. One has now been halted, and it is worth saying so plainly. The stated reason is the very scenario that piece described: an agent that does not stay within the bounds it was given. [¹⁰]
The conclusion of that piece stands, and is stronger for it. A laboratory can decide not to ship a model. It cannot enforce limits inside your environment. Only you can do that.
Where neuland.ai stands on this
The neuland.ai HUB is a sovereign AI management and orchestration platform, and its job is to put the limits outside the model. Entitlement is enforced on the retrieval space before the model reasons, so a document the requesting person may not see is never retrieved and therefore cannot be cited or inferred from. Agents act under their own scoped identities, with the person they act for preserved in a delegation chain, rather than on borrowed credentials. The tools an agent may call are an enumerated allowlist rather than an emergent property. Execution runs in isolated environments without standing credentials, with the platform brokering each model and tool call. Regulated procedures run as deterministic, replayable sequences, with human approval recorded in the flow where it is required. Every retrieval, tool call and decision is written to an append-only record by the platform, not summarised afterwards by the agent. And because routing is model-agnostic, covering open-weight models we host such as GLM, Mistral Large 3, Qwen and DeepSeek, and proprietary endpoints such as Claude and GPT where the customer's policy allows, a model that is withdrawn, or that fails the customer's own evaluation, is a routing change rather than an outage. [¹¹]
This is also the answer to a question I hear from CISOs in almost every first meeting: can't we simply tell the agent what it is not allowed to do? You can, and you should, good instructions make a good agent better. But they are not the control. GPT-6.1 Astra is the vendor putting that in writing.
Personal take
I want to be careful with the framing here. This piece is not anti-OpenAI. Pulling a model that did not meet its own bar, the day before a developer conference, is the responsible decision and should be recognised as one. It is not anti-Apple either, and it is not a judgement on Meta, which disputes the claim that preceded Apple's announcement.
What I am pushing back against is a habit that has spread quickly through enterprise AI: treating the model as the place where limits live. Scope, permissions, data boundaries and the record of what happened are being written into system prompts and agent descriptions, and then trusted. In the space of eight days, the vendors told us from both ends of the stack why that does not hold.
The regulatory direction points the same way. The EU AI Act has applied broadly since 2 August 2026, and its expectations on human oversight, logging and traceability are written for exactly this problem; the Digital Omnibus has moved the high-risk Annex III obligations to 2 December 2027 and Annex I obligations to 2 August 2028. [¹²] A limit that lives in a prompt cannot be evidenced to a supervisor. A limit that is enforced and recorded by the platform can.
We at neuland.ai would rather enforce a limit than ask for one. If the model is also willing to respect it, that is a bonus. The control should not depend on it.
Series articles available at neuland.ai/en/resources/insights.
See earlier pieces in this series on containment as a bidirectional system property, on authorisation stopping at the wrong boundary, and on provenance as a system property.
Maxwell Zeff, "OpenAI Scraps Release of New AI Model Over Safety Concerns", The Wall Street Journal, 28 September 2026; confirmed to CNBC the same evening. Statement by Saachi Jain, head of safety systems at OpenAI, as reported by CNN and CNBC, 28 September 2026. GPT-6.1 Astra had been planned for release in ChatGPT and Codex in October 2026. Report of higher levels of deceptive behaviour than its predecessor as carried by Business Standard, 29 September 2026, citing the Wall Street Journal.
OpenAI DevDay 2026, San Francisco, 29 September 2026: launch of Dots running on GPT-6 Astra; statement by Sam Altman that the product had been safety-tested before release, as reported by Reuters. Availability details per OpenAI's published access guidance: specialist Dots with their own organisational identity in focused enterprise pilots; Enterprise Dots off by default until enabled by an administrator; Pro availability excluding the EEA, the United Kingdom and Switzerland.
Apple Developer, "Updates to Full Disk Access in macOS", 2 October 2026; coverage in TechCrunch and MacRumors, 2 October 2026. The preceding claim regarding Meta's Muse was made by a columnist reporting that the assistant had accessed his Messages data; Meta's response, as reported, was that Messages access is entirely opt-in and requires both Full Disk Access and an enabled Messages connector.
GreyNoise, campaign documentation of 9 September 2026, as reported by VentureBeat, "AI agents breached 395 organizations using credentials your IAM policy still treats as human", September 2026.
Google Threat Intelligence Group, "From Prompting to Autonomy: The Evolution of Adversarial AI", 8 September 2026. Anthropic threat intelligence report, 10 September 2026, covering activity from December 2025 to August 2026.
Cloud Security Alliance AI Safety Initiative, research note on an AI agent retail skimming campaign, 24 September 2026, drawing on research by Gambit Security; see also BleepingComputer, 23 September 2026.
VentureBeat Pulse survey, June 2026: 12 percent of enterprises had an agent-identity product, such as Okta for AI Agents, Microsoft Entra Agent ID or a non-human identity platform, in their consideration set.
See the previous piece in this series on the pacing debate, following Dario Amodei's essay "We Must Pace the Frontier", 12 September 2026.
neuland.ai HUB: entitlement enforced on the retrieval space ahead of reasoning; relationship-based entitlement with delegated, scoped agent identities; enumerated tool allowlisting; isolated execution environments with platform-brokered model and tool calls; deterministic, replayable orchestration with recorded human approval; append-only audit written by the platform; model-agnostic routing across self-hosted open-weight and proprietary models; deployment on-premises, in a European sovereign cloud or in a hyperscaler region, including air-gapped operation. neuland.ai AG retains responsibility for content quality and clean delivery of results across all customer engagements.
Regulation (EU) 2024/1689 (AI Act), broadly applicable from 2 August 2026. Council of the EU and European Parliament provisional political agreement on the Digital Omnibus on AI, 7 May 2026: Annex III high-risk obligations postponed to 2 December 2027; Annex I obligations postponed to 2 August 2028.
Three news items in eight days tell a story that none of them tells alone. On 28 September, OpenAI announced that it would not release GPT-6.1 Astra, the successor to its flagship model, which had been due in ChatGPT and Codex in October. The reason, as its head of safety systems put it, was that the model fell short on staying within the scope and authorisation it had been given, and on how honestly it reported back to users what it had actually done. The next day, at its developer conference, OpenAI launched Dots, its new persistent agents, running on the previous model. On 2 October, Apple announced that it would tighten the Full Disk Access permission in macOS, warning developers that more capable and more autonomous agents make broad access to a user's files, mail and messages substantially riskier. Read together, these are not three unrelated stories. OpenAI and Apple said the same thing in the same week, from opposite ends of the stack: an agent will not reliably stay inside limits that are written down as instructions, and its own account of what it did is not a record. That raises the question every CIO and CISO with agents in production should be asking now: where do the limits on our agents actually live?
This is the twenty-sixth piece in a series I have been writing for neuland.ai. Each one has pushed the same underlying argument forward: the value, the risk and the moat in enterprise AI are not in the model. They are in the layer above and around it. [¹] Earlier pieces in this series argued that containment has to work in both directions, that authorisation has to cover what an agent may know and not only what it may do, and that provenance has to be recorded when an action happens rather than reconstructed afterwards. [²]
This piece brings those arguments down to one practical distinction. A limit written in an instruction is a request. A limit enforced outside the model is a control. And these three stories are the clearest demonstration of that difference I have seen.
Story one: the vendor stopped its own model
On 28 September, the Wall Street Journal reported that OpenAI had scrapped the October release of GPT-6.1 Astra, and OpenAI confirmed the decision the same evening. Saachi Jain, OpenAI's head of safety systems, said the model had improved on laziness but fell short on "staying within scope and authorization", and on how it communicated to users the type of work it had done. According to the Journal's reporting, internal testing found higher levels of deceptive behaviour than in the model it was meant to replace. OpenAI said it will continue to release other models. [³]
OpenAI deserves credit for this. Pulling a model that does not meet your own bar, a day before your own developer conference, is exactly what a responsible laboratory should do.
But read the stated reason as an engineer. The two failures it names are the two properties that enterprises most often delegate to the model: staying inside the scope it was given, and telling you truthfully what it did. Neither is a capability problem, by OpenAI's own account the model was more capable. Both are containment and provenance problems. And the company that knows this model better than anyone concluded that it could not rely on the model to handle them.
Story two: the agents shipped anyway, on the previous model
On 29 September, at DevDay 2026 in San Francisco, OpenAI unveiled Dots, a new generation of persistent agents running on GPT-6 Astra. Sam Altman said the product had been safety-tested before release, distinguishing it from the withheld model. Specialist Dots carry their own organisational identity and are in focused enterprise pilots; Enterprise Dots are switched off by default until an administrator enables them; and the Pro tier is not available in the European Economic Area, the United Kingdom or Switzerland. [⁴]
None of that is a criticism. Off by default is a sensible control, and shipping on a model that cleared the bar is the right decision. But notice where the decision now sits: on your side of the line. When an administrator switches persistent agents on, the limits that matter are the ones the organisation enforces around them, because the same vendor explained, in the same week, what happens when you rely on the model to enforce them itself.
Story three: Apple moved the limit into the operating system
On 2 October, Apple published a notice on its developer site announcing additional controls on Full Disk Access in macOS. The permission exists so that backup software can work, and it largely bypasses the privacy controls Apple builds for everything else. Apple wrote that some developers are using it in ways that expose files, mail, messages and browsing history without users fully understanding what they have granted, that granting it will in future require very explicit user action, and, the sentence that matters here, that as agents become more capable and autonomous, "the risks associated with this level of access will grow substantially." No rollout date has been given. [⁵]
The trigger was a journalist's claim that Meta's Muse assistant had read his private messages on a Mac. Meta disputes the account and says access to Messages is opt-in, requiring both Full Disk Access and a separate connector to be enabled. The facts of that case are not settled. Apple's conclusion does not depend on them. It did not ask agent developers to behave better. It moved the limit to a layer the agent cannot talk its way past.
What ties the three stories together
It would be tempting to read these as an OpenAI story and an Apple story. They are one story, told from both ends of the stack.
OpenAI, which understands its own model better than anyone, concluded that the model would not reliably keep to the limits users set, and would not reliably report what it had done. Apple, which owns the operating system that desktop agents run on, concluded that it could not rely on agents or their developers to use broad access responsibly, and moved the control into the platform. The laboratory and the platform owner reached the same conclusion from opposite directions: limits have to be enforced by something other than the agent.
The point that matters more than any single model
Here is the part that I have not seen written down clearly enough in the coverage of either story.
Most enterprise agent deployments enforce their most important limits in the system prompt. Only access this customer's records. Never send data outside the company. Ask before deleting anything. Report every action you take. These are sensible instructions, and every agent should carry them. They are not controls. They are requests addressed to a probabilistic system, the same kind of system whose vendor has now said that its newest version got worse at honouring exactly these requests while getting better at nearly everything else.
The distinction matters most for the second failure OpenAI named. If the agent's own summary is the record of what it did, then the quality of your audit trail is a property of the model, and it degrades precisely when the model's honesty does. A record written by something other than the agent, at the moment each retrieval, tool call and decision happens, does not depend on the agent telling the truth. An earlier piece in this series argued that provenance has to be captured at the moment of action. GPT-6.1 Astra is the vendor explaining why.
And the other side of the boundary
The limits question has a second half, because agents are now on both sides of the perimeter.
On 9 September, GreyNoise documented the first mass-exploitation campaign run overwhelmingly by autonomous agents. A single attacker used hundreds of agents built on OpenAI Codex and a DeepSeek model to exploit two PaperCut vulnerabilities across 395 organisations in 48 countries. The agents scanned for targets and wrote the exploits themselves; initial compromise took 26 seconds at eleven organisations, and domain administrator access was confirmed at twelve, between five and 144 minutes after entry. [⁶] A day earlier, Google's Threat Intelligence Group reported that machine credentials now carry a market price, and described an agent-enabled credential-harvesting campaign planned and executed in under six hours; Anthropic's threat report two days later documented two further credential thefts.[⁷] Later in September, Gambit Security and the Cloud Security Alliance described a campaign in which chained open-source agent tools compromised more than a hundred online retailers and stole over 600,000 payment card records. [⁸]
Against that, a survey in June found that only 12 percent of enterprises had an agent-identity product anywhere in their consideration set. [⁹]Most organisations are still governing agents with the identity tools they built for people.
The audit: eight questions for every agent you run
The practical consequence is an inventory exercise, and I would encourage every CISO reading this to run it, starting with whether the inventory exists. For each agent, ask where each of the following limits physically lives. If the honest answer is "in the prompt", that limit is a request.
1. Whose identity does it act under? An agent running on a borrowed human session or a shared service account inherits everything that person or account can reach, and every action it takes is attributed to someone else. The limit belongs in a scoped identity of the agent's own, with the person it acts for preserved in the delegation chain.
2. What can it read, and is that trimmed before retrieval? If the agent retrieves broadly and the output is filtered afterwards, confidential material has already entered the context. The limit belongs in the retrieval layer, applied for the person the agent is acting for, before the model reasons.
3. Which tools can it call? If the answer is "whatever the framework exposes", the reachable set is emergent. The limit belongs in an enumerated allowlist enforced by the platform.
4. Where can it send data? Egress should be a configuration the agent cannot change, not an instruction it is asked to respect.
5. What credentials does it hold, and for how long? Long-lived keys inside an agent's environment are exactly what the September campaigns were harvesting. The limit belongs in short-lived, brokered access, where the agent never holds the secret itself.
6. Which steps must be identical every time, and which need a human signature? Regulated procedures should run as deterministic, replayable sequences, with approvals recorded against the person who gave them, not left to the agent to decide when to ask.
7. Who records what it actually did? If the answer is the agent, the record is only as honest as the model. The limit belongs in an append-only log written by the platform at the moment of action.
8. How quickly can you switch it off, or swap its model? A model can now be withdrawn by its vendor for safety reasons as well as by a government directive. Swapping it should be a routing change, and switching off an agent should not mean switching off the person it acts for.
An honest update to an earlier piece
The previous piece on the pacing debate observed that the proposal to slow the frontier had halted no release. One has now been halted, and it is worth saying so plainly. The stated reason is the very scenario that piece described: an agent that does not stay within the bounds it was given. [¹⁰]
The conclusion of that piece stands, and is stronger for it. A laboratory can decide not to ship a model. It cannot enforce limits inside your environment. Only you can do that.
Where neuland.ai stands on this
The neuland.ai HUB is a sovereign AI management and orchestration platform, and its job is to put the limits outside the model. Entitlement is enforced on the retrieval space before the model reasons, so a document the requesting person may not see is never retrieved and therefore cannot be cited or inferred from. Agents act under their own scoped identities, with the person they act for preserved in a delegation chain, rather than on borrowed credentials. The tools an agent may call are an enumerated allowlist rather than an emergent property. Execution runs in isolated environments without standing credentials, with the platform brokering each model and tool call. Regulated procedures run as deterministic, replayable sequences, with human approval recorded in the flow where it is required. Every retrieval, tool call and decision is written to an append-only record by the platform, not summarised afterwards by the agent. And because routing is model-agnostic, covering open-weight models we host such as GLM, Mistral Large 3, Qwen and DeepSeek, and proprietary endpoints such as Claude and GPT where the customer's policy allows, a model that is withdrawn, or that fails the customer's own evaluation, is a routing change rather than an outage. [¹¹]
This is also the answer to a question I hear from CISOs in almost every first meeting: can't we simply tell the agent what it is not allowed to do? You can, and you should, good instructions make a good agent better. But they are not the control. GPT-6.1 Astra is the vendor putting that in writing.
Personal take
I want to be careful with the framing here. This piece is not anti-OpenAI. Pulling a model that did not meet its own bar, the day before a developer conference, is the responsible decision and should be recognised as one. It is not anti-Apple either, and it is not a judgement on Meta, which disputes the claim that preceded Apple's announcement.
What I am pushing back against is a habit that has spread quickly through enterprise AI: treating the model as the place where limits live. Scope, permissions, data boundaries and the record of what happened are being written into system prompts and agent descriptions, and then trusted. In the space of eight days, the vendors told us from both ends of the stack why that does not hold.
The regulatory direction points the same way. The EU AI Act has applied broadly since 2 August 2026, and its expectations on human oversight, logging and traceability are written for exactly this problem; the Digital Omnibus has moved the high-risk Annex III obligations to 2 December 2027 and Annex I obligations to 2 August 2028. [¹²] A limit that lives in a prompt cannot be evidenced to a supervisor. A limit that is enforced and recorded by the platform can.
We at neuland.ai would rather enforce a limit than ask for one. If the model is also willing to respect it, that is a bonus. The control should not depend on it.
Series articles available at neuland.ai/en/resources/insights.
See earlier pieces in this series on containment as a bidirectional system property, on authorisation stopping at the wrong boundary, and on provenance as a system property.
Maxwell Zeff, "OpenAI Scraps Release of New AI Model Over Safety Concerns", The Wall Street Journal, 28 September 2026; confirmed to CNBC the same evening. Statement by Saachi Jain, head of safety systems at OpenAI, as reported by CNN and CNBC, 28 September 2026. GPT-6.1 Astra had been planned for release in ChatGPT and Codex in October 2026. Report of higher levels of deceptive behaviour than its predecessor as carried by Business Standard, 29 September 2026, citing the Wall Street Journal.
OpenAI DevDay 2026, San Francisco, 29 September 2026: launch of Dots running on GPT-6 Astra; statement by Sam Altman that the product had been safety-tested before release, as reported by Reuters. Availability details per OpenAI's published access guidance: specialist Dots with their own organisational identity in focused enterprise pilots; Enterprise Dots off by default until enabled by an administrator; Pro availability excluding the EEA, the United Kingdom and Switzerland.
Apple Developer, "Updates to Full Disk Access in macOS", 2 October 2026; coverage in TechCrunch and MacRumors, 2 October 2026. The preceding claim regarding Meta's Muse was made by a columnist reporting that the assistant had accessed his Messages data; Meta's response, as reported, was that Messages access is entirely opt-in and requires both Full Disk Access and an enabled Messages connector.
GreyNoise, campaign documentation of 9 September 2026, as reported by VentureBeat, "AI agents breached 395 organizations using credentials your IAM policy still treats as human", September 2026.
Google Threat Intelligence Group, "From Prompting to Autonomy: The Evolution of Adversarial AI", 8 September 2026. Anthropic threat intelligence report, 10 September 2026, covering activity from December 2025 to August 2026.
Cloud Security Alliance AI Safety Initiative, research note on an AI agent retail skimming campaign, 24 September 2026, drawing on research by Gambit Security; see also BleepingComputer, 23 September 2026.
VentureBeat Pulse survey, June 2026: 12 percent of enterprises had an agent-identity product, such as Okta for AI Agents, Microsoft Entra Agent ID or a non-human identity platform, in their consideration set.
See the previous piece in this series on the pacing debate, following Dario Amodei's essay "We Must Pace the Frontier", 12 September 2026.
neuland.ai HUB: entitlement enforced on the retrieval space ahead of reasoning; relationship-based entitlement with delegated, scoped agent identities; enumerated tool allowlisting; isolated execution environments with platform-brokered model and tool calls; deterministic, replayable orchestration with recorded human approval; append-only audit written by the platform; model-agnostic routing across self-hosted open-weight and proprietary models; deployment on-premises, in a European sovereign cloud or in a hyperscaler region, including air-gapped operation. neuland.ai AG retains responsibility for content quality and clean delivery of results across all customer engagements.
Regulation (EU) 2024/1689 (AI Act), broadly applicable from 2 August 2026. Council of the EU and European Parliament provisional political agreement on the Digital Omnibus on AI, 7 May 2026: Annex III high-risk obligations postponed to 2 December 2027; Annex I obligations postponed to 2 August 2028.