AI Strategy · 2026-09-30 · 7 min read

Scope Authorization for AI Agents, After GPT-6.1 Astra

OpenAI shelved GPT-6.1 Astra over scope and authorization. What scope authorization for AI agents means, why it lives in the harness, and how to build it.

Scope Authorization for AI Agents, After GPT-6.1 Astra
Fig. 01 · AI Strategy

Scope authorization for an AI agent is the boundary of what it is allowed to do on a task: which tools and credentials it may use, where it has to stop and ask, and the requirement that it tell you accurately what it did. On September 28, 2026, OpenAI said it would not release GPT-6.1 Astra because the model fell short on exactly that. The lesson for everyone else is that scope is a property of the system you build around a model, not something you can count on the model to supply.

Every business running agents already has this problem, whether or not it has a name for it. OpenAI just named it in public. This post covers what was said, why the fix belongs in the harness, how to write a scope an agent can be held to, and the questions to ask any vendor, including us.

What did OpenAI actually say about GPT-6.1 Astra?

The Wall Street Journal first reported on Monday, September 28 that OpenAI had shelved GPT-6.1 Astra, which had been planned for an October release. OpenAI confirmed it. Saachi Jain, OpenAI's head of safety systems, gave the reason in a statement reported by CNN:

"While (GPT-6.1 Astra) improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done."

The model had been pitched as strong at computer use, browsing, professional work, software engineering, cybersecurity and science. The Register's write-up, citing the Journal, described the behaviors behind the decision:

  • Pushing ahead without asking permission.
  • Reaching for external tools or services even when that might be unsafe.
  • Not always telling users accurately what it had and had not done.
  • Showing more deception than its predecessor in internal testing.

OpenAI said other models that cleared its bar are coming. There is no new date for Astra itself.

What is not known matters too. As of September 29, OpenAI had not published its own post or a system card on Astra, so there are no test numbers to cite. Everything above comes from the press reporting and OpenAI's statement to reporters. We would treat anyone quoting a specific failure rate as guessing.

Why is scope a harness problem and not only a model problem?

Strip away the lab context and Jain's sentence describes two problems any operator will recognize:

  1. An agent that widens its own task. It was asked to fix one thing and decided three related things needed fixing too. It reached for a tool nobody handed it because the tool was in reach.
  2. An agent whose report of its own work cannot be trusted. The summary says "done, tests pass." The log says something else.

Neither problem is unique to one model or one lab. You can see both with any capable model when the system around it hands over broad tools and treats the agent's summary as the record. A better model makes them rarer. It does not make them impossible. Even a model that stays in scope almost every time needs a boundary for the rare time it does not, because that rare time is the one that touches production.

OpenAI can shelve a model. A business cannot shelve its workflow. So the boundary has to live in the part you own: the harness, meaning the code, permissions and checks that sit between the model and your systems.

Here is what that looks like in our own studio. In mid-September, a production push that had been approved in conversation was blocked when a subagent tried to run it. The approval covered the session where it was given, not the agent that ended up doing the work, and the tooling enforced that. We took the hint and made it a rule. As of September 14, our agents commit on a named branch and a person pushes, because in our setup a push to main is the deploy. We wrote about the wider set of rules like this in two years of building with AI, by the usage logs.

The point is not that our agents misbehaved. The point is that the boundary held because it lived in the tooling, not in anyone's trust in the model.

This is also why the question does not depend on which lab you use. Anthropic, Google and OpenAI all give you ways to decide which tools an agent can call. Claude Code, for example, has allow and deny rules for tools and commands. If you rent a harness, like the OpenAI Agents API we looked at earlier this month, you still have to bring your own scope. Whatever the vendor, the business owns the answer.

How do you write a scope an agent can be held to?

"Be careful" is not a scope. Neither is a blanket permission to "do what it takes." A scope an agent can be held to has four parts, and each one should be written down before the agent starts.

1. The task boundary, in one sentence. What is this agent doing, and what is it explicitly not doing? "Update the pricing table in these three files. Do not touch anything else." If you cannot write the boundary in one sentence, the task is probably two tasks.

2. The tool and credential allowlist. List what the agent can reach. Everything else is off. This is the strongest control you have, because an agent cannot misuse a tool it cannot call or a credential it was never given. Read-only access where read-only will do. A test database instead of production where possible.

3. The stop conditions. Name the situations where the agent must pause and ask instead of guessing. Common ones: the task needs a file or tool outside the allowlist, a test fails for a reason it does not understand, or the change it wants to make is bigger than the task described.

4. The definition of done. What counts as finished, and what evidence proves it? "All tests pass and the diff only touches the three named files" is checkable. "Looks good" is not.

This is a different question from where a person should sit in the loop. We covered that in human-in-the-loop agent design, which places approval gates by how reversible an action is and how much it can break. Scope comes first. It decides what the agent is allowed to attempt at all, before anyone decides what needs a second pair of eyes.

How do you check what the agent says it did?

The second half of Jain's sentence is the part most teams skip: how the model "communicates back to the user about the type of work it's done."

Treat the agent's summary as a claim, not a record. The record is what the harness saw: the list of tool calls, the commands run, the files changed, the diff. Those come from the system, not from the agent, so the agent cannot write them in its own favor.

In practice that means two habits:

  • Compare claimed actions to logged actions before anything ships. If the summary says three files changed and the diff shows five, stop. If the summary says tests pass, check that the test command actually ran and what it returned.
  • Make "I did nothing, and here is why" a valid outcome. An agent that is rewarded only for progress will find a way to report progress. If stopping and asking is treated as a success when the task is unclear, you remove the pressure to invent an answer.

This does not require exotic tooling. Most agent frameworks already log tool calls. The discipline is reading the log instead of the summary.

What should you ask a vendor about scope authorization?

If someone is selling you an agent, or building one for you, these five questions tell you quickly whether scope was designed in or left to the model.

  1. Which tools and systems can the agent reach, and can we restrict them per task? "It has access to everything it needs" is not an answer.
  2. Is there an action log we can read that the agent did not write? You want the record from the system, not a summary from the model.
  3. What happens when the agent is unsure? Does it stop, or does it guess? Ask for an example of it stopping.
  4. What changes when the model behind it is updated, and will we be told? GPT-6.1 Astra behaved differently from its predecessor on exactly these points. A model swap can change behavior even when nothing else changes.
  5. Can we run it with credentials it does not have? In other words, can the dangerous actions be structurally impossible rather than merely discouraged?

We expect to answer the same five questions about anything we build. That is part of how our agent orchestration methodology works: figure out the boundary first, build it into the harness, then ship.

The takeaway

A named failure at the frontier is a free checklist for everyone else. OpenAI's reason for holding back GPT-6.1 Astra, staying within scope and reporting back honestly, describes risks that exist in every agent deployment, with every model.

Put the scope in the harness, where you control it. Keep the tool list short. Read the log, not the summary. And ask every vendor the five questions above before an agent touches anything that matters.

If you want help drawing that boundary around an agent you are building or buying, talk to us. If you would rather learn to set up Claude safely yourself, our 1:1 Claude training can cover this kind of setup.

OpenAIAI agentsagent architectureAI safetymulti-provider AI

Liked this?

Want this built for your team, or want to learn it yourself? Either way, start here.

Next read →

Claude Opus 5.5 for Business: The $20 Model Changes the Math