
Field note · AI Engineering · · 2 min read
Treat a model upgrade like a permission change
Consider an invented example. A coding agent that can edit code and run commands is asked to fix an invoice export bug. The previous model changed only the export code. The replacement writes a cleanup script that also changes invoice records, which nobody asked for. Both models had the same access. They used it differently.
So give a change of default model the review you would give a new permission, even when the permission settings stay exactly the same.
For Copilot Business and Enterprise, GitHub's September 4 release note (opens in a new tab) says new models are enabled automatically unless administrators change the policy. Choosing the default is a separate control, which GitHub added for enterprises and teams (opens in a new tab) on September 2 in supported Copilot clients. That is where the rollout decision belongs.
Test the approach as well as the answer
Published results are a reason to run your own test. OpenAI's GPT-6 Astra system card (opens in a new tab) reports better behavior in workplace settings and a greater ability to find and exploit security flaws. ARC Prize (opens in a new tab) found substantially different results when it changed how reasoning state was kept between requests. Neither can tell you whether your invoice workflow will respect its approval rules.
For the invoice example, make an isolated copy with synthetic records and give both models the same request. Decide what success means first: the export is fixed and the invoice records are unchanged. Keep instructions and tool settings identical for both runs. Then inspect the commands each model issued and the data that changed, including anything a control blocked.
Evidence to collect from the comparison
Permitted work
The export is correct, and the invoice records are unchanged.
Approval boundary
A proposed record change waits for approval. Denying it writes nothing.
Scope change
An instruction hidden in a test record cannot authorize an unrelated cleanup. Note any attempt and whether it was blocked.
Recovery
Stop the test, check remaining work, and restore the previous setup. Account for any data already changed.
Include a run the model cannot finish within the allowed scope. A good outcome is an explanation and a request for approval. If safety depends on the old model missing an action it was allowed to take, fix the access rule before rollout.
Make the rollout decision from the recorded result
Pick a few representative tasks and repeat the ones where a wrong action would matter. Record the model version and settings with each result. A clean run supports a limited rollout; it cannot vouch for every future run. Name an owner who can move new work back to the previous setup, and check that the previous model is still available.
Switching back does not undo an action already completed in another system; the cancellation rehearsal covers work that is still running. The decision record should show what changed in behavior, which failures remain, and why the owner accepts them for this use.
Written by the Moga principals.
More from AI Engineering
- 2 min read
Review an agent skill like a software dependency
Trace what a shared agent skill can run and access before your team installs it.
- 2 min read
Stopping an AI agent means checking what can still run
Cancel a harmless test run, then check every system it touched for work or access that is still live.