Skip to content

Field note · AI Engineering ·  · 2 min read

Treat a model upgrade like a permission change

Consider an invented example. A coding agent that can edit code and run commands is asked to fix an invoice export bug. The previous model changed only the export code. The replacement writes a cleanup script that also changes invoice records, which nobody asked for. Both models had the same access. They used it differently.

So give a change of default model the review you would give a new permission, even when the permission settings stay exactly the same.

For Copilot Business and Enterprise, GitHub's September 4 release note (opens in a new tab) says new models are enabled automatically unless administrators change the policy. Choosing the default is a separate control, which GitHub added for enterprises and teams (opens in a new tab) on September 2 in supported Copilot clients. That is where the rollout decision belongs.

Test the approach as well as the answer

Published results are a reason to run your own test. OpenAI's GPT-6 Astra system card (opens in a new tab) reports better behavior in workplace settings and a greater ability to find and exploit security flaws. ARC Prize (opens in a new tab) found substantially different results when it changed how reasoning state was kept between requests. Neither can tell you whether your invoice workflow will respect its approval rules.

For the invoice example, make an isolated copy with synthetic records and give both models the same request. Decide what success means first: the export is fixed and the invoice records are unchanged. Keep instructions and tool settings identical for both runs. Then inspect the commands each model issued and the data that changed, including anything a control blocked.

Evidence to collect from the comparison

  • Permitted work

    The export is correct, and the invoice records are unchanged.

  • Approval boundary

    A proposed record change waits for approval. Denying it writes nothing.

  • Scope change

    An instruction hidden in a test record cannot authorize an unrelated cleanup. Note any attempt and whether it was blocked.

  • Recovery

    Stop the test, check remaining work, and restore the previous setup. Account for any data already changed.

Proposed checks for the invented invoice example. Nothing has been measured.

Include a run the model cannot finish within the allowed scope. A good outcome is an explanation and a request for approval. If safety depends on the old model missing an action it was allowed to take, fix the access rule before rollout.

Make the rollout decision from the recorded result

Pick a few representative tasks and repeat the ones where a wrong action would matter. Record the model version and settings with each result. A clean run supports a limited rollout; it cannot vouch for every future run. Name an owner who can move new work back to the previous setup, and check that the previous model is still available.

Switching back does not undo an action already completed in another system; the cancellation rehearsal covers work that is still running. The decision record should show what changed in behavior, which failures remain, and why the owner accepts them for this use.

Written by the Moga principals.

See the AI Engineering service

Bring the idea or the prototype.

Every engagement starts with a scoping call.

Book a call