Skip to content

Field note · AI Engineering ·  · 3 min read

Give a coding agent work small enough to verify

Picture a coding agent asked to “build the publishing workflow” for a hypothetical editorial tool. The agent must settle questions nobody wrote down, such as what counts as a saved draft and what happens when a save fails. You discover its answers in review, spread through one large change.

Hand over a slice instead: one behavior the agent can finish and you can verify, with its boundaries and evidence agreed before any code changes.

Write the brief from the current code

Read the current code and tests before writing the brief. Last week's plan may ask for work that already shipped. If it calls for a drafts table that already exists, fix the plan before the agent reads it.

A sensible first slice lets an editor save a draft and get the same content back after a reload. Its brief also names what must keep working and what waits for later. Write ambiguous cases, such as a failed save, as Given/When/Then examples so the agent does not have to guess.

That slice runs from the editor screen to the database. Crossing those layers can still be manageable when a reviewer can check one behavior end to end. Google's code review guidance (opens in a new tab) puts the right size at one self-contained change with its related tests, and suggests smaller full-stack features as one way to split work. Google also counts lines and files, so one behavior spread across fifty files deserves a second look.

A brief for the draft-save slice

The Moga mascot traces a continuous paper path from a blank editor card through a storage compartment and back to the card.

The draft-save slice runs from the editor to storage and back to the editor.

  • Outcome

    An editor saves a draft, reloads the page, and gets the same content back.

  • If saving fails

    Given unsaved changes, when the save fails, then the text stays in the editor and a message says the draft was not saved.

  • Boundaries

    Published posts and editing permissions work as before. Approval routing and scheduled publication wait for later slices.

  • Evidence

    Tests cover save/reload, failed saves and access. Browser checks cover the editor and published posts. Record the results.

Illustrative brief for the hypothetical editorial tool.

Check each slice before starting the next

Small slices help only if each one is checked as it lands. DORA's guidance on small batches (opens in a new tab) warns that AI tools are often optimized to generate large, complete features, and that machine-generated code may take more review effort per line than human-written code. Regrouping small batches before testing, it adds, delays feedback on defects. A week of slices saved for one review is that kind of regrouping.

The brief's evidence tells the agent what it must demonstrate. Anthropic's best practices for Claude Code (opens in a new tab) note that without a check it can run, the agent's only signal is that the work looks done. The guide recommends stating excluded work and finishing the specification with a check of the whole feature.

Protect those checks while the agent works. If the editor loses text on a failed save, first watch a test catch that loss. Run it again after the fix. If changing a test's expectation would change the agreed behavior, revisit the brief. Adapt our two-account test to check both an allowed save and a refused save at the server. Then compare the diff with the brief and remove unrelated changes.

When the slice is done, update the documentation that owns any changed behavior or decision, such as a decision record or the agent's instructions file. Point the next brief to that updated record. Record what was verified and what is still open, keeping a local pass separate from a merge or a deploy, since each needs its own check.

The Moga mascot pencils a correction onto a paper outline so it matches the finished piece beside it, while a blank sheet waits to one side.
Update the record to match what was built, and keep the open work in view.

A typo fix needs no brief, and Anthropic's guide suggests skipping the plan when you could describe the diff in one sentence. An unfamiliar integration may need a short investigation before it can be sliced, with a question to answer and a point where it stops.

Use what each slice teaches you in the next brief. A case the tests missed, or a decision the agent had to guess, usually belongs there before the agent starts again.

Written by the Moga principals.

See the AI Engineering service

Bring the idea or the prototype.

Every engagement starts with a scoping call.

Book a call