How agents in Codex actually decide what to do

The loop
An agent gets your instruction plus whatever context it can gather. It picks an action: read a file, search the codebase, run a command, propose an edit. It sees the result. Then it decides again. That repeats until it thinks the task is done or it runs out of room.
Everything people find surprising about agents in Codex comes out of that loop. It reads the wrong file because the search returned the wrong file. It writes a fix that does not compile because it never ran the compiler. It gives up on a flaky test because the second run also failed and it concluded the code was broken.
Context is the scarce resource
The active context usually cannot hold an entire large repository. It holds a window, and every command output, file read, and stack trace it pulls in eats that window. A long session can lose useful detail as output accumulates and earlier context is compacted. Check what information remains available before assuming the agent remembers an earlier constraint.
Practical consequences:
- Start a fresh session per task. Do not run a four hour conversation.
- Save long test output to a log, preserve the test command’s exit status, and inspect the relevant failure section. A pipeline ending in
tailcan otherwise hide the test failure status. - Point it at the file when you know the file. "Fix the retry in
src/queue/worker.ts" beats "fix the retry bug" by a wide margin. - Ask for a plan before a large change, then approve or correct the plan. Cheap to fix at that stage, expensive after four hundred lines of diff.
What good instructions look like
The strongest signal you can give an agent is a way to check itself. A reproducing test gives the agent useful feedback, but a pass does not establish that the requirements or test itself are correct. If the task has no verification, the agent will produce code that looks right, declare success, and be wrong roughly as often as an unverified human patch.
So the pattern we teach is to state the check in the prompt. "Write a failing test that reproduces this, then make it pass, then run the full suite." Three clauses. Inspect that test to ensure it fails for the reported behavior and does not merely mirror the implementation.
The limits worth knowing
Agents are weak on anything that requires knowing what your users expect. They are weak on ambiguity, and they resolve ambiguity silently by picking one reading and building it. They also tend toward addition. Ask for a fix and you get a new branch in a conditional, not the deletion that would have been the right answer.
Watch for the agent that has been stuck for several turns. If it has tried the same approach three times with variations, stop it. Read the failure yourself, tell it the constraint it is missing, and restart. More turns alone do not supply a missing constraint.
Do this next
Take your last three merged pull requests. Re-run each one as an agent task in a scratch branch and compare the diffs against what your team actually shipped. You will learn more about where agents in Codex help you in that hour than in a week of reading about them.
If you want help putting this into practice, talk to us.
Where does your team stand?
Each team member completes the proficiency matrix individually. You receive a PDF with the team baseline and a recommended next step.
Assess your team