Choosing between Codex models on real work

The axis that matters
Model names and version numbers change often enough that memorising them is wasted effort. The axis underneath does not change. On one end you have faster, cheaper models that answer well when the task is local and the context is small. On the other you have slower, more expensive ones that reason for longer and keep a multi-step plan coherent across many tool calls.
Most teams pick one, put it in a config file, and never revisit it. That is where the money goes, and where the frustration comes from too.
Match the Codex model to the shape of the task
A starting hypothesis to test on your own tasks:
- Rename, reformat, add a test to an existing suite, write a commit message: try the faster model and check correctness
- Explain unfamiliar code, summarise a diff, answer a question about the repo: compare the faster model against a known-correct answer
- Debug something with a non-obvious cause across several files: the stronger model, and evaluate the diagnosis against a reproducible failure
- Design a change that touches a boundary you care about: stronger model, and read the plan before approving anything
- Long refactors with many mechanical steps: fast model, but check for silently skipped cases
The tell that you are on the wrong model is repetition. If the agent tries the same fix twice with small variations, it has lost the thread. Stop, switch up, and restate the problem rather than paying for a third attempt.
Reasoning effort is a second dial
Alongside model choice, Codex lets you ask for more or less deliberation before it acts. Higher effort can increase latency and token use; whether it improves the result depends on the task and model. Measure correctness and total task cost rather than assuming more effort is always better.
Treat the two dials together. A strong model at low effort and a fast model at high effort are different tools, and neither is a substitute for stating the problem clearly. No model setting rescues an instruction like "fix the login bug" with no reproduction.
What the choice will not fix
Model selection does not fix missing repo context. If the agent does not know your conventions, a larger model will violate them more fluently. Write the instruction file first. Before comparing models, confirm that both receive the relevant ownership rules and verification commands.
Try this
Take five tasks you did this week with the agent. Rerun two of them on the cheaper model and two on the stronger one. Time them and read the diffs. Use the observed results to decide which task categories justify the additional cost. Encode that split as a team habit and move on.
If you want help putting this into practice, talk to us.
Where does your team stand?
Each team member completes the proficiency matrix individually. You receive a PDF with the team baseline and a recommended next step.
Assess your team