Setting a codex SLI your team will actually watch

Why a codex SLI is worth defining
Service level indicators come from operations: pick a measurable signal, watch it, argue about it with data instead of anecdote. Apply the same idea to agent-assisted work. A codex SLI is a number you agree to track about the output of the agent, not about how much people like it.
Without a baseline, a memorable success or failure can dominate the decision. Measure comparable tasks before drawing a conclusion about the tool.
Four indicators worth the effort
- First-diff acceptance rate. Of agent-produced changes, what fraction got merged without the author rewriting the approach. Not without any edits. Without a rewrite.
- Review time per merged line. Compare review effort on similar tasks; lines are a noisy denominator and a small change can still need substantial design review.
- Rework within fourteen days. Count linked corrective changes for defects introduced by the original change. Merely editing the same file does not establish rework. This is where confidently wrong work shows up.
- Test suite pass on first push. Cheap to collect from CI, and it moves quickly when your instruction file improves.
Four is enough. Teams that pick twelve indicators check none of them by month two.
How to collect it without building a platform
Do not start with a dashboard. Start with a label. Tag agent-assisted pull requests with something like agent-assisted and query them later. A single command gets you a baseline:
gh pr list --label agent-assisted --state merged --limit 100 --json number,createdAt,mergedAt,additions,deletions
That gives you cycle time and size. Record whether the approach required a rewrite at merge, then follow up after fourteen days for linked corrective changes. The command does not measure active review time or future rework by itself.
The indicators that mislead
Lines of code generated is the worst one. It rewards verbosity, and agents are already verbose. Number of tasks run is close behind. Both go up when the tool is being used badly.
Suggestion acceptance from an editor completion counts a keystroke, not a decision, so it drifts far from anything a director cares about. Time saved, self-reported, is a survey of optimism. Compare survey responses with cycle time and quality evidence; they measure different things.
Be honest about a limit here too. None of these indicators separate a good change from a change that merely passed. A codex SLI tells you where to look. It does not tell you the code is right.
Where does your team stand?
Each team member completes the proficiency matrix individually. You receive a PDF with the team baseline and a recommended next step.
Assess your team