Back to Research

Agents API computer use: build a browser agent step by step

Build a first browser agent with OpenAI Agents API computer use: create the session, approve website origins, send a task, handle sign-in safely, and clean up.

Railroad Cut (La Tranchée), landscape painting by Paul Cézanne (1867).
DesignRogier MullerSeptember 29, 20266 min read

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.

Agents API computer use lets an agent operate a browser in an OpenAI-hosted environment, so it can test a website, collect information, or use an application through its UI. You enable it by adding a computer_use tool and an openai_hosted environment with the desktop turned on. Your application then approves each website the browser wants to open. The computer use guide is the source for every parameter below.

OpenAI announced the feature at DevDay on 29 September 2026. This walkthrough builds the smallest useful browser agent in JavaScript, using only the documented request fields. It reads one public page and reports what it found.

Who can use OpenAI Agents API computer use?

The Agents API has been in public beta since 10 September 2026. Every request in the docs sends the OpenAI-Beta: agents=v1 header. The DevDay recap says computer use is available through the API, and in Codex and ChatGPT Work on Pro 500 and Enterprise.

Item What the docs say
Status Agents API public beta
Where the browser runs An OpenAI-hosted environment
Data residency United States only, per the Agents API overview
Zero Data Retention Not supported by the Agents API
Billing Model usage at the selected model's API rates, OpenAI tools at standard rates, hosted sandboxes at standard container rates
Idle cleanup A hosted sandbox can be deleted after an hour without activity or keep-alives

The ZDR line matters. If your organization requires Zero Data Retention for API traffic, the Agents API is not an option for that workload.

The same launch brought Codex's multi-agent capabilities, tool search, tool calling and context compaction to the Agents API. The overview describes it as access to the Codex harness through an OpenAI-managed API. This article sticks to the browser.

Create a browser session

Install the openai SDK and export OPENAI_API_KEY, following the Agents API quickstart. The JavaScript examples also use prompt-sync for terminal input. Creating a session does not start a task. This is the docs example, unchanged:

import OpenAI from "openai";

const client = new OpenAI();
const session = await client.beta.agents.sessions.create({
  agent: {
    model: "gpt-6-astra",
    instructions:
      "Read public documentation in the browser. Do not sign in or change any website data. Report the page title and URL you find.",
    tools: [{ type: "computer_use", include_screenshots: true }],
  },
  environment: {
    type: "openai_hosted",
    desktop: { enabled: true },
    network: { access: "enabled" },
  },
});
console.log("Session ID:", session.id);

include_screenshots only controls whether screenshots appear in API output. The agent observes the browser either way. Screenshots can contain account data, so show them only to authorized users and keep them out of logs.

Approve each website origin

Plan for this step before anything else. The browser asks for approval before it opens each new website origin, public sites included. Setting network.access to enabled does not approve anything.

When the stream emits agent.session.requires_action, retrieve the session and read its required_actions. Pending entries of type computer_use_approval_request with a nested request.type of browser_origin_access carry the origin and an optional reason. Show them to the user, collect approve, deny or cancel, and send the decision back with the same request_id. In the docs helper, approval is the pending action and decision is the user's choice:

await client.beta.agents.sessions.events.create(sessionId, {
  events: [
    {
      type: "agent.session.input.computer_use_approval_request_result",
      request_id: approval.request_id,
      response: { type: "browser_origin_access", decision },
    },
  ],
});

The docs example defaults the prompt to deny, which is the right default to copy. Origin approval is per origin, not per action. It will not stop a purchase or a delete on a site you already approved. If you need that guarantee, restrict the browser to resources that cannot perform those actions, or use a browser runtime you control. The docs also say to treat website content as untrusted: a page cannot grant permission or override the user's instructions.

Send the task and follow the root turn

Open the event stream first, then send the task, so you catch the first progress events. In the docs example, the task message is:

await client.beta.agents.sessions.events.create(session.id, {
  events: [
    {
      type: "agent.session.input.message",
      input: [
        {
          role: "user",
          content: [
            {
              type: "input_text",
              text: "Open https://developers.openai.com in the browser. Find the Agents API quickstart, then report its page title and URL.",
            },
          ],
        },
      ],
    },
  ],
});

Loop over the stream and handle four cases. Print agent.session.turn.output_text.done text. Handle agent.session.requires_action with the approval code above. Treat agent.session.failed and agent.session.environment.failed as fatal. Stop on agent.session.turn.completed only when event.turn.subagent_id is null, because that is the main agent's turn. In the docs example, a failed or cancelled subagent turn does not end the loop.

Browser steps show up as computer_use_call items with an id, turn_id, title, status and optional screenshot output. They describe tool operations, not the final answer.

Sign-in without passing credentials to the model

Tasks such as reading a private repository's issues need an authenticated browser. Only the main agent can request sign-in, and the flow supports email addresses, passwords and verification codes. Passkeys and QR-code sign-in are not supported.

Sign-in arrives as a computer_use_approval_request with request.type set to browser_authentication. Render the form from its fields and options. Show the credential_origin and cancel if the user cannot verify it. Submit with action: "submit" through the same approval event. Those values stay out of the model input and out of saved history. Function-tool results do not, so never collect a password with a function tool.

A 202 response means the submission was accepted, not that sign-in worked. Disable automatic SDK retries for credential submissions. Authentication requests expire after five minutes.

Tighten the network and clean up

For a real task, switch network.access from enabled to restricted and list the hosts in allowed_domains. The hosted environment docs accept 1 to 100 exact host names, with no wildcards, protocols, paths or ports. Include the domains a page needs for resources and redirects.

When the root turn completes, fails or is cancelled, call client.beta.agents.sessions.delete(session.id) to request environment cleanup.

Build it with Codex

The quickest route is to let Codex write the event loop from the guide and keep the review on your side. Paste the guide's URL into a Codex task and ask for a script that creates the session, prompts for origin decisions with deny as the default, and cancels every sign-in request. Then read the diff line by line against the docs. Our methodology applies here as it does to any delegated change: Codex drafts, you check each request field against the source.

Run the public-docs task above first. Once it reports the quickstart title and URL, switch the network to restricted and point it at your own staging site.

Further reading

Related training topics

Learn more

Learn more

Learn more

Learn more

Review is one step in the methodology.

Related research

Continue through the research archive

Practise Review with the team

Book a date if you already want one.

See training