AI
Playwright MCP vs Browser Use vs Stagehand: Browser Agents
October 202612 min read

Playwright MCP is a tool layer: it exposes browser actions such as browser_click and browser_snapshot to an MCP client and leaves planning to that client. Browser Use is a Python agent framework whose Agent class runs the whole perceive, decide and act loop itself. Choose Playwright MCP to give an existing agent a browser, and Browser Use when you want the framework to drive an open-ended task.
Not in v4. The Stagehand migration guide states that v4 drives Chromium over the Chrome DevTools Protocol with no Playwright dependency, so a Playwright Page cannot be passed to act. Moving an existing Playwright flow to Stagehand v4 means porting it, including adding explicit waits, because v4 does not auto-wait the way Playwright does.
No. Playwright MCP works from structured accessibility snapshots, where every interactive element carries a ref the model passes to tools like browser_click. Coordinate-based tools are available only if you enable them with the --caps vision flag. Screenshot-based tools such as the OpenAI computer tool or Gemini computer use are the better fit for canvas-rendered pages.
Keep raw secrets out of the model context. Browser Use replaces them with placeholders through sensitive_data and warns if allowed_domains is not set, and Stagehand passes them as act variables so only the placeholder name reaches the model. Turn vision off on screens that show credentials, and restore logins from a storage-state file instead of having the model type them.
No. The Playwright MCP README states that its allowed and blocked origin lists are not a security boundary and do not affect redirects. Enforce the allow list at the network or proxy layer, run each session in an isolated container, and require human confirmation before purchases, data transmission or destructive changes, as the OpenAI computer use guide recommends.

Key Takeaway
Playwright MCP, Browser Use and Stagehand differ mainly in who owns the agent loop. Playwright MCP exposes accessibility-tree tools and leaves planning to your client. Browser Use runs its own Python agent loop. Stagehand v4 offers act, extract and observe over CDP, without Playwright. OpenAI and Gemini computer-use tools work from screenshots instead.
Consider an accounts-payable team at an Indonesian distributor. Forty suppliers each run their own invoice portal, none of them offers an API, and every morning someone logs into each one to check which invoices are paid, disputed or still open. It is the textbook case for a browser agent, and the first search, playwright mcp vs browser use, returns four very different kinds of tool that all claim to solve it.
They do not solve the same problem. This post compares Playwright MCP, Browser Use, Stagehand and the computer-use tools from OpenAI and Google on the questions that decide a production build: who owns the loop, whether the model reads the DOM or looks at pixels, what a repeat run costs, and where the safety controls actually live. Every API name and default below comes from each project's current documentation, cited at the end.
A browser agent is a loop: look at the page, decide the next step, act, look again. The tools in this comparison split cleanly by which part of that loop they give you and which part they leave to you.
The further down that list you go, the more of the loop is your own code, and the more of the reliability is your responsibility. That is not a ranking. A framework that owns the loop is fastest to a demo; owning the loop yourself is what lets you put an approval step exactly where an action becomes irreversible.
The table sets the four approaches side by side on the dimensions that change an architecture decision. The vendor column groups OpenAI and Gemini because their shape is the same: screenshot in, actions out.
| Dimension | Playwright MCP | Browser Use | Stagehand v4 | Computer-use tools |
|---|---|---|---|---|
| Who owns the loop | Your MCP client or agent | The framework's Agent class | Your code, step by step | Your code runs the actions the model returns |
| What the model perceives | Accessibility snapshot with element refs; no vision model needed | DOM state plus screenshots; use_vision defaults to True | Page content per call, resolved to a selector | Screenshots only |
| How an element is targeted | A ref from the last snapshot, or a selector | Element index chosen by the agent | Natural-language instruction, or a replayed observed action | Pixel coordinates; Gemini uses a 0 to 1000 grid |
| Language and runtime | Node MCP server over stdio, any MCP client | Python 3.11 or later, plus a TypeScript edition | TypeScript, Python and Go, over CDP | Any language that calls the Responses or Gemini API |
| Repeat-run cost | Every step is a model turn in your client | Every step is a model call | Observed actions replay with no LLM call; server cache on Browserbase | Every step is a model call with an image |
| Secrets handling | --secrets masks matching text in responses, a convenience only | sensitive_data placeholders, scoped per domain | variables in act; only the placeholder name reaches the model | Yours to build; typing sensitive data counts as transmission |
| Licence and hosting | Apache-2.0 from Microsoft, runs locally or in Docker | MIT library; optional cloud browsers at 0.02 USD per browser-hour | MIT; local browser or Browserbase sessions | Hosted model API; you host the browser or VM |
| Best fit | Giving an existing coding or chat agent a browser | Open-ended tasks where the path is not known in advance | Repeatable flows that must survive layout drift | Canvas apps, desktop apps, pages with poor accessibility |
Two rows carry most of the decision. Perception decides which sites work at all: an accessibility tree is cheap and precise on a well-built form and close to useless on a canvas-rendered dashboard, where only screenshots help. Repeat-run cost decides whether the flow is affordable at forty portals every morning, and only Stagehand has a built-in way to stop paying the model for steps it has already planned.
Playwright MCP is a Model Context Protocol server from Microsoft that wraps Playwright. Its README describes it as working through structured accessibility snapshots, bypassing the need for screenshots or visually tuned models. The model reads a tree of roles and names, each interactive node tagged with a ref, and passes that ref to tools such as browser_click, browser_type and browser_fill_form.
// .mcp.json (or your client's mcpServers block)
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": [
"@playwright/mcp@latest",
"--headless",
"--isolated", // profile lives in memory, gone on close
"--storage-state", "./portal-auth.json", // log in once, outside the agent
"--allowed-origins", "https://portal.supplier.example",
"--secrets", "./.portal.env" // masks matching text in tool responses
]
}
}
}
// What the model actually reads: an accessibility snapshot, not pixels.
// browser_snapshot returns a tree; every interactive node carries a ref.
// - heading "Open invoices" [level=1]
// - textbox "Invoice number" [ref=e14]
// - button "Search" [ref=e15]
// - row "INV-2026-0912 Rp 48.250.000 Unpaid" [ref=e31]
//
// The model then calls a tool with that ref as the target:
// browser_type { "target": "e14", "text": "INV-2026-0912" }
// browser_click { "target": "e15", "element": "Search button" }
//
// The planning loop is YOURS (or your MCP client's). Playwright MCP
// executes one tool call and returns the new snapshot. Nothing more.The server never plans. It runs the tool it was asked to run and reports the new state, so the quality of the agent is entirely the quality of the MCP client and model behind it. That is a strength when the client is already a capable agent, and the reason Playwright MCP is the common way to give Claude Code, Cursor or VS Code a browser. Coordinate-based tools such as browser_mouse_click_xy exist, but only when you opt in with --caps vision.
Read the flag documentation closely before relying on any of it for safety. The README states that --allowed-origins and --blocked-origins do not serve as a security boundary and do not affect redirects, that the secrets option is a convenience and not a security feature, and, in its security section, that Playwright MCP is not a security boundary. A persistent profile is the default, and only one browser instance can use it at a time, so parallel clients need --isolated or separate --user-data-dir paths.
If the caller is a coding agent rather than a chat agent, read the Playwright MCP README's own advice first. It recommends the separate Playwright CLI with skills for coding agents, because CLI invocations avoid loading large tool schemas and verbose accessibility trees into the context window. Use MCP when the agent needs persistent browser state across many reasoning steps.
Browser Use is an MIT-licensed Python library, with a TypeScript edition, whose Agent class owns the loop. You supply a task, a model and optionally a Browser and custom Tools; agent.run executes steps until the model calls done or max_steps runs out. The example below gives the agent a narrow custom tool that writes a draft note into the ERP rather than letting it touch the ERP screen.
import asyncio
from browser_use import ActionResult, Agent, Browser, ChatBrowserUse, Tools
tools = Tools()
# Give the agent a door into YOUR system instead of letting it type into the ERP UI.
@tools.action(description="Record a supplier invoice status as a DRAFT note in the ERP.")
def save_invoice_status(invoice_no: str, status: str) -> ActionResult:
draft_id = erp_client.create_draft_note(invoice_no, status) # your code
return ActionResult(extracted_content=f"draft {draft_id} saved")
async def main():
agent = Agent(
task="Log in with x_user / x_pass, open Open invoices, "
"read the status of INV-2026-0912 and save it with save_invoice_status.",
llm=ChatBrowserUse(model="bu-2-0"),
browser=Browser(allowed_domains=["https://portal.supplier.example"]),
tools=tools,
# The model only ever sees the placeholder names x_user and x_pass.
sensitive_data={"https://portal.supplier.example":
{"x_user": PORTAL_USER, "x_pass": PORTAL_PASS}},
use_vision=False, # default is True: screenshots would show the password field
max_failures=3, # default 5
step_timeout=120, # seconds, default 180
)
history = await agent.run(max_steps=25)
# is_successful() is what the AGENT reports. Check the ERP draft yourself.
print(history.is_successful(), history.final_result())
asyncio.run(main())The defaults are tuned for getting the task done, not for caution. run accepts max_steps with a default of 500 in the current source, use_vision is True, max_failures and max_actions_per_step are both 5, and step_timeout is 180 seconds. For a back-office job, lower the step budget, turn vision off whenever credentials are on screen, and pass allowed_domains: the library itself logs a warning when sensitive_data is set without it, because a prompt-injected page could otherwise read the secrets back.
The same agent can also run as a local MCP server with uvx and the browser-use CLI extra, which inverts the relationship with Playwright MCP: your chat client delegates a whole browsing sub-task to Browser Use instead of driving each click. Note that history.is_successful reports what the agent claims. The Browser Use documentation says so directly, and recommends verifying important external actions independently.
Stagehand, from Browserbase, sits between hand-written selectors and a free-running agent. You write the flow; act performs one natural-language step, extract returns data validated against a Zod schema, and observe returns candidate actions with a selector, description, method and arguments. Passing one of those observed actions back into act replays it without an LLM call, which is how a flow planned once becomes cheap to rerun.
Most comparisons still describe Stagehand as a layer on top of Playwright. That is out of date for v4. Its migration guide states that v4 drives Chromium over the Chrome DevTools Protocol with no Playwright dependency, so a Playwright Page cannot be passed to act and moving a flow means porting it. v4 also removed agent, enableCaching and cacheDir.
import { localBrowser, Stagehand } from "@browserbasehq/stagehand";
import { z } from "zod";
const browser = await localBrowser.launch();
const stagehand = await Stagehand.create({
browser,
model: { modelName: "openai/gpt-5.6-sol", apiKey: process.env.OPENAI_API_KEY },
});
const [page] = await browser.context.pages();
await page.goto("https://portal.supplier.example/login");
// Wrong: one act() with "log in, open invoices and search for 0912".
// Right: one atomic step per act(). Secrets go in as variables; only the
// placeholder name reaches the model provider.
await stagehand.act("type %user% into the username field", { variables: { user: PORTAL_USER } });
await stagehand.act("type %pass% into the password field", { variables: { pass: PORTAL_PASS } });
await stagehand.act("click the sign in button");
// v4 has NO Playwright-style auto-waiting: wait explicitly before the next step.
await page.waitForSelector("table.invoices");
// observe() plans once; act(action) replays the returned action with no LLM call.
const { data: searchSteps } = await stagehand.observe("find the invoice search box and its submit button");
for (const step of searchSteps) await stagehand.act(step);
const { data } = await stagehand.extract(
"extract every open invoice row",
z.object({ invoices: z.array(z.object({
number: z.string(), amountIdr: z.number(), status: z.string(),
})) }),
);
await stagehand.close();
await browser.close();The port is not only a rename. Stagehand.create replaces new Stagehand plus init, act, extract and observe move onto the Stagehand instance, and every primitive returns a data and metadata pair. The detail that breaks ported scripts most quietly is waiting: v4 resolves a selector once and throws if it is not there yet, unlike Playwright's auto-waiting, so explicit waitForSelector calls come back.
Caching changed in the same release. The local cache directory is gone; caching is now server-side on Browserbase, enabled by default there, keyed on the instruction, page content and call options, and deliberately not on the model. On a local browser, the observe-then-act replay pattern is the cost lever you control, and the docs recommend one atomic step per act call over chaining several steps into one instruction.
The vendor tools skip the DOM entirely. OpenAI's computer tool in the Responses API returns a computer_call holding an ordered, batched list of actions, drawn from click, double_click, drag, move, scroll, keypress, type, wait and screenshot. Your code executes them and answers with a computer_call_output screenshot. Gemini's computer_use tool works the same way, with coordinates normalised to a 0 to 1000 grid that you scale to the real screen.
# ── OpenAI: the computer tool in the Responses API ──
response = client.responses.create(
model="gpt-6.1-sol",
tools=[{"type": "computer"}],
input="Open the Open invoices page and read the status of INV-2026-0912.",
)
# The output holds a computer_call with an ORDERED, batched actions array:
# {"type": "computer_call", "call_id": "call_002", "actions": [
# {"type": "click", "button": "left", "x": 405, "y": 157},
# {"type": "type", "text": "INV-2026-0912"}]}
# Allowed: click, double_click, drag, move, scroll, keypress, type, wait, screenshot.
for action in call.actions:
if not policy.allows(action): # your allow list, your confirmation step
raise NeedsHuman(action)
execute(page, action) # Playwright, PyAutoGUI, a VM ...
response = client.responses.create(
model="gpt-6.1-sol",
tools=[{"type": "computer"}],
previous_response_id=response.id,
input=[{"type": "computer_call_output", "call_id": call.call_id,
"output": {"type": "computer_screenshot",
"image_url": f"data:image/png;base64,{png_b64}",
"detail": "original"}}],
)
# ── Gemini: the computer_use tool, coordinates on a 0-1000 grid ──
tools = [{"type": "computer_use", "environment": "browser"}]
def to_px(x, y, width=1440, height=900): # 1440x900 is the suggested browser size
return int(x / 1000 * width), int(y / 1000 * height)
# A function_call may arrive with safety_decision = require_confirmation.
# Treat that as a hard stop for a human, not a hint.Two details matter for anyone building on them now. OpenAI's guide recommends code execution, where the model writes Playwright or PyAutoGUI scripts, over the structured computer tool for its newest model, and keeps a migration section for integrations built on the older computer-use-preview. Gemini marks computer use as a Preview capability that may contain errors and security vulnerabilities, and its responses can carry a safety_decision of require_confirmation, which your loop must turn into a real pause for a human.
The consumer products built on this idea have not been stable either. Operator launched in January 2025 and shut down on 31 August 2025 after ChatGPT agent absorbed it, and ChatGPT agent was itself removed from ChatGPT in early August 2026 without advance notice. For a business process, that history argues for building on the API layer you control rather than on any one hosted assistant.
Every tool in this comparison says, in its own documentation, that it will not keep a hostile page or a confused model away from your data. The controls have to sit in your environment. A minimal checklist for a supplier-portal agent:
Treat everything on screen as untrusted input. OpenAI's guide states that text in a page, document or tool result cannot grant permission or override the user's instructions. A supplier portal can carry a note telling the agent to email invoices elsewhere, and only a loop that you control can refuse it.
For the accounts-payable scenario, the choice follows from how known the path is and who will maintain it. A reasonable split:
Whichever runs the browser, keep writes away from it. The agent reads the portal and hands structured data to your own code, which creates a draft in the ERP for a person to post. That keeps the browser agent replaceable, which matters in a category where Stagehand dropped Playwright and its agent method within a single major version.
The useful rule is to pick by loop ownership, not by benchmark headlines. Give the loop to Browser Use when the path is unknown, keep it yourself with Stagehand when the path repeats, lend tools to an existing agent with Playwright MCP, and reach for screenshots only when there is no DOM worth reading. In every case, the security boundary is yours to build.