browser-use/browser-use

Open-source Python library and CLI that lets AI agents control a local or cloud browser, with an optional hosted agent API and BU2 browser-optimized model.

  • 117k GitHub stars
  • Python
  • ⚖️ MIT
  • 🎯 Intermediate
uv add browser-use
browser-use/browser-use preview image

What it is

Browser Use is an open-source browser agent that lets an LLM navigate and act on the web. It is offered three ways: a fully hosted cloud (agent plus browser), a CLI that gives an existing coding agent browser access, and a Python library (Python >= 3.11) for running the agent locally with your choice of model. The CLI and library can each connect to a local or cloud browser.

Who it's for

  • Developers who want to run a browser-controlling AI agent from their own Python code
  • Users of coding agents such as Claude Code, Codex, Hermes or OpenClaw who want to give them browser access
  • Teams that want hosted agents and stealth cloud browsers without managing the infrastructure
  • Developers building Claude-driven browser workflows with the Anthropic SDK

Requirements

Requirements

  • Python >= 3.11 for the Python library
  • uv (the source uses it to install and run)
  • An LLM API key such as OPENAI_API_KEY, or BROWSER_USE_API_KEY for the BU2 model or cloud browser
  • Toolsets for Claude: an Anthropic SDK version that includes anthropic.tools.browser
  • Bash in the Claude integration requires a Linux or macOS host with /bin/bash (WSL on Windows)

Setup

  1. Install the library

    Install Browser Use with uv (Python >= 3.11). If starting a new project, run uv init --python 3.12 first.

    bash
    uv add browser-use
  2. Add API keys to .env

    Add your OpenAI API key. The Browser Use API key is optional and is needed for the BU2 model or a cloud browser.

    bash
    # .env
    OPENAI_API_KEY=your-key
    # BROWSER_USE_API_KEY=your-key  # Optional: BU2 model or cloud browser
  3. Run the agent script

    After saving the quickstart code as agent.py, run it. The agent opens a browser, looks up the repository, and prints its answer.

    bash
    uv run agent.py

Examples

Quickstart agent

python
python
import asyncio

from browser_use import Agent, Browser, ChatBrowserUse, ChatOpenAI
from dotenv import load_dotenv

load_dotenv()

async def main():
    llm = ChatOpenAI(model='gpt-5.6-luna', reasoning_effort='xhigh')
    # llm = ChatBrowserUse(model='bu-2-0')  # Use BU2 instead; requires BROWSER_USE_API_KEY
    agent = Agent(
        task="Find the number of stars of the browser-use repo",
        llm=llm,
        # browser=Browser(use_cloud=True),  # Use a cloud browser; requires BROWSER_USE_API_KEY
    )
    history = await agent.run()
    print(history.final_result())

if __name__ == "__main__":
    asyncio.run(main())

What it does: Runs an agent with an OpenAI model to find the repo's star count. Comments show how to switch to BU2 or a cloud browser.

CLI setup prompt for your coding agent

Prompt
prompt
Install or upgrade browser-use to the latest stable version with uv using Python 3.12, run `browser-use skill install` to register the skill, and connect it to my browser. If setup or connection fails, follow https://github.com/browser-use/browser-harness/blob/main/install.md.

Expected output: Paste into Claude Code, Codex, Hermes, OpenClaw, or another agent to set up the CLI path.

Custom tool

python
python
import asyncio
from datetime import datetime, timezone

from browser_use import ActionResult, Agent, ChatBrowserUse, Tools
from dotenv import load_dotenv

load_dotenv()
tools = Tools()

@tools.action(description='Get the current date and time in UTC.')
def get_current_time() -> ActionResult:
    return ActionResult(extracted_content=datetime.now(timezone.utc).isoformat())

async def main():
    agent = Agent(
        task="What is the current UTC time?",
        llm=ChatBrowserUse(model='bu-2-0'),
        tools=tools,
    )
    history = await agent.run()
    print(history.final_result())

if __name__ == "__main__":
    asyncio.run(main())

What it does: Registers a Python function as an agent tool via Tools and passes it to the Agent. Uses BROWSER_USE_API_KEY from .env.

Using Claude or Gemini through ChatBrowserUse

python
python
from browser_use import Agent, ChatBrowserUse

llm = ChatBrowserUse(model='anthropic/claude-sonnet-4-6')  # or 'google/gemini-3-pro'
agent = Agent(task='...', llm=llm)

What it does: ChatBrowserUse accepts provider-prefixed model IDs through the Browser Use gateway, using BROWSER_USE_API_KEY.

Toolsets for Claude

python
python
import os

from anthropic import AsyncAnthropic
from browser_use.integrations.toolsets_for_claude import Bash, BrowserUse

task = 'Open example.com and report its page title.'
driver = BrowserUse()  # Or BrowserUse(use_cloud=True)
bash = Bash(output_dir='outputs')

async with driver, AsyncAnthropic() as client:
    runner = client.beta.messages.tool_runner(
        model=os.environ['ANTHROPIC_MODEL'],
        max_tokens=32_768,
        max_iterations=100,
        tools=[driver, bash],
        messages=[{'role': 'user', 'content': task}],
    )
    result = await runner.until_done()

What it does: Uses Browser Use as the driver for Claude's browser toolset, plus Bash. The snippet runs inside an async function.

Pros & cons

Pros

  • Pro:Three adoption paths (hosted cloud, CLI, Python library) cover different levels of infrastructure control
  • Pro:Model choice is flexible: ChatOpenAI, ChatAnthropic, ChatGoogle, ChatBrowserUse, or local models via Ollama
  • Pro:Supports custom tools, custom system prompt extension or override, and local or cloud browsers
  • Pro:The library is MIT-licensed and free to use

Cons

  • Con:Model inference, ChatBrowserUse and hosted browsers are paid separately from the free library
  • Con:No browser configuration guarantees CAPTCHAs can be avoided or solved; results depend on the site
  • Con:Cloud profile sync transfers cookies only, not local storage, IndexedDB or extensions, so some sites may require signing in again

Images