What it is
Browser Use is an open-source browser agent that lets an LLM navigate and act on the web. It is offered three ways: a fully hosted cloud (agent plus browser), a CLI that gives an existing coding agent browser access, and a Python library (Python >= 3.11) for running the agent locally with your choice of model. The CLI and library can each connect to a local or cloud browser.
Who it's for
- Developers who want to run a browser-controlling AI agent from their own Python code
- Users of coding agents such as Claude Code, Codex, Hermes or OpenClaw who want to give them browser access
- Teams that want hosted agents and stealth cloud browsers without managing the infrastructure
- Developers building Claude-driven browser workflows with the Anthropic SDK
Requirements
Requirements
- Python >= 3.11 for the Python library
- uv (the source uses it to install and run)
- An LLM API key such as OPENAI_API_KEY, or BROWSER_USE_API_KEY for the BU2 model or cloud browser
- Toolsets for Claude: an Anthropic SDK version that includes anthropic.tools.browser
- Bash in the Claude integration requires a Linux or macOS host with /bin/bash (WSL on Windows)
Setup
Install the library
Install Browser Use with uv (Python >= 3.11). If starting a new project, run
uv init --python 3.12first.bashuv add browser-useAdd API keys to .env
Add your OpenAI API key. The Browser Use API key is optional and is needed for the BU2 model or a cloud browser.
bash# .env OPENAI_API_KEY=your-key # BROWSER_USE_API_KEY=your-key # Optional: BU2 model or cloud browserRun the agent script
After saving the quickstart code as agent.py, run it. The agent opens a browser, looks up the repository, and prints its answer.
bashuv run agent.py
Examples
Quickstart agent
pythonimport asyncio
from browser_use import Agent, Browser, ChatBrowserUse, ChatOpenAI
from dotenv import load_dotenv
load_dotenv()
async def main():
llm = ChatOpenAI(model='gpt-5.6-luna', reasoning_effort='xhigh')
# llm = ChatBrowserUse(model='bu-2-0') # Use BU2 instead; requires BROWSER_USE_API_KEY
agent = Agent(
task="Find the number of stars of the browser-use repo",
llm=llm,
# browser=Browser(use_cloud=True), # Use a cloud browser; requires BROWSER_USE_API_KEY
)
history = await agent.run()
print(history.final_result())
if __name__ == "__main__":
asyncio.run(main())What it does: Runs an agent with an OpenAI model to find the repo's star count. Comments show how to switch to BU2 or a cloud browser.
CLI setup prompt for your coding agent
PromptInstall or upgrade browser-use to the latest stable version with uv using Python 3.12, run `browser-use skill install` to register the skill, and connect it to my browser. If setup or connection fails, follow https://github.com/browser-use/browser-harness/blob/main/install.md.Expected output: Paste into Claude Code, Codex, Hermes, OpenClaw, or another agent to set up the CLI path.
Custom tool
pythonimport asyncio
from datetime import datetime, timezone
from browser_use import ActionResult, Agent, ChatBrowserUse, Tools
from dotenv import load_dotenv
load_dotenv()
tools = Tools()
@tools.action(description='Get the current date and time in UTC.')
def get_current_time() -> ActionResult:
return ActionResult(extracted_content=datetime.now(timezone.utc).isoformat())
async def main():
agent = Agent(
task="What is the current UTC time?",
llm=ChatBrowserUse(model='bu-2-0'),
tools=tools,
)
history = await agent.run()
print(history.final_result())
if __name__ == "__main__":
asyncio.run(main())What it does: Registers a Python function as an agent tool via Tools and passes it to the Agent. Uses BROWSER_USE_API_KEY from .env.
Using Claude or Gemini through ChatBrowserUse
pythonfrom browser_use import Agent, ChatBrowserUse
llm = ChatBrowserUse(model='anthropic/claude-sonnet-4-6') # or 'google/gemini-3-pro'
agent = Agent(task='...', llm=llm)What it does: ChatBrowserUse accepts provider-prefixed model IDs through the Browser Use gateway, using BROWSER_USE_API_KEY.
Toolsets for Claude
pythonimport os
from anthropic import AsyncAnthropic
from browser_use.integrations.toolsets_for_claude import Bash, BrowserUse
task = 'Open example.com and report its page title.'
driver = BrowserUse() # Or BrowserUse(use_cloud=True)
bash = Bash(output_dir='outputs')
async with driver, AsyncAnthropic() as client:
runner = client.beta.messages.tool_runner(
model=os.environ['ANTHROPIC_MODEL'],
max_tokens=32_768,
max_iterations=100,
tools=[driver, bash],
messages=[{'role': 'user', 'content': task}],
)
result = await runner.until_done()What it does: Uses Browser Use as the driver for Claude's browser toolset, plus Bash. The snippet runs inside an async function.
Pros & cons
Pros
- Pro:Three adoption paths (hosted cloud, CLI, Python library) cover different levels of infrastructure control
- Pro:Model choice is flexible: ChatOpenAI, ChatAnthropic, ChatGoogle, ChatBrowserUse, or local models via Ollama
- Pro:Supports custom tools, custom system prompt extension or override, and local or cloud browsers
- Pro:The library is MIT-licensed and free to use
Cons
- Con:Model inference, ChatBrowserUse and hosted browsers are paid separately from the free library
- Con:No browser configuration guarantees CAPTCHAs can be avoided or solved; results depend on the site
- Con:Cloud profile sync transfers cookies only, not local storage, IndexedDB or extensions, so some sites may require signing in again
Images
