agent-browser ยท diff

git:20260121.3a8e21e to git:20260207.1f08065

268 added, 137 removed. Audit A to A.

---
name: agent-browser
- description: Browser automation using Vercel's agent-browser CLI. Use when you need to interact with web pages, fill forms, take screenshots, or scrape data. Alternative to Playwright MCP - uses Bash commands with ref-based element selection. Triggers on "browse website", "fill form", "click button", "take screenshot", "scrape page", "web automation".
+ description: Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, test web applications, or extract information from web pages.
+ allowed-tools: Bash(agent-browser:*)
---
- # agent-browser: CLI Browser Automation
+ # Browser Automation with agent-browser
- Vercel's headless browser automation CLI designed for AI agents. Uses ref-based selection (@e1, @e2) from accessibility snapshots.
+ ## Quick start
- ## Setup Check
+ ```bash
+ agent-browser open <url> # Navigate to page
+ agent-browser snapshot -i # Get interactive elements with refs
+ agent-browser click @e1 # Click element by ref
+ agent-browser fill @e2 "text" # Fill input by ref
+ agent-browser close # Close browser
+ ```
+ ## Core workflow
+
+ 1. Navigate: `agent-browser open <url>`
+ 2. Snapshot: `agent-browser snapshot -i` (returns elements with refs like `@e1`, `@e2`)
+ 3. Interact using refs from the snapshot
+ 4. Re-snapshot after navigation or significant DOM changes
+
+ ## Commands
+
+ ### Navigation
+
```bash
- # Check installation
- command -v agent-browser >/dev/null 2>&1 && echo "Installed" || echo "NOT INSTALLED - run: npm install -g agent-browser && agent-browser install"
+ agent-browser open <url> # Navigate to URL (aliases: goto, navigate)
+ # Supports: https://, http://, file://, about:, data://
+ # Auto-prepends https:// if no protocol given
+ agent-browser back # Go back
+ agent-browser forward # Go forward
+ agent-browser reload # Reload page
+ agent-browser close # Close browser (aliases: quit, exit)
+ agent-browser connect 9222 # Connect to browser via CDP port
```
- ### Install if needed
+ ### Snapshot (page analysis)
```bash
- npm install -g agent-browser
- agent-browser install # Downloads Chromium
+ agent-browser snapshot # Full accessibility tree
+ agent-browser snapshot -i # Interactive elements only (recommended)
+ agent-browser snapshot -c # Compact output
+ agent-browser snapshot -d 3 # Limit depth to 3
+ agent-browser snapshot -s "#main" # Scope to CSS selector
```
- ## Core Workflow
+ ### Interactions (use @refs from snapshot)
- **The snapshot + ref pattern is optimal for LLMs:**
+ ```bash
+ agent-browser click @e1 # Click
+ agent-browser dblclick @e1 # Double-click
+ agent-browser focus @e1 # Focus element
+ agent-browser fill @e2 "text" # Clear and type
+ agent-browser type @e2 "text" # Type without clearing
+ agent-browser press Enter # Press key (alias: key)
+ agent-browser press Control+a # Key combination
+ agent-browser keydown Shift # Hold key down
+ agent-browser keyup Shift # Release key
+ agent-browser hover @e1 # Hover
+ agent-browser check @e1 # Check checkbox
+ agent-browser uncheck @e1 # Uncheck checkbox
+ agent-browser select @e1 "value" # Select dropdown option
+ agent-browser select @e1 "a" "b" # Select multiple options
+ agent-browser scroll down 500 # Scroll page (default: down 300px)
+ agent-browser scrollintoview @e1 # Scroll element into view (alias: scrollinto)
+ agent-browser drag @e1 @e2 # Drag and drop
+ agent-browser upload @e1 file.pdf # Upload files
+ ```
- 1. **Navigate** to URL
- 2. **Snapshot** to get interactive elements with refs
- 3. **Interact** using refs (@e1, @e2, etc.)
- 4. **Re-snapshot** after navigation or DOM changes
+ ### Get information
```bash
- # Step 1: Open URL
- agent-browser open https://example.com
+ agent-browser get text @e1 # Get element text
+ agent-browser get html @e1 # Get innerHTML
+ agent-browser get value @e1 # Get input value
+ agent-browser get attr @e1 href # Get attribute
+ agent-browser get title # Get page title
+ agent-browser get url # Get current URL
+ agent-browser get count ".item" # Count matching elements
+ agent-browser get box @e1 # Get bounding box
+ agent-browser get styles @e1 # Get computed styles (font, color, bg, etc.)
+ ```
- # Step 2: Get interactive elements with refs
- agent-browser snapshot -i --json
+ ### Check state
- # Step 3: Interact using refs
- agent-browser click @e1
- agent-browser fill @e2 "search query"
+ ```bash
+ agent-browser is visible @e1 # Check if visible
+ agent-browser is enabled @e1 # Check if enabled
+ agent-browser is checked @e1 # Check if checked
+ ```
- # Step 4: Re-snapshot after changes
- agent-browser snapshot -i
+ ### Screenshots & PDF
+
+ ```bash
+ agent-browser screenshot "body" path.png # Viewport screenshot to file
+ agent-browser screenshot --full "body" path.png # Full page to file
+ agent-browser screenshot "body" # Viewport screenshot (auto-saved to tmp)
+ agent-browser screenshot ".modal" modal.png # Screenshot specific element
+ agent-browser pdf output.pdf # Save as PDF
```
- ## Key Commands
+ > **v0.9.1 gotcha**: `screenshot` requires a CSS selector as the first positional arg.
+ > Use `"body"` for full-viewport captures. Omitting the selector causes
+ > `Validation error: selector: Expected string, received null`.
- ### Navigation
+ ### Video recording
```bash
- agent-browser open <url> # Navigate to URL
- agent-browser back # Go back
- agent-browser forward # Go forward
- agent-browser reload # Reload page
- agent-browser close # Close browser
+ agent-browser record start ./demo.webm # Start recording (uses current URL + state)
+ agent-browser click @e1 # Perform actions
+ agent-browser record stop # Stop and save video
+ agent-browser record restart ./take2.webm # Stop current + start new recording
```
- ### Snapshots (Essential for AI)
+ Recording creates a fresh context but preserves cookies/storage from your session. If no URL is provided, it
+ automatically returns to your current page. For smooth demos, explore first, then start recording.
+ ### Wait
+
```bash
- agent-browser snapshot # Full accessibility tree
- agent-browser snapshot -i # Interactive elements only (recommended)
- agent-browser snapshot -i --json # JSON output for parsing
- agent-browser snapshot -c # Compact (remove empty elements)
- agent-browser snapshot -d 3 # Limit depth
+ agent-browser wait @e1 # Wait for element
+ agent-browser wait 2000 # Wait milliseconds
+ agent-browser wait --text "Success" # Wait for text (or -t)
+ agent-browser wait --url "**/dashboard" # Wait for URL pattern (or -u)
+ agent-browser wait --load networkidle # Wait for network idle (or -l)
+ agent-browser wait --fn "window.ready" # Wait for JS condition (or -f)
```
- ### Interactions
+ ### Mouse control
```bash
- agent-browser click @e1 # Click element
- agent-browser dblclick @e1 # Double-click
- agent-browser fill @e1 "text" # Clear and fill input
- agent-browser type @e1 "text" # Type without clearing
- agent-browser press Enter # Press key
- agent-browser hover @e1 # Hover element
- agent-browser check @e1 # Check checkbox
- agent-browser uncheck @e1 # Uncheck checkbox
- agent-browser select @e1 "option" # Select dropdown option
- agent-browser scroll down 500 # Scroll (up/down/left/right)
- agent-browser scrollintoview @e1 # Scroll element into view
+ agent-browser mouse move 100 200 # Move mouse
+ agent-browser mouse down left # Press button
+ agent-browser mouse up left # Release button
+ agent-browser mouse wheel 100 # Scroll wheel
```
- ### Get Information
+ ### Semantic locators (alternative to refs)
```bash
- agent-browser get text @e1 # Get element text
- agent-browser get html @e1 # Get element HTML
- agent-browser get value @e1 # Get input value
- agent-browser get attr href @e1 # Get attribute
- agent-browser get title # Get page title
- agent-browser get url # Get current URL
- agent-browser get count "button" # Count matching elements
+ agent-browser find role button click --name "Submit"
+ agent-browser find text "Sign In" click
+ agent-browser find text "Sign In" click --exact # Exact match only
+ agent-browser find label "Email" fill "user@test.com"
+ agent-browser find placeholder "Search" type "query"
+ agent-browser find alt "Logo" click
+ agent-browser find title "Close" click
+ agent-browser find testid "submit-btn" click
+ agent-browser find first ".item" click
+ agent-browser find last ".item" click
+ agent-browser find nth 2 "a" hover
```
- ### Screenshots & PDFs
+ ### Browser settings
```bash
- agent-browser screenshot # Viewport screenshot
- agent-browser screenshot --full # Full page
- agent-browser screenshot output.png # Save to file
- agent-browser screenshot --full output.png # Full page to file
- agent-browser pdf output.pdf # Save as PDF
+ agent-browser set viewport 1920 1080 # Set viewport size
+ agent-browser set device "iPhone 14" # Emulate device
+ agent-browser set geo 37.7749 -122.4194 # Set geolocation (alias: geolocation)
+ agent-browser set offline on # Toggle offline mode
+ agent-browser set headers '{"X-Key":"v"}' # Extra HTTP headers
+ agent-browser set credentials user pass # HTTP basic auth (alias: auth)
+ agent-browser set media dark # Emulate color scheme
+ agent-browser set media light reduced-motion # Light mode + reduced motion
```
- ### Wait
+ ### Cookies & Storage
```bash
- agent-browser wait @e1 # Wait for element
- agent-browser wait 2000 # Wait milliseconds
- agent-browser wait "text" # Wait for text to appear
+ agent-browser cookies # Get all cookies
+ agent-browser cookies set name value # Set cookie
+ agent-browser cookies clear # Clear cookies
+ agent-browser storage local # Get all localStorage
+ agent-browser storage local key # Get specific key
+ agent-browser storage local set k v # Set value
+ agent-browser storage local clear # Clear all
```
- ## Semantic Locators (Alternative to Refs)
+ ### Network
```bash
- agent-browser find role button click --name "Submit"
- agent-browser find text "Sign up" click
- agent-browser find label "Email" fill "user@example.com"
- agent-browser find placeholder "Search..." fill "query"
+ agent-browser network route <url> # Intercept requests
+ agent-browser network route <url> --abort # Block requests
+ agent-browser network route <url> --body '{}' # Mock response
+ agent-browser network unroute [url] # Remove routes
+ agent-browser network requests # View tracked requests
+ agent-browser network requests --filter api # Filter requests
```
- ## Sessions (Parallel Browsers)
+ ### Tabs & Windows
```bash
- # Run multiple independent browser sessions
- agent-browser --session browser1 open https://site1.com
- agent-browser --session browser2 open https://site2.com
+ agent-browser tab # List tabs
+ agent-browser tab new [url] # New tab
+ agent-browser tab 2 # Switch to tab by index
+ agent-browser tab close # Close current tab
+ agent-browser tab close 2 # Close tab by index
+ agent-browser window new # New window
+ ```
- # List active sessions
- agent-browser session list
+ ### Frames
+
+ ```bash
+ agent-browser frame "#iframe" # Switch to iframe
+ agent-browser frame main # Back to main frame
```
- ## Examples
+ ### Dialogs
- ### Login Flow
+ ```bash
+ agent-browser dialog accept [text] # Accept dialog
+ agent-browser dialog dismiss # Dismiss dialog
+ ```
+ ### JavaScript
+
```bash
- agent-browser open https://app.example.com/login
- agent-browser snapshot -i
- # Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Sign in" [ref=e3]
- agent-browser fill @e1 "user@example.com"
- agent-browser fill @e2 "password123"
- agent-browser click @e3
- agent-browser wait 2000
- agent-browser snapshot -i # Verify logged in
+ agent-browser eval "document.title" # Run JavaScript
```
- ### Search and Extract
+ ## Global options
```bash
- agent-browser open https://news.ycombinator.com
- agent-browser snapshot -i --json
- # Parse JSON to find story links
- agent-browser get text @e12 # Get headline text
- agent-browser click @e12 # Click to open story
+ agent-browser --session <name> ... # Isolated browser session
+ agent-browser --json ... # JSON output for parsing
+ agent-browser --headed ... # Show browser window (not headless)
+ agent-browser --full ... # Full page screenshot (-f)
+ agent-browser --cdp <port> ... # Connect via Chrome DevTools Protocol
+ agent-browser -p <provider> ... # Cloud browser provider (--provider)
+ agent-browser --proxy <url> ... # Use proxy server
+ agent-browser --headers <json> ... # HTTP headers scoped to URL's origin
+ agent-browser --executable-path <p> # Custom browser executable
+ agent-browser --extension <path> ... # Load browser extension (repeatable)
+ agent-browser --help # Show help (-h)
+ agent-browser --version # Show version (-V)
+ agent-browser <command> --help # Show detailed help for a command
```
- ### Form Filling
+ ### Proxy support
```bash
- agent-browser open https://forms.example.com
+ agent-browser --proxy http://proxy.com:8080 open example.com
+ agent-browser --proxy http://user:pass@proxy.com:8080 open example.com
+ agent-browser --proxy socks5://proxy.com:1080 open example.com
+ ```
+
+ ## Environment variables
+
+ ```bash
+ AGENT_BROWSER_SESSION="mysession" # Default session name
+ AGENT_BROWSER_EXECUTABLE_PATH="/path/chrome" # Custom browser path
+ AGENT_BROWSER_EXTENSIONS="/ext1,/ext2" # Comma-separated extension paths
+ AGENT_BROWSER_PROVIDER="browserbase" # Cloud browser provider
+ AGENT_BROWSER_STREAM_PORT="9223" # WebSocket streaming port
+ AGENT_BROWSER_HOME="/path/to/agent-browser" # Custom install location (for daemon.js)
+ ```
+
+ ## Example: Form submission
+
+ ```bash
+ agent-browser open https://example.com/form
agent-browser snapshot -i
- agent-browser fill @e1 "John Doe"
- agent-browser fill @e2 "john@example.com"
- agent-browser select @e3 "United States"
- agent-browser check @e4 # Agree to terms
- agent-browser click @e5 # Submit button
- agent-browser screenshot confirmation.png
+ # Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]
+
+ agent-browser fill @e1 "user@example.com"
+ agent-browser fill @e2 "password123"
+ agent-browser click @e3
+ agent-browser wait --load networkidle
+ agent-browser snapshot -i # Check result
```
- ### Debug Mode
+ ## Example: Authentication with saved state
```bash
- # Run with visible browser window
- agent-browser --headed open https://example.com
- agent-browser --headed snapshot -i
- agent-browser --headed click @e1
+ # Login once
+ agent-browser open https://app.example.com/login
+ agent-browser snapshot -i
+ agent-browser fill @e1 "username"
+ agent-browser fill @e2 "password"
+ agent-browser click @e3
+ agent-browser wait --url "**/dashboard"
+ agent-browser state save auth.json
+
+ # Later sessions: load saved state
+ agent-browser state load auth.json
+ agent-browser open https://app.example.com/dashboard
```
- ## JSON Output
+ ## Sessions (parallel browsers)
- Add `--json` for structured output:
+ ```bash
+ agent-browser --session test1 open site-a.com
+ agent-browser --session test2 open site-b.com
+ agent-browser session list
+ ```
+ ## JSON output (for parsing)
+
+ Add `--json` for machine-readable output:
+
```bash
agent-browser snapshot -i --json
+ agent-browser get text @e1 --json
```
- Returns:
- ```json
- {
- "success": true,
- "data": {
- "refs": {
- "e1": {"name": "Submit", "role": "button"},
- "e2": {"name": "Email", "role": "textbox"}
- },
- "snapshot": "- button \"Submit\" [ref=e1]\n- textbox \"Email\" [ref=e2]"
- }
- }
+ ## Debugging
+
+ ```bash
+ agent-browser --headed open example.com # Show browser window
+ agent-browser --cdp 9222 snapshot # Connect via CDP port
+ agent-browser connect 9222 # Alternative: connect command
+ agent-browser console # View console messages
+ agent-browser console --clear # Clear console
+ agent-browser errors # View page errors
+ agent-browser errors --clear # Clear errors
+ agent-browser highlight @e1 # Highlight element
+ agent-browser trace start # Start recording trace
+ agent-browser trace stop trace.zip # Stop and save trace
+ agent-browser record start ./debug.webm # Record video from current page
+ agent-browser record stop # Save recording
```
- ## vs Playwright MCP
+ ## Deep-dive documentation
- | Feature | agent-browser (CLI) | Playwright MCP |
- |---------|---------------------|----------------|
- | Interface | Bash commands | MCP tools |
- | Selection | Refs (@e1) | Refs (e1) |
- | Output | Text/JSON | Tool responses |
- | Parallel | Sessions | Tabs |
- | Best for | Quick automation | Tool integration |
+ For detailed patterns and best practices, see:
- Use agent-browser when:
- - You prefer Bash-based workflows
- - You want simpler CLI commands
- - You need quick one-off automation
+ | Reference | Description |
+ |-----------|-------------|
+ | [references/snapshot-refs.md](references/snapshot-refs.md) | Ref lifecycle, invalidation rules, troubleshooting |
+ | [references/session-management.md](references/session-management.md) | Parallel sessions, state persistence, concurrent scraping |
+ | [references/authentication.md](references/authentication.md) | Login flows, OAuth, 2FA handling, state reuse |
+ | [references/video-recording.md](references/video-recording.md) | Recording workflows for debugging and documentation |
+ | [references/proxy-support.md](references/proxy-support.md) | Proxy configuration, geo-testing, rotating proxies |
- Use Playwright MCP when:
- - You need deep MCP tool integration
- - You want tool-based responses
- - You're building complex automation
+ ## Ready-to-use templates
+
+ Executable workflow scripts for common patterns:
+
+ | Template | Description |
+ |----------|-------------|
+ | [templates/form-automation.sh](templates/form-automation.sh) | Form filling with validation |
+ | [templates/authenticated-session.sh](templates/authenticated-session.sh) | Login once, reuse state |
+ | [templates/capture-workflow.sh](templates/capture-workflow.sh) | Content extraction with screenshots |
+
+ Usage:
+ ```bash
+ ./templates/form-automation.sh https://example.com/form
+ ./templates/authenticated-session.sh https://app.example.com/login
+ ./templates/capture-workflow.sh https://example.com ./output
+ ```