Skip to content
Effloow
← Back to Articles
AI INFRASTRUCTURE ARTICLES ·2026-10-04 ·BY EFFLOOW EDITORIAL ·12 MIN READ

Browser Use vs Skyvern Self-Hosted: Setup, Trade-offs, and Completion Verification

We audited Browser Use and Skyvern’s self-hosted installation paths, execution models, failure boundaries, and cost structures, and show why agent-reported success alone does not establish workflow completion.
Browser Automation AI Agents Self-Hosted Playwright Reliability
SHARE
Illustration for Browser Use vs Skyvern Self-Hosted: Setup, Trade-offs, and Completion Verification
Illustration: AI-assisted. Editorial policy

Why we evaluated Browser Use and Skyvern

We did not bring Browser Use and Skyvern into our lab to see whether an agent could click a button in a polished demo. We wanted to answer a more expensive question: can either system complete a recurring login-to-download workflow without creating a permanent human repair queue?

Agent-reported done is not verified completion

A strict schema receipt still passed for a missing and a corrupted file, so a workflow should count as complete only when page state, business identifier, and artifact size and digest are checked independently of the agent's own success claim.

Our target shape was deliberately ordinary:

  1. Open a locally hosted application.
  2. Authenticate with a test account.
  3. complete a three-stage form.
  4. Submit the transaction.
  5. Download the generated file.
  6. Prove that the downloaded artifact is the expected file.

That last condition matters. An agent saying “done” is not evidence that the workflow finished correctly. A successful click is not a successful transaction, and a valid JSON response is not proof that a file exists.

The architectural distinction also turned out to be less absolute than the usual “DOM versus vision” shorthand suggests.

Browser Use builds its page representation primarily from browser structure, including the DOM and accessibility information exposed through Chrome. Screenshots can be disabled, enabled, or requested when needed. That lets us configure Browser Use for DOM-first operation, but it is not DOM-only.

Skyvern’s self-hosted execution loop captures a screenshot, combines visual interpretation with a simplified DOM, asks the configured model for the next action, and executes that action through Playwright. That makes it vision-forward, but not vision-only.

In practical terms, both products run a perception-action loop. Each additional loop can add model latency, input volume, another chance to select the wrong element, and another opportunity for a retry to multiply cost. The meaningful comparison is therefore not whether one product “uses AI.” Both do. We care about:

  • verified workflow completion rate;
  • wall-clock time per verified completion;
  • model requests and input volume per run;
  • retries and terminal failure reasons;
  • human interventions per 100 runs;
  • infrastructure and inference cost per successful task.

We also separated the product shapes. Browser Use is primarily an embeddable MIT-licensed agent library. Skyvern is a larger AGPL-licensed workflow platform with an API server, web interface, database, browser runtime, workflow state, and operational views.

That distinction affects engineering cost more than a synthetic leaderboard score. Browser Use gives us an embeddable component. Skyvern gives us a broader platform for managing browser workflows.

For adjacent infrastructure evaluations, we keep our current shortlist in our tools collection. The rest of this review focuses on what we could verify without pretending that an incomplete run ledger was a production benchmark.

Hands-On Walkthrough: Setup, Execution & Output

We treated installation reproducibility as the first test. If a team cannot recreate the environment from pinned commands, any reliability percentage is suspect.

Browser Use installation path

For Browser Use, we use Python 3.11 or newer and take uv add browser-use as the documented setup path. The following pip-based setup is a proposed reproduction path, not an installation test we completed:

python3.12 -m venv .venv-browser-use
source .venv-browser-use/bin/activate

python -m pip install --upgrade pip
pip install browser-use

# Make the browser dependency explicit instead of relying on first-run behavior.
python -m playwright install chromium

We would make Chromium provisioning an explicit build step rather than assume that package installation downloads a compatible browser binary. We did not test first-run browser provisioning here. Our production image would pin the package, verify the browser installation procedure for that version, cache the browser layer, and fail the build if Chromium could not launch.

The minimal agent shape is straightforward:

import asyncio
from browser_use import Agent, ChatOpenAI

async def main() -> None:
    llm = ChatOpenAI(
        model="your-model",
        base_url="http://127.0.0.1:8000/v1",
        api_key="local-development-key",
    )

    agent = Agent(
        task=(
            "Open http://demo-app:8080, sign in with the supplied test "
            "credentials, complete the three-step request form, download "
            "the receipt, and return its filename and visible request ID."
        ),
        llm=llm,
    )

    history = await agent.run()
    print(history.final_result())

if __name__ == "__main__":
    asyncio.run(main())

A local OpenAI-compatible endpoint removes provider charges, but it does not make inference free. GPU depreciation, power, model-serving memory, queueing, and operator time still belong in the cost model.

Skyvern self-hosted path

For Skyvern, we used the full-stack deployment boundary rather than mistaking the lightweight Python package for the server.

git clone https://github.com/Skyvern-AI/skyvern.git
cd skyvern

cp .env.example .env
# Configure the selected LLM provider or local Ollama endpoint in .env.

docker compose up

The important trap is that plain pip install skyvern is the remote API SDK, not the complete local platform. The server-capable package is:

pip install "skyvern[server]"

For the web interface, API server, browser runtime, and PostgreSQL together, Docker Compose remains the practical starting point. We budgeted at least the documented 4 GB of RAM for evaluation and would follow the recommendation of 8 GB or more for production. The 8 GB figure is a recommendation, not a mandatory deployment minimum or a memory measurement from our workload.

For our Skyvern self-hosting plan, we account for the API server, web UI, PostgreSQL, browser runtime, and a configured LLM provider. Self-hosting removes a per-task platform charge, but we still have to operate the API, browser capacity, database, model endpoint, backups, updates, and any required proxy service.

The completion contract we actually executed

We did not have a defensible N-run vendor trace containing success, duration, model usage, and artifact hashes. Rather than invent those numbers, we tested the most important measurement boundary: whether structured output alone can falsely mark a workflow complete.

We pinned Pydantic 2.10.6 and checked an application-defined strict completion receipt against a real local fixture file. We tested Pydantic validation and our proposed completion contract, not Browser Use or Skyvern integrations. Matching the fixture’s size and digest established only that its bytes agreed with the expected receipt—not that a browser workflow completed or a website delivered the correct business document. The contract and artifact checks were:

from hashlib import sha256
from pathlib import Path
from typing import Literal

from pydantic import BaseModel, ConfigDict

class DownloadReceipt(BaseModel):
    model_config = ConfigDict(strict=True)

    status: Literal["completed"]
    path: str
    byte_count: int
    sha256: str

def verify_artifact(receipt: DownloadReceipt) -> dict:
    artifact = Path(receipt.path)
    if not artifact.exists():
        return {"verified": False, "reason": "missing_artifact"}

    payload = artifact.read_bytes()
    size_matches = len(payload) == receipt.byte_count
    digest_matches = sha256(payload).hexdigest() == receipt.sha256

    return {
        "verified": size_matches and digest_matches,
        "reason": "matched" if size_matches and digest_matches else "artifact_mismatch",
        "observed_bytes": len(payload),
        "size_matches": size_matches,
        "digest_matches": digest_matches,
    }

Our test output was unambiguous:

{
  "pydantic": "2.10.6",
  "genuine": {
    "strict_schema_accepted": true,
    "artifact_verified": true
  },
  "missing_artifact": {
    "strict_schema_accepted": true,
    "artifact_verified": false,
    "reason": "missing_artifact"
  },
  "corrupted_artifact": {
    "strict_schema_accepted": true,
    "artifact_verified": false,
    "reason": "artifact_mismatch"
  },
  "numeric_string": {
    "lax_schema_accepted": true,
    "strict_schema_accepted": false,
    "error_type": "int_type"
  }
}

Strict validation rejected a numeric string where an integer was required. It also rejected missing required fields and an invalid status. However, schema validation could not determine whether the referenced file existed or matched its declared digest.

Our completion numerator must therefore be:

verified_success =
    expected_final_page_state
    AND expected_business_identifier
    AND artifact_exists
    AND artifact_size_matches
    AND artifact_digest_matches

Anything weaker inflates the apparent reliability of both products.

What we could not verify and deployment limitations

The first failure was methodological: we could not produce a valid head-to-head reliability score from the available run evidence.

We had install paths, architectural boundaries, pricing inputs, and a completed validation experiment. We did not have repeated Browser Use and Skyvern executions with comparable model settings, task seeds, token accounting, retries, and artifact verification. Publishing “9/10 versus 8/10” under those conditions would be fiction.

That absence exposed several production traps.

Browser installation is a separate dependency

Installing the Python package is not equivalent to provisioning a runnable Chromium binary. For a reproduction build, we would verify whether python -m playwright install chromium is required by the pinned Browser Use version and add a browser launch smoke test.

Our workaround is simple: pin Browser Use, pin the container base image, install Chromium during the build, and never permit a production worker to download browser assets on first request.

Skyvern’s packages support two different deployment models

pip install skyvern does not create the API server, UI, browser runtime, and database stack. We would use skyvern[server] only for a controlled source-based development path. For team evaluation, we would start with Compose and pin every image digest.

We would also inspect and explicitly configure the maximum step count, task timeout, and retry policy before submitting real work. An unbounded retry loop does not ensure resilience and can make costs unpredictable.

Self-hosting does not automatically keep inference local

Skyvern can send screenshots to a configured remote model on every perception step. Browser Use can likewise send page state and optional screenshots to an external model. Running the orchestrator on our own server does not change that data boundary.

For sensitive workflows, we would either use an approved private endpoint or a local model that had passed the same task suite. We would not assume a smaller local model preserved cloud-model reliability.

CAPTCHA and bot detection remain hard boundaries

For self-hosted Skyvern, CAPTCHA handling can require manual intervention. Browser Use’s managed browser offering adds stealth, proxies, and CAPTCHA services, but no browser configuration guarantees universal success.

That distinction changes the operating model. A workflow that occasionally waits for a person is not unattended automation. We would classify it as “human-assisted” and account for intervention time separately.

The same rule applies to authentication state. Browser Use profile synchronization covers cookies, but not every form of local storage, IndexedDB state, or extension-backed authentication. A supposedly warm session can still return to the login screen.

Retry loops can disguise selector mistakes

An agent can select a plausible but incorrect element, observe no expected transition, and retry with a slightly different action. Without a hard step budget and a terminal-state validator, the run may look active while making no business progress.

Our workaround is to record a failure taxonomy rather than one aggregate error:

Failure class Required evidence
Wrong element or selector hallucination Intended control, selected control, screenshot, DOM reference
Navigation or page timeout URL, elapsed time, timeout boundary
Authentication failure Login state, challenge type, session source
CAPTCHA or bot block Challenge screenshot and intervention requirement
Retry exhaustion Step count, repeated action signature, last changed state
False completion Agent result versus independent page and artifact checks
Artifact failure Missing file, wrong size, or digest mismatch
Model failure Invalid action schema, refusal, or unavailable endpoint

This is the level of instrumentation we would require before involving a customer workflow. Teams needing help building that harness can review our AI infrastructure services or contact us with the target volume and error budget.

Scale, Latency & Cost vs. Alternatives

We cannot honestly report measured success rate, wall-clock time, token count, or dollars per completed task for the full login-to-download workflow. Those values were not established by the executable experiment.

What we can compare is the cost structure and the operational work each option leaves with us.

Decision factor Browser Use OSS Skyvern self-hosted Deterministic Playwright
Primary shape Python agent library Workflow platform and API Browser automation library
License MIT AGPL v3 Apache 2.0
Grounding DOM and accessibility first; optional vision Screenshot plus simplified DOM Explicit selectors and assertions
Initial setup Lowest Highest Low
UI and run history We build it Included in the stack We build it
Database requirement Application choice PostgreSQL in Compose Application choice
Model cost Per agent step Per agent step, often with image input None for scripted steps
Browser infrastructure We operate it or buy managed browsers We operate it We operate it
CAPTCHA in self-hosted mode No universal guarantee Manual intervention may be required External solver or manual path
Artifact verification Application responsibility Workflow plus application checks Direct application checks
Best fit Agent capability embedded in code Operations-owned recurring workflows Stable, high-volume sites
Verified benchmark result here Not established Not established Not run

For managed Browser Use browsers, we budget $0.02 per browser-hour, with model usage and traffic accounted for separately. At that rate, browser time alone is rarely the dominant cost for a short task. Model calls, retries, proxy traffic, failed runs, and repair labor can be much larger.

We use this formula:

cost_per_verified_success =
    total_browser_cost
    + total_model_cost
    + total_proxy_cost
    + allocated_infrastructure_cost
    + human_intervention_cost
    ------------------------------------------------
    verified_success_count

The denominator is the dangerous part. If a task costs $0.04 per attempt but only 80% of attempts pass independent verification, the attempt cost understates reality:

$0.04 / 0.80 = $0.05 per verified success

That example is arithmetic, not our measured vendor result.

For self-hosting, our break-even comparison would be:

monthly_self_hosted_cost =
    compute
    + model inference
    + database and storage
    + proxy service
    + monitoring
    + engineering maintenance

monthly_managed_cost =
    successful task volume
    × managed cost per verified success

If self-hosting costs $1,500 per month after infrastructure and engineering allocation, and a managed service costs $0.15 more per verified workflow, the nominal break-even point is 10,000 verified workflows per month. Reliability differences can move that threshold dramatically.

A deterministic Playwright script remains the baseline we would beat. If we control the target application or its selectors rarely change, adding an LLM to every click usually adds avoidable latency and reliability risks. The better hybrid is to script stable login, navigation, submission, and artifact checks while reserving the agent for ambiguous page interpretation.

Skyvern is more attractive when operations staff need workflow visibility, stored state, conditional blocks, and reruns without editing Python. Browser Use is more attractive when we already own the surrounding queue, credentials, telemetry, and validation layer.

Our Final Verdict: When to Deploy, When to Skip

We would not recommend one product for every use case.

Browser Use is the better engineering component. It is easier to embed, has a permissive license, and lets us control the orchestration around the agent. We would choose it when our team already has production primitives for queues, secrets, traces, retries, browser isolation, and result validation.

Skyvern self-hosted is the better workflow platform. We would choose it when the operational interface, run history, workflow authoring, and centralized task state justify running a larger stack. We would review the AGPL implications before modifying and exposing the system over a network.

Deploy Browser Use if:

  • we want browser control inside an existing Python service;
  • engineers own the workflow and surrounding infrastructure;
  • we can explicitly bootstrap and pin Chromium;
  • we want DOM-first operation with selective vision;
  • we can implement independent completion checks;
  • we are comfortable building our own queue, credential, and observability layers.

Deploy Skyvern self-hosted if:

  • operations teams need a visual workflow and run investigation interface;
  • we accept running an API, UI, browser runtime, and PostgreSQL;
  • we have at least the documented memory floor and a scaling plan;
  • we can configure step ceilings, timeouts, retries, and model routing;
  • we understand that CAPTCHA handling may require a person;
  • the AGPL license is compatible with the intended deployment.

Hold off or avoid both if:

  • the target site has a stable API we can call directly;
  • deterministic Playwright already completes the workflow reliably;
  • every run must finish without human intervention on bot-protected sites;
  • screenshots or page contents cannot leave our network, and we cannot run a suitable model within that boundary;
  • we cannot verify the final business state independently;
  • the business requires a reliability percentage that has not been measured on its own workflow.

Our strongest conclusion is not that vision beats the DOM or that one agent is categorically more reliable. It is that browser-agent success must be externally verified.

The strict Pydantic receipt helped prevent malformed output, but it still accepted a perfectly structured reference to a missing or corrupted file. We only caught those false completions by checking the artifact itself.

That is the production gate we would apply to both Browser Use and Skyvern: no workflow counts as complete merely because the agent says it is. Until repeated runs produce traceable success, latency, usage, intervention, and artifact data, we have a promising automation system—not a reliability benchmark.

Get the next one
in your inbox.

One short weekly dispatch with new guides, tools, and what we tested. No spam, unsubscribe anytime.

Get weekly AI tool reviews & automation tips

Join our newsletter. No spam, unsubscribe anytime.

More in Articles

Tools you can use