Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
build browser agent JevJev Claude Code tutorialPlaywright AI agent

How to Build a Fast Browser Agent with Jev and Claude Code

A practical guide to building a fast browser automation agent using Jev's judgment model, Playwright or Jev Ultrafast, and Claude Code.

Edited by Luis Chavez-Mattos, Director of Product RSS
How to Build a Fast Browser Agent with Jev and Claude Code

What makes a Jev-powered browser agent different?

A Jev-powered browser agent splits the work in two. Jev, a judgment model, makes the fast, cheap, repetitive decisions about what to click or type on a page. Claude Code handles the planning, writing, and verification. Instead of feeding screenshots to a vision model and waiting for it to reason through every click, you give Jev a numbered list of page elements and a tightly scoped question, and it returns an answer plus a confidence score in a fraction of a second. That division of labor is what turns a browser agent from something that takes 10 to 20 minutes on a 2-minute task into something that can complete the same task in seconds.

TL;DR

  • Jev is a judgment model, not a chat model: it takes a text “state” and a question, then returns an answer from a defined set of options along with a confidence score, but it never generates free text.
  • Three question types cover almost everything: a yes/no call (“null”) for deciding whether to act, a “choice” for picking one option from a list (routing decisions), and a “score” for ranking or sorting data.
  • Speed comes from skipping pixels entirely: tools like Playwright or Jev Ultrafast extract every clickable or typeable element on a page, number them, and hand that numbered list to Jev as structured data instead of a screenshot.
  • Claude only steps in when needed: when Jev’s confidence drops below a set threshold (the example used 0.7) or Jev flags the page as blocked, Claude reasons about what happened and either reformulates the question or writes the actual text to type.
  • Verification matters: Jev can report “done” on a page that isn’t actually a completed task, so a final check from Claude against the original goal catches false positives before the agent stops.
  • Real-world efficiency gains are documented: browser.com rebuilt their agent on this approach and cut browser calls for a Google Flights search from 1,092 down to 101, roughly a 90% reduction.
  • Building one yourself is fast: with a Type Safe API key, the Jev skill, and a clear prompt, Claude Code can scaffold a working browser agent in well under 15 minutes, though you’ll still need to patch common failure cases by hand.
REMY IS NOT
  • ✕a coding agent
  • ✕no-code
  • ✕vibe coding
  • ✕a faster Cursor
IT IS
✓a general contractor for software

The one that tells the coding agents what to build.

How does Jev actually decide what to click?

Jev works by taking two inputs: a “state” (the text description of whatever it needs to judge) and a question with a limited set of valid answers. For browser automation, the state is the list of interactive elements on the current page, each assigned a number. A page might have a link to pricing labeled “1,” a free trial button labeled “2,” an email input field labeled “3,” and so on, up to roughly a hundred elements.

Jev then gets asked two choice questions in the same call: what type of action should happen (click, type, select, scroll up, scroll down, wait, done, or blocked), and which numbered element that action should target (with a “none” option if nothing fits). Both answers come back with a confidence score. If both scores clear the threshold you’ve set, the automation layer (Playwright or Jev Ultrafast) executes the action in code, no AI involved at that step, waits for the page to settle, re-reads the new element list, and sends the next question set back to Jev. That loop, read page, ask Jev, act, repeat, is the entire engine. It’s inexpensive (the source cites roughly 4 cents per million tokens in, with output effectively free) and runs fast enough to make thousands of decisions in seconds.

Where does Claude Code fit into the loop?

Claude Code handles everything Jev can’t. Jev is explicitly weak at reasoning and can’t produce written output, so Claude takes over in three situations. First, when Jev’s confidence falls below the acceptance threshold or it reports the page as blocked, Claude reasons about the goal and the current state and either sends Jev a revised question or intervenes directly. Second, when the action type is “type” and the text to enter isn’t already specified in the goal, Jev can identify the right input box but can’t generate what goes inside it, so Claude writes that text. Third, and most importantly, when Jev reports “done,” Claude performs an independent verification pass rather than accepting the claim outright. The browser.com team’s own documentation reportedly flags this same requirement: an agent’s self-reported “done” state needs a separate check before you trust it.

This plan-collect-ask-act-verify sequence repeats continuously. Claude plans, Playwright or Jev Ultrafast collects the page state, Jev answers the structured questions, code executes the resulting action, and Claude periodically verifies progress against the original goal.

How do you actually build one in Claude Code?

Two prerequisites: a Type Safe API key stored in your environment file, and the Type Safe (Jev) skill installed so Claude knows how to phrase and batch questions the way Jev expects. Batching matters for speed, since sending multiple well-formed questions in a single call is what keeps the loop fast.

Plans first. Then code.

PROJECTYOUR APP
SCREENS12
DB TABLES6
BUILT BYREMY
1280 px · TYP.
yourapp.msagent.ai
A · UI · FRONT END

Remy writes the spec, manages the build, and ships the app.

From there, the build is a single detailed prompt. Describe the goal (for example, “navigate pricing pages with annual toggles for SaaS tools and screenshot each one”), specify whether to use Playwright or the Jev Ultrafast library (which skips some manual browser setup), and let Claude Code generate the full plan: connecting to the browser layer, formatting the numbered element state, sending the choice and null questions to Jev, handling the confidence thresholds, and wiring in Claude’s verification step. Running on a capable Claude model, this produced a working browser agent from one prompt in roughly 20 minutes in the source demonstration, closely replicating a public demo originally shown on Twitter without access to its underlying code.

What are the common failure traps?

The process is not plug-and-play. Several recurring issues surfaced during testing, and each required adding more specific questions to the Jev request rather than assuming the agent would generalize.

Cookie banners and consent popups trip up the agent repeatedly unless you explicitly add a question like “is this a cookie or preferences banner, yes or no” and hard-code the accept action when it fires.

Iframes and file uploads are largely invisible to this method. Playwright can’t see inside embedded frames, so payment forms or embedded signup widgets won’t appear in Jev’s numbered element list. Upload buttons have the same blind spot. Even the browser.com team lists these as out of scope for now.

False positives on “done” are a real risk. In one test, an agent was asked to visit ten blog pages across four sites and screenshot the ten most recent posts. It reported completion in 77 seconds with all ten pages marked done, but a review of the screenshots showed most weren’t actual blog posts, just intermediate pages the agent mistook for the target. This is exactly why a Claude-driven verification pass against the original goal, rather than trusting Jev’s self-report, is necessary before the loop terminates.

Limited memory is another constraint: the agent in the source example only retained visibility into its last five actions, which becomes a problem on longer, multi-step tasks where earlier context would help it avoid repeating mistakes.

Is building a Jev browser agent worth it?

For tasks that involve many small, repetitive navigation decisions, routing, cookie handling, clicking through multi-step flows, the speed and cost advantage over a pure vision-model agent is substantial, especially given the documented 10x reduction in browser calls on a flight search benchmark. But it’s not a “set it and forget it” system. You need to anticipate common page patterns (consent banners, iframes, false completion states) and explicitly account for them in your question design. Treat the first build as a working draft, not a finished product, and plan to iterate on the question set as you hit edge cases specific to the sites you’re automating.

Frequently Asked Questions

What is Jev used for in a browser agent?

Jev acts as the fast decision-maker. It reads a numbered list of elements on a web page and answers narrowly scoped yes/no, choice, or score questions about what action to take and which element to target, returning a confidence score with each answer.

Can Jev replace Claude or another LLM entirely?

No. Jev can’t write text, hold a conversation, or reason through ambiguous situations. It’s a classifier, not a generator. Claude (or a similar model) is still needed for planning, writing input text, and verifying that a task is genuinely complete.

Do I need Playwright, or can I use something else?

Playwright is one option for driving the browser and extracting page elements. Jev Ultrafast, an open-source framework associated with browser.com, is an alternative that can reduce setup time while achieving similar results.

Why does this approach reduce the number of browser calls?

Because Jev answers structured questions from text data instead of analyzing screenshots, there’s no image processing or open-ended reasoning on every step. Browser.com’s own reported figures show calls dropping from 1,092 to 101 on a Google Flights search using this combined method.

What’s the biggest risk when building one of these agents?

Trusting a “done” signal without verification. Jev can mark a task complete based on a page that only superficially matches the goal, so a secondary check against the original objective is necessary before the agent stops running.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.