Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Jev judgment modelJev browser agentbrowser.com Jev ultrafast

What Is Jev? The Judgment Model Making Browser Agents Fast

Jev is a cheap, instant judgment model that speeds up browser agents by classifying page elements instead of reasoning through them.

Edited by Luis Chavez-Mattos, Director of Product RSS
What Is Jev? The Judgment Model Making Browser Agents Fast

What is Jev?

Jev is a judgment model, not a language model. It doesn’t write text, hold a conversation, or generate paragraphs. Instead, it takes in a block of text (called the “state”) along with a question or list of questions, and it returns an answer chosen from a fixed set of options, plus a confidence score for that answer. Because it never generates free text, it can respond in a fraction of a second and at a fraction of the cost of a model like Claude. It was built by a founder who worked on the methods behind ChatGPT at OpenAI, and its narrow job description is exactly what makes it useful for browser automation.

TL;DR

  • Jev answers questions, it doesn’t write text. It classifies input data against yes/no, multiple-choice, or scored questions and returns a confidence level with each answer.
  • Speed comes from skipping reasoning entirely. There are no screenshots, no pixel analysis, and no generative thinking, just numbered page elements matched against a predefined question set.
  • Browser agents built on Jev have reported decisions in roughly a third of a second per click, letting agents work through dozens of page interactions in the time a typical LLM-driven agent takes to process one.
  • browser.com rebuilt its agent around Jev and Jev Ultrafast, an open-source framework for driving Chrome, and completed a Google Flights search in 7 seconds while cutting the number of browser calls from roughly 1,092 down to about 101 for that task.
  • Jev doesn’t replace Claude or other LLMs. It’s weak at reasoning and can’t produce text, so it handles the thousands of small “what do I click next” decisions while Claude handles planning, typing content, and verifying that a task actually finished.
  • The combined pattern is plan with Claude, classify with Jev, verify with Claude. When Jev’s confidence drops too low or it reports being “blocked,” control passes back to an LLM to reason through the problem.
  • It’s not failure-proof. Cookie banners, iframes, file uploads, false “done” signals, and limited memory of past actions are known weak points that need extra guardrails.
Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

How does Jev actually work?

Jev takes two inputs: a state (the text data describing what’s happening) and a question set. The question set comes in three shapes:

  • A null (yes/no): used when the answer determines whether to take an action.
  • A choice (pick one from a list): used when the answer determines which path or team to route to.
  • A score (a described scale): used when the answer determines how to rank or sort a set of items.

For every question, Jev returns an answer plus a confidence score. There’s no generated prose anywhere in that output, which is why it can run so cheaply: input tokens cost a few cents per million, and output is effectively free because the “output” is just a label and a probability, not generated text.

That narrowness is the whole point. A general-purpose LLM has to read everything on a page, reason about what matters, and write out a decision in natural language before any code can act on it. Jev skips the writing and the open-ended reasoning. It just classifies.

How does Jev make browser agents faster?

The trick is in how the web page gets turned into Jev’s input. A tool like Playwright, or the open-source Jev Ultrafast framework built alongside browser.com, drives Chrome and records everything interactable on the page: links, buttons, text inputs, dropdowns. Each one gets a number. A page might produce up to 100 numbered elements.

That numbered list, plus the user’s goal, gets sent to Jev with two choice questions: what action type should happen (click, type, select, scroll up, scroll down, wait, done, or blocked), and which numbered element that action applies to (with an explicit “none” option if nothing fits).

Because the input is structured numbered data rather than a screenshot, there’s no image processing and no open-ended generation. Jev just matches the goal against the options and returns an action type, a target element, and a confidence score for each. If both confidence scores clear a set threshold (a value like 0.7 was used as an example), the browser automation layer executes the action directly, no AI involved at that step. It’s pure deterministic code: click, wait briefly for the page to settle, re-read the new page state, and send the next question set to Jev.

This loop, plan once with an LLM, then classify-and-execute repeatedly with Jev, is why the call count collapses. browser.com’s own reported comparison on a Google Flights search showed total browser calls dropping from around 1,092 to roughly 101 using this combined approach, about a tenth of the previous volume, while still completing the task in seconds rather than minutes.

Why combine Jev with Claude instead of using either alone?

Jev and Claude solve different problems. Jev is fast and cheap but can’t reason well and can’t produce text. Claude (or another capable LLM) can reason and write, but is slower and more expensive to call on every single micro-decision a browser agent makes.

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

The combined architecture splits the work along those lines:

  • Claude plans. It interprets the overall goal and sets up the task.
  • Jev classifies. For each page, it answers the “what to click” and “where to click” questions quickly and cheaply, looping through as many pages as the task requires.
  • Claude intervenes when needed. If Jev’s confidence drops below the threshold, or if Jev reports the page is “blocked,” the loop stops and hands control to Claude. Claude reasons about what happened and either reformulates the question set for Jev or takes over that step directly.
  • Claude writes. Jev can identify that a text box needs input, but it can’t generate what to type. If the content isn’t already present in the stated goal, Claude supplies the actual text.
  • Claude verifies completion. When Jev reports “done,” that signal isn’t trusted automatically. An independent check, run through Claude, confirms the goal was actually achieved. Browser-use’s own documentation also calls for this kind of independent verification of “done” states.

In short, Jev handles volume; Claude handles judgment calls and language. That division is what lets a browser agent make thousands of small decisions per task without thousands of expensive, slow LLM calls.

Is Jev worth using, and where does it fall short?

Jev-based agents are meaningfully faster than screenshot-and-reasoning agents, but they aren’t flawless out of the box. Several recurring problems showed up in testing:

  • Cookie banners and consent popups can trip up the flow unless the question set explicitly asks whether a given element is a consent banner and instructs the script to accept it.
  • Iframes and embedded elements, like payment forms, signup widgets inside embedded frames, or file upload buttons, are often invisible to the Playwright-style scraping step, meaning Jev never sees them as options. Browser-use’s own team lists these as current limitations.
  • False “done” signals can occur, where Jev reports a task complete even though the actual target (like a specific blog post) was never reached. This is specifically why a Claude-based verification step matters before trusting the result.
  • Limited memory means the agent may only retain awareness of a handful of its most recent actions, which can cause it to lose track of progress on longer, multi-step tasks.

None of these are fatal, but they mean a working Jev-based agent requires iteration: adding questions, tightening the confidence thresholds, and layering in explicit instructions for known edge cases rather than expecting it to infer everything from a vague goal statement.

Frequently Asked Questions

What is Jev used for in AI browser agents?

Jev is used to decide, at each step of a browser task, what action to take (click, type, scroll, wait) and on which page element to take it. It replaces slower, screenshot-based reasoning with fast classification over structured page data.

Is Jev the same as an LLM like Claude or GPT?

Other agents start typing. Remy starts asking.

YOU SAID "Build me a sales CRM."
01 DESIGN Should it feel like Linear, or Salesforce?
02 UX How do reps move deals — drag, or dropdown?
03 ARCH Single team, or multi-org with permissions?

Scoping, trade-offs, edge cases — the real work. Before a line of code.

No. Jev only returns classifications and confidence scores from a predefined set of options, it cannot generate free text or hold a conversation. It’s typically paired with an LLM like Claude, which handles planning, writing text input, and verifying results.

Why is Jev cheaper and faster than using Claude for browser automation?

Jev skips image processing and open-ended text generation entirely. It works from numbered lists of page elements and answers narrow questions, which is far less computationally expensive than having an LLM read a screenshot and reason in natural language about what to do next.

What is Jev Ultrafast?

Jev Ultrafast is an open-source framework, built alongside browser.com, for driving a Chrome browser from code, similar in role to Playwright. It’s used to extract page elements and feed them into Jev’s question-answering loop.

Does Jev work reliably on every website?

Not perfectly. Known weak points include cookie consent banners, content inside iframes, file upload interfaces, false “task complete” signals, and limited memory of prior steps on long multi-step tasks. These generally require extra guardrails or additional questions built into the agent’s logic.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.