Clef 27B vs Gemma, Llama: Which Decision Model Actually Wins?
Clef 27B, Cloudflare's multimodal decision model, goes up against Gemma and Llama on security, classification, and image tasks. Here's how it stacks up.

What is Clef 27B and why does it matter?
Clef 27B is a 27 billion parameter decision model from Cloudflare, released under an Apache 2.0 license. Unlike a chatbot, it doesn’t generate text token by token. You feed it a situation (as text, an image, video frames, or JSON), give it a set of typed questions with possible answers, and it returns calibrated probabilities for each option in a single forward pass. No generation overhead, no waiting on streamed tokens, just a fast structured answer.
That puts Clef in the same category as other “decision models” that have been compared head to head before, including Gemma-based variants, Llama-based ones, and a model referred to as “Jev” in earlier testing. What separates Clef from that pack is multimodality. It doesn’t just read text prompts. It can look at an image, watch a short video clip, and still return the same kind of structured, probability-based output.
TL;DR
- Clef 27B is a non-generative decision model that returns calibrated probabilities across multiple questions in one forward pass instead of writing tokens like a chatbot.
- It caught a social engineering test that smaller models failed, flagging a fake “emergency access” request with a 76% social engineering probability and only a 5.9% chance of granting access.
- Multimodality is the real differentiator, with Clef handling static images, AI-generated video clips, foreign-language news screenshots, and a chemistry diagram all through the same probability-output format.
- On a broader benchmark sweep covering tool retrieval, API banking, and clinical classification, Clef reportedly led by wide margins, outperforming Jev, which had been described as the prior gold standard in those categories.
- Jev still wins on “when to escalate to a human”, the one task in the comparison that leans most heavily on human-style judgment rather than pattern classification.
- A model called Lia ranked at the bottom across the tested categories, consistent with earlier results showing it’s capable for its size but outclassed by larger decision models.
- Running the full 27B model locally consumed just under 54GB of VRAM with KV cache included, so this isn’t a model you casually run on a laptop GPU.
One coffee. One working app.
You bring the idea. Remy manages the project.
How does Clef 27B actually make decisions?
Clef’s core mechanic is simple to describe and unusual in practice. You give it a scenario, written as plain text, a JSON payload, an image, or frames extracted from video. Then you attach one or more typed questions, each with a fixed set of possible answers. Clef processes all of it together and returns a probability for every option on every question, computed in a single pass.
A demonstrated example: a support ticket describing an API outage that’s blocking orders. Two questions get asked at once: which team should handle it, and is it urgent? Clef answers both in parallel within the same pass, returning “technical” at 100% confidence and “urgent: yes” at 100% confidence. There’s no sequential generation, no back-and-forth, just two independent answers computed together.
This design matters for anyone building classification or routing systems. A chatbot-style model has to generate an answer as text and then you parse it, which adds latency and the risk of malformed output. A model like Clef skips that step entirely and hands back structured, calibrated numbers you can act on directly.
How did Clef perform on the security test?
The security test used was a recurring one in this line of comparisons: a social engineering scenario where a supposed employee requests emergency access, claims a manager approved it, but offers nothing verifiable. It’s designed to trip up models that default to being helpful rather than skeptical.
In earlier rounds of testing, models described as Lia and OpenJ fell for a version of this trap, while Clef and Jev caught it. Clef returned a 5.9% probability of granting access and a 76% probability that the request was social engineering, with a recommendation to escalate to security. That result puts Clef in the same tier as Jev on this specific task, reinforcing the pattern that larger decision models tend to have sharper instincts for this kind of ambiguous, high-stakes request.
The takeaway for builders: model size still correlates with judgment quality on adversarial inputs. If you’re deploying a decision model anywhere near access control, fraud flags, or approval workflows, the small efficient option isn’t automatically the safe one.
Is Clef 27B worth it for multimodal tasks?
This is where Clef separates itself from the field. In testing, it was pointed at several very different inputs and asked to reason about them using the same probability-output approach:
A staged image of three people in a workplace relationship triangle produced nuanced, if obviously speculative, emotional readings: a 94.5% probability one person was stressed about work rather than the men around her, a 92.7% probability one man “wasn’t in love, just using her,” and only a 47.7% score on a money-motivation question for a second character. It’s not ground truth, but it shows Clef can parse implied social dynamics from a still image, not just objects and text.
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
An 8-second AI-generated video of a man alone in a frozen forest holding a candle got read as a likely ritual or memorial scene (72.2% and 60.2% confidence respectively) with only a 16% danger probability, and Clef correctly inferred the man survives. For a silent clip with no dialogue, that’s a reasonable inference task, not just object detection.
A screenshot of a Uzbek-language news site was correctly identified as Uzbek with 95.9% confidence, correctly tagged the topic as the Ukraine conflict at 97.8%, and correctly read the tone as neutral at 93.1%. It missed slightly on a UN-related detail, scoring that question at 39.6% when it should likely have been higher, but overall handled a language and script most models struggle with.
A university-level chemistry titration curve (a diprotic acid being neutralized with NaOH) got read with high confidence across the board: titration curve type at 99.1%, diprotic acid identification at 97.8%, and equivalence point count at 96.7%.
Taken together, these results suggest Clef isn’t just a text classifier with an image bolted on. It’s handling genuinely different modalities (static images, motion in video, non-Latin script, scientific diagrams) through one consistent decision-making interface.
How does Clef compare to Gemma, Llama and Jev on benchmarks?
Across a broader benchmark sweep covering tool retrieval, API-style banking tasks, and clinical classification, Clef was reported to lead by a wide margin, including ahead of Jev, which had held the top spot in prior comparisons. The one category where Jev still came out ahead was “when to escalate to a human,” a task that depends on softer judgment calls rather than clean classification boundaries.
Lia, a smaller model in the same comparison set, placed at the bottom across these categories. That’s consistent with how it performed in earlier security testing too: solid for its size, but clearly in a different performance tier than the larger 27B-class models.
For builders choosing between decision models, the practical signal here is that raw parameter count still buys you something in this category, particularly on tasks that require picking up on subtle cues (deception, understatement, risk) rather than straightforward pattern matching.
Frequently Asked Questions
What makes Clef 27B different from a normal chatbot model?
Clef doesn’t generate text. It takes a scenario plus a set of typed questions with fixed answer options, and returns a calibrated probability for each option in one forward pass. There’s no token-by-token generation and no parsing of free text output.
Can Clef 27B process images and video, not just text?
Yes. It accepts text, JSON, images, and video frames as input, and applies the same probability-based decision output across all of them. It was tested on a staged relationship photo, an AI-generated video clip, a foreign-language news screenshot, and a chemistry diagram.
How does Clef 27B compare to Jev on security tasks?
On a social engineering test involving a fake emergency access request, Clef and Jev both correctly flagged the risk and recommended escalation, while smaller models in the comparison did not. On a broader benchmark set, Jev reportedly only outperformed Clef on tasks involving when to escalate a decision to a human.
Is Clef 27B open source?
Yes, it’s released under an Apache 2.0 license, making it freely usable and modifiable.
What hardware do you need to run Clef 27B locally?
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
Running the full 27 billion parameter model locally, including KV cache, consumed just under 54GB of VRAM in testing, so it requires a substantial GPU setup rather than consumer-grade hardware.
