Automate Form and Document Review with Open Image Decision Models
How open image decision models like Image-4B and Jev Omni let RPA bots judge scans, forms, and screenshots without custom training or OCR.

What problem are image decision models actually solving?
Most business processes break down into two kinds of steps: actions and decisions. Robotic process automation (RPA) has spent two decades getting good at the actions, clicking buttons, copying fields, moving files between systems. The decisions are a different story. Questions like “is this form signed?” or “does this photo actually match the refund claim?” have traditionally needed a human to look at the image and judge it, because fixed rules can’t read a page the way a person can. Image decision models are a new class of open vision models built specifically to answer that kind of yes/no or multiple-choice question directly from an image, without custom training for each company or use case.
TL;DR
- Image decision models take a raw image (a scan, PDF page, screenshot, or photo) plus a set of typed questions and return structured answers with confidence scores, no OCR step required.
- RPA has automated the “rectangles” (clicks, copies, data entry) for years, but the “diamonds” (judgment calls) have mostly stayed with humans because older deep learning approaches needed heavy fine-tuning per use case.
- UiPath’s trajectory illustrates the gap between RPA’s promise and reality: the company went public around a $30-40 billion valuation but has since fallen to roughly $6.5 billion in market cap, partly because the underlying models required more customization than customers expected.
- Open models like Image-4B and Jev Omni now rank near the top of image decision benchmarks, and one of them was reportedly built by a single person in 15 days using about $1,200 of rented GPU time.
- No single model wins every question type. In hands-on testing, one model handled classification-style questions (what’s missing?) better, while the other was stronger on direct true/false questions (is it signed?), even on the same document.
- Confidence thresholds let you route uncertain answers to a human or a second model, which turns the old binary of “trust the bot or don’t” into a tunable escalation path.
- The workflow skips OCR entirely. The model reasons over the raw image itself to decide what’s filled in, signed, legible, or missing, then downstream logic (like auto-generated follow-up emails) runs on those structured answers.
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
How do image decision models actually work?
An image decision model takes two inputs: an image and a set of predefined questions, similar in structure to how text-based decision models work. Instead of feeding the model a transcript or extracted text, you feed it the image directly, whether that’s a scanned tax form, a screenshot of a completed web transaction, or a photo a customer uploaded for a refund claim. The model then answers each question with a probability distribution across the possible options.
That last part matters. Rather than returning a flat “yes” or “no,” a well-built image decision model returns a confidence score for every possible answer, including an explicit “unknown” option. That unknown category is a safety valve: instead of forcing a guess when the image is ambiguous or the handwriting is illegible, the model can admit uncertainty, which is far more useful in a production pipeline than a confident wrong answer.
Crucially, there’s no OCR step in this pipeline. The model isn’t reading extracted text, it’s interpreting the image itself, the same way a person would glance at a form and notice a blank signature line or a handwritten date. That’s a meaningful architectural difference from older document-automation approaches that relied on text extraction followed by rule-based logic.
Why hasn’t RPA solved this already?
RPA vendors have been selling automation for judgment calls for years, and the industry got big doing it. UiPath is the clearest example: the company first drew attention around 2018, went public by 2021 at a valuation in the $30-40 billion range, and has since seen its stock fall from the high $70s to roughly $13-14 a share, pushing its market cap under $6.5 billion.
The core issue wasn’t that the product was badly built. It’s that the deep learning techniques underlying most RPA “decision” features needed heavy fine-tuning for each specific company and use case. A model trained to detect signatures on one company’s insurance claim form often didn’t generalize to another company’s tax form or loan application. That meant every new document type was effectively a new modeling project, which undercut the promise of “automate everything” that drove the initial hype and valuations.
Open image decision models aim at exactly that gap: general-purpose classifiers that can handle a new form type reasonably well out of the box, without a bespoke training run for every client.
What does this look like in practice?
A practical demonstration of this approach involved a “form decision inspector” built to review PDFs and images of forms, running them through two open models side by side: Image-4B and Jev Omni, both of which rank near the top of an image-focused decision benchmark. The setup lets you drop in a document (for example, a US tax form) and ask a batch of questions simultaneously: Is the signature field filled in? Is the date of birth present? Is the postal address complete? What type of form is this? How legible is the handwriting?
Both models ran the same form and mostly agreed, but not always. On one test document, one model correctly reported that a signature was missing when asked directly as a classification question (“what’s missing?”) but incorrectly said “signed: true” when asked the same thing as a yes/no question. Rephrasing the field from a true/false check to a multiple-choice classification (“typed name only” vs. “handwritten signature”) fixed the discrepancy for that model. The other model handled both phrasings correctly from the start.
The takeaway from that kind of testing: model choice and question phrasing both matter, and they interact. A model that’s strong at open classification questions might be weaker at binary true/false checks on the exact same visual evidence, even though the underlying image content is identical.
Is this reliable enough to replace human review?
Not uniformly, and that’s by design rather than a flaw. The practical pattern is to set a confidence threshold (for example, 90%) below which the system automatically routes the document to a human reviewer instead of trusting the model’s answer. Because each question returns a probability rather than a flat label, that threshold can be tuned per question. A “is this a tax form?” classification might tolerate a lower confidence bar than “is the signature present?” on a legal document.
This is also where running two models in parallel earns its keep. When both models agree with high confidence, you can act automatically (send a reminder email, route the document, flag it as complete). When they disagree, or when either model returns a low-confidence or “unknown” answer, that’s a natural signal to escalate. In the demo, disagreements between the two models surfaced exactly the cases where human review would add the most value, like a signature field that was ambiguous enough to trip up one model but not the other.
The practical implication for anyone building this kind of pipeline: budget time to test multiple models against your specific document types and question phrasings before trusting any single model’s output blindly. A model that performs well on one benchmark or one kind of question doesn’t automatically generalize to every field you need to check.
What can you build with this today?
Because the output is structured (a label plus a confidence score per question), it plugs directly into ordinary conditional logic. A missing-signature result can trigger an automatically generated follow-up email asking the sender to sign and resubmit, no large language model needed to write that email, just an if/then branch keyed off the image model’s answer. The same pattern extends to invoice routing, scan quality checks, refund-claim photo verification, or any process where a human currently has to eyeball a document before it can move to the next step.
The broader shift is that the “diamond” steps in a process map, the judgment calls, no longer require a custom-trained classifier for every company and every form type. Open, general-purpose image decision models can take a first pass at a wide range of document and image review tasks, with confidence scoring and human-in-the-loop fallback handling the cases they can’t resolve on their own.
Frequently Asked Questions
What is an image decision model?
It’s a type of AI model that takes an image (a scan, screenshot, or photo) along with a set of predefined questions and returns structured answers with confidence scores, without needing a separate OCR or text-extraction step.
How is this different from traditional RPA?
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
Traditional RPA automates fixed, rule-based steps like clicking buttons or copying data. It historically struggled with judgment calls that required interpreting an image, such as checking whether a form is signed. Image decision models are built specifically to handle those judgment calls directly from visual input.
Do these models require fine-tuning for each company’s documents?
Not necessarily. Part of the appeal of general-purpose open image decision models is that they aim to handle a range of document types out of the box, unlike earlier deep learning approaches that often needed custom training per use case. Testing and question-phrasing adjustments are still recommended, since models vary in which question types they handle best.
Can one model handle every type of question equally well?
No. Testing across models shows that a given model might be strong on classification-style questions (what’s missing from this form?) while being weaker on direct true/false questions (is this signed?), even on the identical document. Running more than one model and comparing results helps catch these gaps.
What happens when the model isn’t confident in its answer?
A confidence threshold can be set so that low-confidence answers get routed to a human reviewer or a second model instead of being accepted automatically. Some models also support an explicit “unknown” response, which reduces the risk of a confidently wrong answer going unchecked.