Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
open source vs closed AIenterprise AI trendsAI industry pivot

Open Weights vs Enterprise Lock-In: Cohere and Reflection AI's Split Bets

Cohere's North 2 locks AI inside enterprise walls while Reflection AI's Beam promises open weights. Here's what the split strategy means.

Edited by Luis Chavez-Mattos, Director of Product RSS
Open Weights vs Enterprise Lock-In: Cohere and Reflection AI's Split Bets

Why are AI companies splitting into “open” and “locked down” camps?

Because the money has to come from somewhere. OpenAI and Anthropic dominate enterprise AI spending, and both reportedly operate at heavy losses despite massive revenue. Investors who funded years of model training now expect returns, and that pressure is reshaping how every other company in the space behaves. Some labs respond by closing off their best models or delaying open releases. Others, especially outside the US, pivot hard into enterprise sales, pitching governance and cost control rather than raw model intelligence. Two releases on the same day, Cohere’s North 2 and Reflection AI’s Beam, show both strategies at once.

TL;DR

  • Cohere’s North 2 is not a new model but a full agentic platform that enterprises can run in the cloud, on premises, or fully air-gapped, aimed at government and regulated industries that need data to never leave their walls.
  • North 2 sells governance, not intelligence: reusable agents, memory, skills, connectors to tools like Teams, Slack, SharePoint and GitHub, admin-level spending limits, and the ability to plug in models from other providers.
  • Reflection AI’s Beam is a 501 billion parameter mixture-of-experts model with only 23 billion active parameters per token, pretrained on 23.8 trillion tokens using a cluster of 6,144 Nvidia GPUs in under four weeks.
  • Weights aren’t out yet: Reflection says Beam will ship under the permissive Apache 2.0 license later this month, meaning every claim so far comes from the company’s own announcement rather than independent testing.
  • Beam leads some open-model benchmarks like SWE-Bench and Terminal-Bench but trails larger open models such as Qwen3-Max and Kimi K2 on raw capability, a gap Reflection itself acknowledges.
  • Both companies are reacting to the same fear: getting locked into a handful of closed providers while nobody has proven AI’s unit economics actually work.
✗ VIBE-CODED APP
Tangled. Half-built. Brittle.
✓ AN APP, MANAGED BY REMY
UIReact + Tailwind✓
APIValidated routes✓
DBPostgres + auth✓
DEPLOYProduction-ready✓
Architected. End to end.

Built like a system. Not vibe-coded.

Remy manages the project — every layer architected, not stitched together at the last second.

What exactly is Cohere’s North 2?

North 2 is a platform, not a standalone model. Its pitch is that enterprises no longer have to trade off between security, intelligence, cost, and control. Instead, they get all four packaged into one deployable system. Organizations can run it fully air-gapped, meaning completely cut off from the internet, which matters enormously for government agencies and regulated industries that can’t risk sensitive data touching an external API.

The platform includes reusable agents, persistent memory, configurable skills, and prebuilt connectors into tools already in enterprise stacks: Microsoft Teams, Slack, SharePoint, GitHub. Admins get a dashboard to set token spending limits per user, turning AI costs into something a finance team can actually forecast. North 2 is also designed to be model-agnostic, so a company can bring its own model rather than being locked into Cohere’s.

That design choice is the real signal. Cohere isn’t trying to win on who has the smartest model. It’s betting that enterprises, especially in regulated sectors, care more about predictable bills, auditability, and data sovereignty than about shaving a few points off a benchmark.

What makes Reflection AI’s Beam different?

Beam is built on a mixture-of-experts (MoE) architecture. Think of it as a hospital with hundreds of specialist doctors: for any given patient, only a handful of specialists get called in. You get the combined knowledge of the whole hospital but only pay for the doctors actually used. That’s why Beam can carry 501 billion total parameters while only activating 23 billion of them for any given token, keeping inference costs down relative to a dense model of similar size.

Reflection says Beam was pretrained on 23.8 trillion tokens across a cluster of 6,144 Nvidia GPUs, completing the run in under four weeks. The company also claims it discarded roughly 95% of raw internet data during cleaning, arguing that standard filtering pipelines would have missed about 1.8 trillion tokens worth keeping. The underlying argument: data quality beats data volume, a claim that lines up with what most serious pretraining teams have been saying for the past two years.

On the infrastructure side, Reflection reports a notably smooth pretraining loss curve, with only nine restarts across the entire run and 92.3% of total time spent on productive training rather than recovering from failures. For a run at this scale, that’s a meaningful engineering claim, since large training runs are notorious for crashes and instability.

How was Beam’s reasoning capability trained?

The more interesting part may be what happened after pretraining. Reflection ran a large reinforcement learning phase, using roughly 10,500 GPUs over four weeks to generate more than 100 million attempts across about one million distinct practice environments. The company describes this as one of the largest open RL runs by any lab, and says benchmark scores were still climbing when the run was stopped, suggesting there was more performance left on the table.

Everyone else built a construction worker.
We built the contractor.

🦺
CODING AGENT
Types the code you tell it to.
One file at a time.
🧠
CONTRACTOR · REMY
Runs the entire build.
UI, API, database, deploy.

This matters because reinforcement learning, not just pretraining scale, has become the dominant lever for improving reasoning, tool use, and agentic behavior in recent frontier models. A company willing to throw this much compute at RL specifically is signaling that it sees that phase, not raw parameter count, as the real differentiator.

How does Beam compare to other open models?

According to Reflection’s own published benchmarks, Beam leads several tests against comparably sized open models like Inkling, Nvidia’s Nemotron, and others, scoring 80.9 on SWE-Bench and around 80 on Terminal-Bench. But the comparison gets less favorable against larger open-weight competitors. On Terminal-Bench, GLM scored around 81, Qwen3-Max reached 86.6, and Kimi K2 posted 88.3, a clear lead. Reflection acknowledges this directly, framing Beam not as the most capable open model overall but as the most efficient Western open-weight option relative to its active parameter count.

That distinction matters. Beam isn’t claiming to beat Kimi K2 or Qwen3-Max on raw capability. It’s claiming a better capability-per-compute tradeoff, which is a different and more modest pitch, but one that could matter a lot for anyone trying to self-host a frontier-adjacent model without a hyperscaler’s budget.

Is the open weights promise worth trusting yet?

Not yet, and that’s worth being blunt about. As of Reflection’s announcement, no weights were actually downloadable. The company says they’ll arrive later this month under an Apache 2.0 license, which is about as permissive as open licensing gets. But every performance number currently circulating comes from Reflection’s own internal testing, not independent verification. Until the files are public and third parties can run their own evaluations, the right posture is cautious interest rather than confidence.

There’s also no stated reason for the delay between announcement and release. It could be final safety testing, infrastructure prep for hosting, or simply sequencing an API launch before the open release. Any of those would be unremarkable, but the gap does mean the open weights claim is currently a promise, not a shipped product.

What does this split say about AI’s business model?

Cohere and Reflection are solving for the same underlying anxiety from opposite directions. Enterprises and governments don’t want to be permanently dependent on two or three closed providers, especially while nobody has proven that frontier AI companies can run profitably at scale. Cohere’s answer is to sell control: keep the data in-house, make costs predictable, let customers swap models if needed. Reflection’s answer is to sell independence: give away the weights so anyone can run, modify, or self-host the model without a vendor relationship at all.

Neither approach guarantees commercial success. Enterprise-focused pivots depend on actually landing large contracts with security-sensitive buyers, a crowded field now that Mistral and others are chasing the same budgets. Open-weight strategies depend on community adoption and the ability to monetize around the free model, through hosting, fine-tuning services, or enterprise support. Both are bets that the current two-horse race between OpenAI and Anthropic won’t last forever, and that there’s room, and money, in building something those two aren’t offering.

Frequently Asked Questions

What is Cohere North 2?

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

North 2 is Cohere’s agentic AI platform designed for enterprise and government deployment. It can run in the cloud, on premises, or fully air-gapped, and includes reusable agents, memory, tool connectors, and admin controls for spending limits. It’s model-agnostic, meaning customers can plug in models from other providers.

What is Reflection AI’s Beam model?

Beam is a 501 billion parameter mixture-of-experts language model from Reflection AI, with only 23 billion parameters active per token. It was pretrained on 23.8 trillion tokens and underwent a large reinforcement learning phase. Reflection plans to release the weights under an Apache 2.0 license.

Is Beam already open source?

Not yet. As of the announcement, the weights had not been released. Reflection says they will publish them later in the month under Apache 2.0, but until that happens, benchmark claims come only from the company itself.

How does Beam compare to models like Kimi K2 or Qwen3-Max?

On several benchmarks, larger open models like Kimi K2 and Qwen3-Max outperform Beam, particularly on tasks like Terminal-Bench. Reflection positions Beam as the most efficient Western open-weight model for its active parameter count rather than the top performer overall.

Why are companies like Cohere focusing on enterprise instead of consumer AI?

Because enterprise and government budgets reward governance, security, and predictable costs over raw model capability. With OpenAI and Anthropic dominating headline model performance while reportedly losing money, companies like Cohere and Mistral are differentiating through control and deployment flexibility instead of competing directly on benchmarks.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.