Ling 3.1 Flash Hands-On: What This Free 560B Model Gets Right and Wrong
Hands-on testing of Ling 3.1 Flash, Ant Group's 560B MoE model, across coding, SQL tuning, SVG design and medical reasoning tasks.

What is Ling 3.1 Flash?
Ling 3.1 Flash is a hybrid reasoning model from Inclusion AI, the research group inside Ant Group. It uses a mixture-of-experts (MoE) design with around 560 billion total parameters, but only about 25 billion are active per token, which keeps inference costs far lower than the raw parameter count suggests. It supports context windows up to a million tokens and is positioned as a coding and agentic model, meaning it’s built to call tools, read files, and carry out multi-step tasks rather than just answer questions in a chat window. At the time of testing, it wasn’t yet open-sourced, but Ant Group had signaled a release within roughly a week.
TL;DR
- Ling 3.1 Flash is a 560B-parameter MoE model with only about 25B active parameters per token, giving it strong efficiency relative to its size.
- It supports a million-token context window, which matters for large codebases, long documents, or multi-file agentic tasks.
- On Ant Group’s own benchmark chart, Ling 3.1 Flash lands at or near the top on cyber security, healthcare, and real-world agentic test suites, though it’s not a clean sweep against Claude Opus on every panel.
- Hands-on testing showed a mixed record on complex agentic coding: a Blender-to-Godot game pipeline produced broken, disoriented gameplay despite the model completing the task end to end.
- It performed notably well on single-shot SVG generation, producing a usable animated night skyline from one prompt.
- On SQL query optimization, it correctly isolated the single highest-impact fix in a tangled query and explained why other common fixes wouldn’t help, which is a harder skill than just listing changes.
- It handled a simulated emergency medical case with a structured, clinically plausible answer, and did reasonably well on a multilingual test across 80 languages, though it struggled with rare and low-resource languages.
- The core value proposition is the same as other Ling models: it’s free and cheap to self-host, not best-in-class, but a capable daily driver for general-purpose and agentic work.
Other agents start typing. Remy starts asking.
Scoping, trade-offs, edge cases — the real work. Before a line of code.
How does Ling 3.1 Flash perform on benchmarks?
Ant Group’s published benchmark chart spans 12 tests covering real-world categories: banking and finance agents, cyber security, codebase questions, and healthcare. Ling 3.1 Flash sits at or near the top on Frontier Suite, Cyber Gym, and HealthBench Professional, and holds its own on agentic-style evaluations like Skill Bench and Multi-Challenge.
It isn’t a universal leader. Claude Opus edges it out on a few panels, including Terminal Bench. It’s also worth noting that the HealthBench result reportedly came from Ant Group’s own evaluation environment rather than an independent one, which is a detail worth keeping in mind when comparing it to scores from neutral third-party benchmarks. The general pattern: Ling does best on tasks that resemble practical, applied work rather than pure reasoning puzzles.
How did it do on a real agentic coding task?
The toughest test involved giving Ling 3.1 Flash an open-ended goal: generate its own 3D assets using Blender, then wire them into a working Godot game, with no step-by-step guidance. This is a genuinely hard agentic workflow, since it requires the model to plan across tools, generate assets, and integrate them into a functioning game loop.
After roughly an hour of autonomous tool use, the model did produce files and a playable game panel, but the result broke down in practice. Movement keys worked, but objects were disoriented, a trampoline asset rendered sideways, and there was no coherent landing or level layout. The takeaway from this test is that Ling 3.1 Flash can execute the mechanics of a long agentic chain (calling tools, writing files, running commands) but the actual spatial and logical coherence of a complex multi-step creative task fell apart. Complex, open-ended agentic goals remain a weak point.
Is Ling 3.1 Flash good at design tasks like SVG generation?
Yes, this was one of the stronger results. Ant Group claims Ling ranks near the top for SVG and mobile app design tasks, and a direct one-shot test backed that up. The prompt asked for a single self-contained HTML file containing an animated SVG scene: a night city skyline viewed from a rooftop. The model generated a script using just over 20,000 completion tokens, and the rendered result was a legitimate night skyline with a blinking tower and layered buildings. It wasn’t as dense or richly detailed as a top-tier generative design tool might produce, but for a single prompt with no iteration, it was a clean, usable output. This suggests Ling 3.1 Flash is reasonably reliable for scoped, well-defined generative front-end tasks, even if it struggles with open-ended multi-tool pipelines.
Can it actually optimize SQL queries?
This test targeted something more specific than raw code generation: given a slow, tangled SQL query, could the model isolate the single highest-impact fix instead of dumping a generic list of ten possible optimizations?
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
The result was strong. Ling correctly identified a correlated subquery that recomputed each opponent’s win rate per row as the real bottleneck, and replaced it with a single pre-aggregated join, including a filter moved into the join logic. It didn’t stop at a theoretical explanation either: it built a synthetic dataset, timed the query before and after, and ran an ablation test showing the subquery accounted for nearly all the runtime cost. It also explained why alternative fixes, like adding indexes or rewriting the outer query, wouldn’t have solved the actual problem. For a former database administrator reviewing the output, this was judged a strong, focused result: one real fix, properly justified, with evidence.
How well does it handle medical reasoning and multilingual tasks?
In a simulated emergency case test (explicitly not a substitute for real clinical judgment), the model was asked to act as a clinical assistant and produce a prioritized consult note. It correctly flagged an inferior heart attack pattern, caught a clinical trap in the case where low blood pressure pointed to right-sided involvement, and recommended checking a V4R lead, a specific and clinically relevant detail. The response stayed structured across five clean sections without rambling, which reflects strong instruction-following even on a domain-specific, high-stakes topic.
The multilingual test asked the model to translate “will you marry me” into 80 languages in a single request, ranging from major world languages to rare ones like Faroese, Baluchi, and Tyrian. Major languages (Spanish, French, German, Russian, Mandarin, Turkish, and others) came back looking correct, and Arabic was correctly rendered in the female grammatical form. Rare and low-resource languages were noticeably weaker, with several outputs (including Tamil, Yoruba, and Basque) described as off or just slightly wrong, suggesting Ling’s language coverage, like most models, thins out quickly outside major language families.
Is Ling 3.1 Flash worth using?
For general-purpose daily use and well-scoped technical tasks like SQL tuning, SVG generation, and structured domain reasoning, Ling 3.1 Flash performs competently and sometimes impressively. It is not positioned as a frontier-beating model, and it clearly loses ground on complex, multi-stage agentic tasks that require sustained coherence across tool calls. Its real selling point is accessibility: it’s free to use, inexpensive to self-host on rented GPU infrastructure, and the active-parameter count (25B out of 560B total) keeps inference costs manageable compared to dense models of similar scale. For teams or individuals looking for a capable, low-cost daily driver rather than a benchmark leader, it’s a reasonable option once it’s publicly released.
Frequently Asked Questions
What is the parameter count of Ling 3.1 Flash?
It has roughly 560 billion total parameters in a mixture-of-experts architecture, but only about 25 billion are active per token, which keeps compute costs closer to a much smaller dense model.
Is Ling 3.1 Flash open source?
At the time of testing it had not yet been released publicly, but Ant Group’s Inclusion AI team indicated an open-source release was expected within about a week.
How does Ling 3.1 Flash compare to Claude Opus?
It performs at or near the top on several benchmark categories like cyber security and healthcare-focused evaluations, but Claude Opus still outperforms it on some panels, including terminal-based agentic tasks.
What are Ling 3.1 Flash’s biggest weaknesses?
Complex, open-ended agentic workflows that require chaining multiple tools over long sessions, like generating 3D assets and assembling them into a working game, is where it breaks down. Rare and low-resource languages are also a weak spot.
What is Ling 3.1 Flash good at?
Remy doesn't write the code. It manages the agents who do.
Remy runs the project. The specialists do the work. You work with the PM, not the implementers.
It handled single-shot SVG and front-end design generation well, isolated the correct high-impact fix in a complex SQL optimization task, and produced a structured, clinically coherent response to a simulated emergency medical case.