Goodbye, Llama? Meta launches new proprietary AI model Muse Spark — first since Superintelligence Labs' formation

Meta has unveiled Muse Spark, its first proprietary AI model developed under the new Superintelligence Labs division led by Alexandr Wang, signaling a potential shift away from its open-source Llama roots.
Meta has been one of the most interesting companies of the generative AI era — initially gaining a loyal and huge following of users for the release of its mostly open source Llama family of large language models (LLMs) beginning in early 2023 but coming to screeching halt last year after Llama 4 debuted to mixed reviews and ultimately, admissions of gaming benchmarks.
That bumpy rollout of Llama 4 apparently spurred Meta founder and CEO Mark Zuckerberg to totally overhaul Meta's AI operations in the summer of 2025, forming a new internal division, Meta Superintelligence Labs (MSL) which he recruited 29-year-old former Scale AI co-founder and CEO Alexandr Wang to lead as Chief AI Officer.
Now, today, Meta is showing us the fruits of that effort: Muse Spark, a new proprietary model that Wang says (posting on rival social network X, used more often by the machine learning community) is "the most powerful model that meta has released," and has "support for tool-use, visual chain of thought, & multi-agent orchestration." He also says it will be the start of a new Muse family of models, raising questions about what will become of Meta's popular lineup and ongoing development of the Llama family.
It arrives not as a generic chatbot, but as the foundation for what Wang calls "personal superintelligence"—an AI that doesn’t just process text but "sees and understands the world around you" to act as a digital extension of the self, echoing Zuckberg's public manifesto for a vision of personal superintelligence published in summer 2025.
However, it is proprietary only — confined for now to the Meta AI app and website, as well as a " private API preview to select users," according to Meta's blog post announcing it — a move likely to rankle the literally billions of users of Llama models and the thousands of developers who relied upon it (some of whom are active participants in rival social network Reddit's r/LocalLLaMA subreddit). In addition, no pricing information for the model has yet been announced.
It's unclear if Meta has ended development on the Llama family entirely. When asked directly by VentureBeat, a Meta spokesperson said in an email: “Our current Llama models will continue to be available as open source,” which doesn’t address the question of development of future Llama models.
Visual chain-of-thought
At its core, Muse Spark is a natively multimodal reasoning model. Unlike previous iterations that "stitched" vision and text together, Muse Spark was rebuilt from the ground up to integrate visual information across its internal logic. This architectural shift enables "visual chain of thought," allowing the model to annotate dynamic environments—identifying the components of a complex espresso machine or correcting a user's yoga form via side-by-side video analysis.
The most significant technical leap, however, is a new "Contemplating" mode. This feature orchestrates multiple sub-agents to reason in parallel, allowing Meta to compete with extreme reasoning models like Google's Gemini Deep Think and OpenAI's GPT-5.4 Pro.
In benchmarks, this mode achieved 58% in "Humanity’s Last Exam" and 38% in "FrontierScience Research," figures that Meta claims validate their new scaling trajectory.
Perhaps more impressive for the company’s bottom line is the model’s efficiency. Meta reports that Muse Spark achieves its reasoning capabilities using over an order of magnitude less compute than Llama 4 Maverick, its previous mid-size flagship. This efficiency is driven by a process called "thought compression". During reinforcement learning, the model is penalized for excessive "thinking time," forcing it to solve complex problems with fewer reasoning tokens without sacrificing accuracy.
Benchmarks reveal a return-to-form
The launch of Muse Spark is framed as a statistical "quantum leap," ending Meta’s year-long absence from the absolute frontier of AI performance.
By reconciling Meta’s official internal data with independent auditing from third-party LLM tracking firm Artificial Analysis, a clear picture emerges: Muse Spark is not just a marginal improvement over the Llama series; it is a fundamental re-entry into the "Top 5" global models.
According to the Artificial Analysis Intelligence Index v4.0, Muse Spark achieved a score of 52. For context, Meta’s previous flagship, Llama 4 Maverick, debuted in 2025 with an **Index score of just 18. **
By nearly tripling its performance, Muse Spark now sits within striking distance of the industry’s most elite systems, trailing only Gemini 3.1 Pro Preview (57), GPT-5.4 (57), and Claude Opus 4.6 (53).
Meta’s official benchmarks suggest that Muse Spark is particularly dominant in multimodal reasoning, specifically where visual figures and logic intersect.
CharXiv Reasoning: In "figure understanding," Muse Spark achieved a score of86.4, significantly outperformingClaude Opus 4.6(65.3),Gemini 3.1 Pro(80.2), andGPT-5.4(82.8).MMMU Pro: Official reports place the model at80.4, while Artificial Analysis’s independent audit measured it at80.5%. This makes it thesecond-most capable vision modelon the market, surpassed only by Gemini 3.1 Pro Preview (83.9% official; 82.4% independent).Visual Factuality (SimpleVQA): Muse Spark scored71.3, placing it ahead of GPT-5.4 (61.1) and Grok 4.2 (57.4), though it narrowly trails Gemini 3.1 Pro (72.4).
These scores validate Meta’s focus on "visual chain of thought," enabling the model to not just recognize objects, but to reason through complex spatial problems and dynamic annotations.
The "Thinking" gear of Muse Spark was put to the test against specialized benchmarks designed to break non-reasoning models.
Humanity’s Last Exam (HLE): In this multidisciplinary evaluation, Meta reports a score of42.8(No Tools) and50.4(With Tools). Independent audits by Artificial Analysis tracked the model at39.9%, trailing Gemini 3.1 Pro Preview (44.7%) and GPT-5.4 (41.6%).GPQA Diamond (PhD Level Reasoning): Muse Spark achieved a formidable89.5, surpassing Grok 4.2 (88.5) but trailing the specialized "max reasoning" outputs of Opus 4.6 (92.7) and Gemini 3.1 Pro (94.3).ARC AGI 2: This remains a notable weak point. Muse Spark scored42.5, far behind the abstract reasoning puzzles solved byGemini 3.1 Pro(76.5) andGPT-5.4(76.1).CritPT (Physics Research): Independent auditing found Muse Spark achieved the5th highest scoreat11%. This marks a substantial lead overGemini 3 Flash(9%) andClaude 4.6 Sonnet(3%).
One of the most striking results from the official data is Muse Spark's performance in the health sector, likely a result of Meta's collaboration with over 1,000 physicians.
HealthBench Hard: Muse Spark achieved42.8, a massive lead overClaude Opus 4.6(14.8),Gemini 3.1 Pro(20.6), and evenGPT-5.4(40.1).MedXpertQA (Multimodal): It scored78.4, comfortably ahead of Opus 4.6 (64.8) and Grok 4.2 (65.8), though it still trails Gemini 3.1 Pro’s top-tier score of 81.3.
Agentic Systems and Efficiency: The "Thought Compression" Effect
While Muse Spark excels at reasoning, its "agentic" performance—executing real-world work tasks—presents a more nuanced picture.
SWE-Bench Verified: Muse Spark scored77.4, trailingClaude Opus 4.6(80.8) andGemini 3.1 Pro(80.6).GDPval-AA Elo: Meta’s official score of1444differs slightly from Artificial Analysis’s recorded1427. In both cases, Muse Spark trailsGPT-5.4(1672) andOpus 4.6(1606), suggesting that while the model "thinks" well, it is still refining its ability to "act" in long-horizon software and office workflows.Token Efficiency: This is where Muse Spark distinguishes itself. To run the Intelligence Index, it used58 million output tokens. In contrast,Claude Opus 4.6required157 milliontokens andGPT-5.4required120 million. This supports Meta's claim of "thought compression" making the model more efficient.
Source: VentureBeat
















