Mistral launches Large 4 with 1.05T parameters, 49B active and 1M-token context

Mistral AI has launched Mistral Large 4, a new open-weight multimodal model built around a granular Mixture-of-Experts architecture with 1.05 trillion total parameters, 49 billion active parameters and a 1-million-token context window.
Released on October 6, 2026 as a Public Preview, Large 4 also integrates a 1.6B-parameter vision encoder and supports function calling, structured outputs, document Q&A, batching and Mistral's Agents and Conversations APIs.
The architecture places Large 4 in an unusual position: its total parameter count crosses the trillion-parameter threshold, while only about 4.7% of those parameters are active for a given token. This allows Mistral to increase model capacity without incurring the compute profile of a similarly sized dense model.
··········
MISTRAL LARGE 4 AT A GLANCE
........
Specification | Mistral Large 4 |
Release date | October 6, 2026 |
Release status | Public Preview |
Architecture | Granular Mixture-of-Experts |
Total parameters | 1.05T |
Active parameters | 49B |
Active share | ≈4.67% |
Vision encoder | 1.6B parameters |
Context window | 1M tokens |
Modalities | Text and vision |
Weights | Open-weight |
Function calling | Supported |
Structured outputs | Supported |
Agents API | Supported |
Document Q&A | Supported |
Batch inference | Supported |
........
The combination of trillion-scale total capacity and 49B active parameters is central to the model's design. Large 4 is not equivalent to running a dense 1.05T-parameter network for every generated token.
··········
ONLY ABOUT 4.7% OF THE MODEL IS ACTIVE PER TOKEN
Mixture-of-Experts models distribute parameters across multiple specialized components and route each token through only part of the network.
For Large 4, Mistral reports 1.05T total parameters and 49B active parameters.
Data Studios calculation:
49B ÷ 1,050B ≈ 4.67%
This means roughly one parameter in 21 is active for a given token when comparing the published active and total parameter counts.
The ratio does not mean Large 4 has the same memory or deployment requirements as a dense 49B model. The complete model weights still need to be stored or distributed across the inference system. The advantage is primarily that the computational path for each token uses a much smaller subset of the total network.
This separation between total capacity and active compute is what allows modern MoE systems to scale model size without scaling inference computation proportionally.
··········
THE 1M-TOKEN CONTEXT WINDOW TARGETS LONG-RUNNING WORKLOADS
Large 4 supports a 1-million-token context window, substantially expanding the amount of information that can be supplied within a single model context.
That capacity is relevant for workloads involving large repositories, collections of documents, lengthy technical records and extended agent histories.
A large context window does not guarantee that every token receives equal effective attention, nor does it eliminate the need for retrieval systems. Processing extremely large contexts also increases inference cost and latency.
It does, however, raise the ceiling for applications that need to preserve substantially more information within the model's immediate working context.
For agents, this can include tool outputs, previous decisions, documents and accumulated interaction history across longer workflows.
··········
LARGE 4 COMBINES TEXT AND VISION IN A GENERAL-PURPOSE MODEL
Mistral describes Large 4 as a general-purpose multimodal model.
Its architecture includes a 1.6B-parameter vision encoder, allowing applications to combine visual information with text rather than relying exclusively on separate image-processing pipelines.
This makes the model applicable to workflows involving documents, screenshots, charts, diagrams and other mixed visual-textual inputs.
The vision encoder represents only a small fraction of the full 1.05T-parameter architecture, but it gives the model a native path for converting visual information into representations that can participate in the broader reasoning process.
For enterprise applications, this is particularly relevant to document workflows where important information is encoded not only as text but also through layout, tables, figures and visual structure.
··········
LARGE 4 IS BUILT FOR TOOL USE AND AGENTIC APPLICATIONS
Large 4 supports function calling and Mistral's Agents and Conversations interfaces.
Function calling allows the model to determine when an external function or API is required and generate the structured arguments needed to invoke it.
That turns the model from a text-generation component into a reasoning layer capable of participating in multi-step software workflows.
........
Capability | Practical use |
Function calling | Invoke external APIs and software functions |
Structured outputs | Return machine-readable responses |
Agents API | Build persistent agentic workflows |
Conversations API | Maintain multi-turn application state |
Document Q&A | Analyze document-based information |
Vision | Process visual and textual inputs |
1M context | Maintain large working contexts |
Batch inference | Process large asynchronous workloads |
........
The combination is more important than any individual feature. Long context, multimodality and tool calling can operate within the same model, allowing an agent to inspect information, reason over it and invoke external systems without switching between separate specialist models for every stage.
··········
PUBLIC PREVIEW IS NOT THE SAME AS GENERAL AVAILABILITY
Mistral has released Large 4 as a Public Preview, rather than as a General Availability model.
Under Mistral's model lifecycle policy, Public Preview represents a near-final stage intended for public use and feedback, but models in this category can still receive silent updates.
They also do not have a guaranteed transition to General Availability.
This distinction matters for production deployments that require stable model behavior across time. Mistral's General Availability models do not receive silent updates, while Public Preview releases can change before their final production version.
Large 4 can therefore be evaluated and integrated now, but organizations requiring reproducible production behavior should account for its preview status.
··········
API PRICING MAKES THE ACTIVE-PARAMETER ARCHITECTURE COMMERCIALLY RELEVANT
Mistral is making Large 4 available through its inference platform rather than limiting the release to downloadable weights.
The model documentation lists API pricing for input, cached input and output tokens, with different displayed rates depending on the inference tier.
This means the MoE architecture has two distinct implications.
For self-hosting, organizations need to consider the complete weight footprint and hardware topology.
For managed inference, the user does not directly manage expert routing or model placement and instead pays according to token usage.
The distinction is important because 49B active parameters should not be interpreted as a 49B deployment footprint. Active parameters describe the computational route through the network, while the complete 1.05T model represents a much larger storage and infrastructure problem.
··········
LARGE 4 REPRESENTS A MAJOR SCALE INCREASE FOR MISTRAL'S OPEN-WEIGHT STRATEGY
Mistral has previously used open-weight releases to compete with substantially larger proprietary model providers while maintaining deployment options outside a closed API.
Large 4 pushes that strategy into the trillion-parameter MoE class.
The architecture combines four characteristics that increasingly define frontier open models:
very large total capacity → sparse activation → multimodality → agent-oriented tool use.
The result is a model designed not only for conversational inference but also for long-context applications, document analysis, software integration and autonomous workflows.
Its open-weight status is particularly significant for organizations that want greater control over deployment, infrastructure and data handling than a closed hosted model permits.
··········
THE 1.05T NUMBER NEEDS TO BE INTERPRETED CORRECTLY
The headline parameter count is large enough to invite direct comparisons with other trillion-parameter systems, but total parameters alone provide limited information about real inference requirements or model quality.
For Large 4, 49B active parameters are at least as important as the 1.05T total figure when interpreting the architecture.
Benchmark performance, routing efficiency, memory bandwidth, quantization, serving topology and workload characteristics all influence how the model behaves in practice.
The 1M context window likewise describes maximum context capacity rather than guaranteed performance across every long-context task.
Large 4 should therefore be evaluated as a sparse multimodal system rather than as a conventional dense trillion-parameter model.
Its significance lies in the combination: Mistral has moved its flagship open-weight architecture beyond one trillion total parameters while keeping per-token activation below 50B and integrating vision, long context and agent tooling into the same model.
··········
FOLLOW US FOR MORE.
DATA STUDIOS
datastudios.org




