top of page

Cloudflare launches Clef-omni with 30B parameters, native audio-video understanding and low-cost AI decisions

1 minute ago
15 min read
Cloudflare Clef-omni launches with 30B parameters, native audio-video understanding and low-cost AI decision scoring - Data Studios

Cloudflare has launched Clef-omni, a 30-billion-parameter multimodal decision model capable of evaluating text, images, audio, and video within a single inference request, extending its open-weight Clef family beyond conventional text classification and visual analysis. Released on October 9, 2026, the model is available through Cloudflare Workers AI at $0.15 per million input tokens, with no output-token charges, and its weights have been published under the Apache 2.0 license.


Built on Alibaba's Qwen3-Omni-30B-A3B-Instruct architecture, Clef-omni uses a Mixture-of-Experts design with approximately 3 billion active parameters. Unlike general-purpose language models that generate conversational responses, it evaluates predefined questions and returns probability scores for the permitted answers. Its principal advantage is the ability to perform these decisions directly on different media types, without requiring developers to construct separate transcription, image-captioning, and text-classification pipelines.


Cloudflare describes applications ranging from customer-service routing and invoice processing to security monitoring, industrial equipment inspection, and autonomous-agent supervision. A single request can combine a photograph, an audio recording, and a video clip, allowing the model to consider visual and acoustic evidence together before producing structured decisions.


The launch also introduces changes across the wider Clef family. Clef-flash now costs $0.038 per million input tokens, down from $0.09, while the original Clef benefits from serving optimizations that Cloudflare reports have improved inference speed by up to twice its previous level.


These improvements create three distinct options for developers: Clef-flash for inexpensive high-volume decisions, Clef for more demanding text and visual classification, and Clef-omni for applications requiring audio-video understanding.


The economic opportunity is substantial for software systems that execute thousands or millions of recurring decisions, but the models' capabilities remain deliberately narrower than those of general-purpose AI assistants. Their value depends on how accurately they classify information, how reliably their probabilities are calibrated, and whether the surrounding application translates those predictions into appropriate actions.


··········


CLEF-OMNI COMBINES A 30B MIXTURE-OF-EXPERTS BACKBONE WITH STRUCTURED DECISION SCORING


The model retains the multimodal comprehension capabilities of Qwen3-Omni while replacing conventional text generation with a specialized mechanism for evaluating predefined answers.


Cloudflare developed Clef-omni from Qwen3-Omni-30B-A3B-Instruct, using the foundation model's language, vision, and audio processing capabilities. The original architecture contains approximately 30 billion parameters distributed across a Mixture-of-Experts network, with around 3 billion activated during inference.


A sparse MoE model selects relevant computational components for each input instead of activating the entire network. This can reduce the arithmetic required for a forward pass, although the full model still needs sufficient memory and infrastructure to store and access its parameters.


Cloudflare retained the comprehension backbone and multimodal encoders while excluding the speech-generation components from its operational inference path. It then introduced a joint schema head, designed to evaluate multiple questions and their possible answers using representations extracted from the input.


The mechanism differs substantially from asking a language model to generate a JSON document.


In a conventional classification workflow, a general-purpose model might generate several tokens to describe the selected category. The surrounding application would then parse that response, check whether it matches the expected schema, and potentially retry if the model returned an invalid answer.


Clef-omni instead computes scores for the permitted options directly. Its decision head processes the input's internal representations, identifies evidence relevant to each question, and produces one score per answer option. A softmax transformation converts those scores into probabilities, allowing the application to select an outcome or apply its own confidence thresholds.


Cloudflare describes a two-stage attention mechanism in which candidate answers first gather relevant evidence from the available context, after which question representations interact with the broader information to calculate their scores.


This architecture also enables several questions to be evaluated in the same inference pass. A customer-service application could simultaneously classify the subject of a request, estimate its severity, and determine whether escalation is appropriate, instead of making three independent model calls.


The model was post-trained using low-rank adaptation techniques, commonly called LoRA, while keeping the foundation backbone frozen. Cloudflare combined label-smoothed cross-entropy training with Brier-score calibration, an approach intended to improve both classification quality and the relationship between predicted confidence and actual correctness.


Probability calibration is relevant because applications often need more than the most likely answer. A model assigning 51% confidence to a security alert should not necessarily trigger the same automated response as one assigning 99%, even when both select the same category.


However, training a model for calibrated predictions does not guarantee that its scores remain reliable across every industry, dataset, or production environment. Organizations still need to measure calibration using representative examples from their own workloads.


........


Technical characteristic

Cloudflare Clef-omni

Release date

October 9, 2026

Developer

Cloudflare

Foundation model

Qwen3-Omni-30B-A3B-Instruct

Architecture

Mixture-of-Experts

Total parameters

Approximately 30 billion

Active parameters

Approximately 3 billion

Inference design

Single-pass structured decision scoring

Supported inputs

Text, JSON, images, audio, video

Supported decisions

Binary, multiple-choice, ordinal scoring

Hosted context window

64,000 tokens

Output

Probabilities and structured decisions

Hosted platform

Cloudflare Workers AI

Open weights

Available on Hugging Face

License

Apache 2.0

Input pricing

$0.15 per million tokens

Output pricing

No additional charge


........


Cloudflare's model documentation also describes local execution using PyTorch and Transformers. In its tested configuration, loading the backbone in bfloat16 required approximately 64 GB of GPU memory, with validation performed on an NVIDIA H200.


This specification illustrates an important distinction between active parameters and deployment requirements. Although approximately 3 billion parameters participate in a given inference operation, the complete model is considerably larger than an ordinary dense 3B model and cannot be assumed to run economically on low-memory consumer hardware.


··········


NATIVE AUDIO AND VIDEO INPUT REMOVES SEPARATE TRANSCRIPTION AND CAPTIONING STEPS


Clef-omni accepts synchronized visual and acoustic information directly, allowing applications to make bounded decisions from recordings without first converting every input into natural-language descriptions.


Multimodal processing normally involves several stages when different models are responsible for understanding different media formats.


A customer-service system might transcribe an audio recording, send the resulting transcript to a language model, and classify the conversation. A video-monitoring application might extract frames, describe their contents using a vision model, and submit those descriptions for further analysis.


These architectures can work effectively, but every intermediate conversion introduces latency, additional infrastructure requirements, and potential information loss.


A transcription system may omit background sounds that are relevant to equipment monitoring, while a video-captioning model may fail to preserve the temporal relationship between what appears on screen and what is heard.


Clef-omni processes these inputs through the multimodal foundation model before applying its decision head. Cloudflare's implementation aligns a video's soundtrack with sampled visual frames so that the model can evaluate the relationship between visual and acoustic events.


Consider an equipment-maintenance company receiving a photograph of an installation, a recording of the machine operating, and a short video showing its moving components. A single request could ask whether the identification label is visible, whether the equipment produces abnormal sounds, and whether a particular component appears to operate correctly.


The output would consist of structured scores rather than a detailed maintenance report. The application could use those scores to identify submissions requiring further inspection or human review.


This distinction defines the product's scope: Clef-omni is intended to decide whether specified conditions are met, rather than produce unrestricted explanations of the entire recording.


The hosted model has explicit media limits. According to the current API documentation, each request accepts up to four images, four audio clips, and two videos, subject to individual file constraints and combined limits.


Audio clips may last up to 300 seconds each, while video clips are limited to 60 seconds each and sampled at approximately two frames per second. Audio and video together are subject to a combined decoded-data limit of 16 MiB.


These limits affect application design. A business reviewing hour-long conversations or continuous surveillance footage must divide the material into smaller segments, select representative intervals, or use additional processing before submitting it to Clef-omni.


Frame sampling also imposes a limitation: an event occurring entirely between sampled frames may not be represented sufficiently in the visual input, even when the model performs well on longer or more persistent activities.


The context window introduces another constraint. Media inputs are converted into model tokens and count toward the 64,000-token capacity alongside text and decision questions. Cloudflare documents that requests exceeding the context through media tokens fail, while lengthy text-state content may be truncated to fit the remaining capacity.


Consequently, a model with native video understanding is not automatically suitable for unrestricted long-form video analysis. Its current hosted configuration is better matched to short recordings, structured inspection tasks, customer-interaction classification, and selected events extracted from longer workflows.


··········


BENCHMARKS SHOW STRONG CLASSIFICATION ACCURACY, BUT CLEF-OMNI DOES NOT ALWAYS BEAT CLEF OR CLEF-FLASH


Cloudflare's published evaluations indicate competitive performance across routing and classification tasks, with meaningful differences depending on the benchmark and the model selected.


The company evaluated Clef-omni against the original Clef, Clef-flash, and TypeSafe's Jev decision model. The tests cover several categories, including tool selection, customer-support classification, product relevance, phishing identification, and specialized workflow decisions.


On BANKING77, which evaluates classification across banking-related customer intents, Clef-omni achieved a macro-F1 score of 94.8%, narrowly exceeding Clef's 94.2%.


It also obtained 97.7% on CLINC150+OOS, compared with 97.43% for Clef. This benchmark includes intent-classification tasks and out-of-scope requests, making it relevant to applications that must recognize when an input does not belong to any expected category.


However, other results favor the existing models. Clef-flash achieved 98.76 on the reported BFCL case-exact evaluation, compared with 98.2 for Clef-omni, while the original Clef performed better on PhishNChips, a phishing-related classification benchmark.


The differences are especially pronounced on the home-appliances evaluation, where Clef-flash substantially outperformed Clef-omni. These results demonstrate that the addition of multimodal capabilities does not produce a universal improvement across every classification workload.


........


Benchmark

Clef-omni

Clef

Clef-flash

BFCL, case exact

98.20

98.47

98.76

BANKING77, macro-F1

94.80

94.20

90.93

CLINC150+OOS, macro-F1

97.70

97.43

66.77

Amazon ESCI, macro-F1

57.80

57.48

57.39

PhishNChips, accuracy

73.20

79.60

75.05

Home appliances, case exact

69.30

82.95

97.73

API-Bank, accuracy

92.70

91.93

93.11


........


These scores are taken from Cloudflare's own evaluation results. They are not independently reproduced measurements, and their interpretation depends on the benchmark-specific evaluation methodology.


The metrics are also not interchangeable. Macro-F1 gives equal importance to individual classes when averaging their F1 scores, whereas classification accuracy measures the proportion of correct predictions under the benchmark's evaluation rules.


For applications with heavily imbalanced categories, macro-F1 can be particularly informative because a model cannot obtain a high score merely by predicting the most common class.


The benchmark results also raise an important question about the commercial value of multimodality.


Most of the reported comparisons involve established classification and decision benchmarks rather than comprehensive tests of audio-video understanding under realistic operating conditions. Performance on banking-intent classification or tool routing does not establish how accurately Clef-omni recognizes simultaneous visual and acoustic events in industrial footage.


A manufacturer considering the model for equipment inspection would need validation datasets representing genuine mechanical faults, background noise, different recording devices, environmental conditions, and situations in which visual and acoustic signals disagree.


Similarly, a company using Clef-omni to classify customer conversations should measure its accuracy across languages, accents, recording quality, overlapping speech, and interactions involving sensitive personal information.


The published results support Clef-omni's suitability for a range of structured decision tasks, but the model's incremental advantage over cheaper alternatives must be established separately for each production workload.


If an application processes only text, Clef-flash may deliver adequate accuracy at a substantially lower price. If it relies on synchronized audio and video, Clef-omni offers capabilities the other hosted Clef variants do not provide in the same form.


··········


CLOUDFLARE REPORTS LOW INFERENCE LATENCY, ALTHOUGH ITS PUBLISHED SPEED FIGURES DIFFER


The model's specialized decision head avoids autoregressive text generation, but end-to-end request latency is higher than the smallest timing figures cited in Cloudflare's announcement materials.


Cloudflare's October 9 changelog reports approximately 20 milliseconds for text decisions, less than 100 milliseconds for image or audio inputs, and around 300 milliseconds for a 21-second video clip with sound.


The company's longer technical announcement provides different measurements: approximately 130 milliseconds at the median for text-only decisions, 150 milliseconds for images, several hundred milliseconds for audio clips, and around 1.5 seconds for a 21-second video with sound.


Both sets of figures originate from Cloudflare. The publicly presented materials do not provide a complete reconciliation establishing whether the differences arise from different measurement boundaries, infrastructure configurations, preprocessing stages, or experimental conditions.


For that reason, the lower numbers should not be treated as guaranteed end-to-end response times for ordinary API customers.


The distinction between pure model execution and service-level latency is important. A production request can include authentication, network transmission, media decoding, preprocessing, inference, serialization, and the return of structured results. Depending on the application, some of these operations may contribute materially to the total time experienced by the caller.


The technical advantage of decision models is that they do not need to generate a lengthy response token by token. Once the input has been processed, the model can evaluate the specified options directly, avoiding the sequential output-generation stage typical of a conversational language model.


However, complex multimodal input still requires significant processing, particularly when recordings contain many frames or extensive acoustic information.


For enterprise deployment, a meaningful performance test should therefore measure the latency of the complete request using the intended media sizes, geographical locations, and concurrency levels. Median response time should be considered alongside higher-percentile measurements, because an application may experience substantial delays even when its typical requests complete quickly.


Cloudflare has separately improved the performance of the original Clef model through serving optimizations, partly involving a move to SGLang.


For inputs of approximately 3,400 tokens, the company reports that Clef's median latency declined from 616 milliseconds to 305 milliseconds, representing a twofold improvement. The 95th-percentile latency decreased from 777 to 531 milliseconds during the same comparison.


These improvements concern the existing Clef model, not a newly trained version of Clef-omni. Cloudflare states that the original model's weights were unchanged, indicating that the gain came primarily from serving infrastructure rather than increased model intelligence.


The example illustrates why inference economics cannot be evaluated through model architecture alone. Scheduling, batching, runtime implementation, hardware utilization, and memory management all influence the cost and latency experienced by customers.


··········


PRICING STARTS AT $0.15 PER MILLION INPUT TOKENS, WITH MEDIA CONVERTED INTO BILLABLE TOKENS


Clef-omni charges only for input processing, while the cheaper Clef-flash provides an alternative for applications that do not require native audio-video classification.


Cloudflare's October 9 changes establish a clear pricing hierarchy within the Clef family. Clef-flash costs $0.038 per million input tokens, Clef-omni costs $0.15, and the original Clef costs $0.24.


All three models return structured decision outputs without separate output-token charges.


The price difference is significant at scale. Clef-omni is approximately 3.95 times more expensive per input token than Clef-flash, but 37.5% cheaper than Clef.


The models also differ in hosted context capacity. Clef-flash now supports 24,000 tokens, reduced from 64,000 as part of the price optimization, while Clef and Clef-omni retain 64,000-token windows.


Cloudflare reports that just 0.24% of hosted Clef-flash requests previously exceeded 24,000 input tokens, providing an operational rationale for the reduction. Its self-hosted weights remain unchanged and can support substantially larger contexts, with Cloudflare documenting up to 256,000 tokens for that deployment option.


........


Model

Input price / 1M tokens

Hosted context

Clef-flash

$0.038

24K

Clef-omni

$0.150

64K

Clef

$0.240

64K


........


Audio and video use the same token-based billing system, rather than a separate charge per minute of recording.


Cloudflare's documentation estimates approximately 780 tokens per minute of audio. Video consumes up to approximately 15,400 visual tokens per minute at maximum resolution, plus around 780 audio tokens per minute when a soundtrack is present.


At 480p, the documentation indicates approximately 8,600 video-frame tokens per minute. A one-minute recording at that resolution with audio would therefore account for approximately 9,380 media tokens.


At the published price of $0.15 per million input tokens, the estimated direct inference expense would be approximately $0.00141 per minute of 480p video with sound, excluding accompanying instructions, additional inputs, infrastructure, and any operational overhead.


For 100,000 one-minute recordings at the same tokenization rate, that produces an estimated model charge of approximately $140.70. At the documented maximum visual tokenization rate, the corresponding media-only estimate would increase to approximately $242.70.


These are calculations derived from Cloudflare's pricing and tokenization rules, not independently measured billing results. Actual charges depend on the encoded media, resolution, processing configuration, and the additional context supplied with each request.


For organizations processing large numbers of support calls, short inspection videos, or automated moderation tasks, such rates may make native multimodal classification economically attractive. However, a direct comparison with alternative pipelines must include the complete cost of transcription, preprocessing, storage, engineering maintenance, and validation.


Cloudflare's pricing advantage is most commercially relevant when Clef-omni can replace several inference stages without producing enough additional classification errors to offset the savings.


A low price per request is not automatically equivalent to a low cost per correct decision. If the model generates unreliable results for a particular application, the expense of human review, repeated processing, or incorrect automated actions may exceed the savings achieved through cheaper inference.


··········


DEVELOPERS CAN USE CLEF-OMNI THROUGH WORKERS AI OR DEPLOY THE OPEN WEIGHTS IN THEIR OWN INFRASTRUCTURE


Cloudflare provides a hosted decision API and downloadable model weights, allowing developers to choose between managed inference and greater control over execution infrastructure.


The hosted model is available through Cloudflare Workers AI using the identifier @cf/cloudflare/clef-omni.


Its API follows the System One structure already used by Clef and Clef-flash, allowing applications to submit a state describing the situation and a collection of typed questions specifying the decisions required.


Three primary question types are supported.


The noul type represents a binary decision, returning the estimated probability of a true outcome. The choice type evaluates a set of named alternatives, returning a selected option and the associated probabilities. The score type evaluates ordered categories and produces a probability-weighted score.


A customer-support application might submit a conversation, an optional audio recording, and three questions: whether the issue is urgent, which department should handle it, and how severe its business impact appears to be.


The application receives structured answers that can be integrated into ticketing systems, dashboards, or workflow engines without parsing a lengthy natural-language explanation.


The API accepts up to 64 typed questions within a single request, providing a way to evaluate several related aspects of the same evidence while avoiding repeated submissions of identical context.


Media files are supplied as embedded data rather than remote URLs in the hosted interface. Developers therefore need to handle file acquisition, encoding, and the applicable size limits before calling the model.


Clef-omni is also compatible with Cloudflare AI Gateway, making it possible to incorporate decision scoring into a broader inference architecture that may use different models for planning, generation, execution, and verification.


For organizations preferring to control their own infrastructure, Cloudflare has published the model weights and associated decision-head implementation on Hugging Face under Apache 2.0.


Self-hosting permits deployment outside Cloudflare's managed inference environment, subject to hardware requirements and the implementation's technical dependencies. It also introduces responsibility for model serving, performance optimization, security updates, monitoring, and infrastructure availability.


The published model card describes a reference implementation using PyTorch and Transformers, with facilities for encoding structured decision requests and applying the joint scoring head. Support for additional serving frameworks is still developing, so compatibility should be validated against the actual release and deployment configuration rather than assumed from the underlying Qwen architecture.


Open-weight availability improves deployment flexibility, but it does not remove the operational complexity of running a 30B-parameter multimodal system. Organizations must compare managed API pricing against the cost of dedicated GPUs, utilization, engineering maintenance, and the volume of requests they expect to process.


The hosted service will generally be simpler for applications with variable workloads or limited infrastructure expertise, while self-hosting may be attractive where organizations require additional control over processing environments, deployment location, or model customization.


··········


CLEF-OMNI'S LONG-TERM VALUE DEPENDS ON RELIABLE DECISIONS, NOT GENERAL-PURPOSE AI GENERATION


Cloudflare is entering a specialized segment of the AI market in which models are optimized to evaluate structured choices rather than generate unrestricted content.


This positioning differs from the commercial strategies of general-purpose assistant providers. A coding model, research assistant, or conversational agent may need to generate substantial text, reason through unfamiliar problems, and plan multi-step operations. A decision model instead receives a bounded set of alternatives and evaluates which outcome is best supported by the available information.


The distinction becomes important in applications where a larger AI system repeatedly needs to select tools, classify incoming information, assess intermediate results, or determine whether a task requires further review.


In such environments, using a frontier conversational model for every classification may impose unnecessary inference costs and latency. A specialized decision model can perform those narrower operations, leaving more computationally intensive reasoning to systems that need it.


Cloudflare has already described internal use cases for its Clef family, including spam identification in public GitHub documentation issues, phishing-related moderation within its EmDash content-management platform, sensitive-information detection, and malicious-domain analysis.


These examples demonstrate the relevance of the decision-model approach to Cloudflare's own services, although they should not all be interpreted as production deployments of Clef-omni specifically. Several began with earlier members of the Clef family.


The addition of native audio and video broadens the potential applications to customer-interaction analysis, inspection workflows, multimedia moderation, and operational monitoring.


Nevertheless, classification accuracy remains dependent on the quality of the available evidence and the suitability of the predefined answer categories. A model asked to determine whether a machine is operating normally may return a confident prediction despite having no reliable evidence about its expected acoustic behavior or mechanical configuration.


Similarly, a moderation system that evaluates a short video may overlook information outside the sampled frames, while a customer-service classifier may misinterpret sarcasm, overlapping speech, or unfamiliar terminology.


These problems require application-level validation, confidence thresholds, and appropriate mechanisms for escalation. In consequential settings, probability scores should inform authorization policies rather than replace them, particularly where an automated decision could affect customers, security operations, financial transactions, or access to services.


The published benchmark results also indicate that model selection should remain workload-specific. Clef-omni performs well on several established classification tests but does not consistently outperform Clef or Clef-flash, while its hosted price sits between those alternatives.


The strongest justification for adopting it is therefore the combination of multimodal input, structured output, and competitive inference cost—not an assumption that it is the most accurate decision model for every task.


Cloudflare's October 9 release also demonstrates that serving infrastructure and commercial packaging can evolve independently of model training. Clef-flash became cheaper partly through adjustments to its hosted context window, while Clef became faster through runtime improvements without a new set of model weights.


For companies deploying AI at scale, these changes are economically meaningful because even modest differences in the cost or latency of repeated decisions can accumulate across millions of transactions.


Clef-omni extends this approach to recordings and visual information, offering developers a way to turn multimodal evidence into bounded, machine-readable decisions without generating intermediate transcripts or lengthy explanations. Its ultimate commercial advantage will depend on whether the resulting classifications remain accurate, well calibrated, and sufficiently inexpensive when evaluated against complete production workflows rather than isolated model calls.


··········


FOLLOW US FOR MORE.


DATA STUDIOS


datastudios.org

bottom of page