AI DevelopmentNew Release10 min readPublished October 9, 2026

Preserve the evidence that a transcript can lose

Clef Omni: Audio and Video Decisions for AI Workflows

Assess Clef Omni for audio and video decisions: native inputs, endpoint differences, reviewable evidence and a fair test against transcription pipelines.

DA
Digital Applied Team
Research and practical guidance
PublishedOctober 9, 2026
Read time10 min
SourcesPrimary documentation

Some business decisions depend on what a recording sounds or looks like, not just the words in its transcript. Clef Omni brings native audio and video input to Cloudflare’s decision-model family. Its useful role is a bounded question over that evidence, with an explicit review path. It is not a substitute for defining what the evidence can actually support.

Product documentation and current access details were checked on October 11, 2026. The workflow examples are proposed evaluation designs, not reports of completed customer deployments.

Key takeaways
  1. 01
    Native input changes the evidence pathAudio and video can reach the decision model without a separate transcript-only bottleneck.
  2. 02
    Verify the endpointCloudflare documents audio and video while current OpenRouter metadata lists text and image inputs.
  3. 03
    Ask bounded questionsUse visible or audible conditions with clear options and a review path.
  4. 04
    Compare complete pipelinesInclude preprocessing, failed requests and human review rather than ranking vendor latency claims.

01 — New capabilityWhat Clef Omni adds to decision models

Cloudflare’s October 9 announcement describes Clef Omni as an open-weight decision model accepting text, images, audio and video. It is built on the comprehension side of Qwen3-Omni-30B-A3B-Instruct and returns schema-bound scores rather than generated speech. The architecture is relevant because evidence can be evaluated together instead of first being reduced to a transcript.

The same announcement reports benchmark and latency results. Those are Cloudflare’s measurements, not an independent replication by Digital Applied. This guide does not use them to declare that native media will always be faster or more accurate than a staged pipeline. The decision depends on what information the task needs and how the whole application prepares and handles it.

A useful example is classifying whether a short product demonstration contains a visible opening action and an audible click. A transcript may contain no words at all, while a single still frame may miss the action. That makes the example a candidate for native media evaluation. It does not establish that the model can diagnose the product’s mechanical condition or approve its safety from the clip.

Words carry the answer
Transcript-based decision
Inspectable text

Route a request by the words spoken when nonverbal media is not necessary for the decision.

Simple evidence
Sound and motion matter
Native media decision
Richer input

Evaluate a bounded audible or visible condition while preserving a path for uncertainty.

Different evidence

02 — Endpoint contractCheck the serving interface before designing the workflow

The Cloudflare model documentation is the relevant starting point for its native serving interface. The announcement demonstrates audio and video inputs alongside text and images. By contrast, the October 11 OpenRouter catalogue metadata advertises text and image input for its route. These should not be treated as interchangeable capability statements.

That difference can have several operational explanations, but this guide does not guess which applies. The practical response is to verify your actual endpoint with the required media type and current documentation. A model name alone is not an integration contract. Do not build an audio workflow around an aggregator field that currently describes only text and images, then assume a failed request is a model-quality problem.

Record the provider, model identifier, input format and any relevant file limits in the trial. Use a small known-good media sample before sending a large evaluation set. If the request fails, preserve the failure category separately from a wrong model decision. Unsupported input, an unreadable file and an incorrect classification are different failures requiring different repairs.

The capability boundary is explicit

Native audio/video support is documented by Cloudflare. The October 11 OpenRouter input metadata is narrower. Confirm the route you will use rather than transferring every model capability to every endpoint.

03 — Decision scopeAsk what the recording can establish

Write the question in terms of observable evidence. Is the package label visible? Does the clip contain a click after the lid closes? Does the speaker explicitly request a return? These are bounded questions. Is the customer honest, is the product safe, or will the machine fail next month are much broader inferences that a short recording may not support.

Keep the answer choices aligned with the evidence quality. A clip with background noise or an obscured subject may require review rather than a forced yes or no. Where the interface returns scores over fixed options, the application still needs a policy for low-confidence or insufficient evidence. An uncertainty path should be part of the workflow design, not added after a confident mistake reaches production.

The decision-model primer explains why typed answers are attractive to software. The same structure can hide a badly framed question if the developer mistakes schema validity for factual support. A neatly returned label is useful only when the task definition, available evidence and downstream consequence fit together.

For the fictional lid example, distinguish three labels in the business policy: audible click present, audible click absent and evidence insufficient. A loud background sound should not become evidence that the lid failed, and a missing beginning should not become evidence that the action never happened. Reviewers should agree on those definitions before labelling the examples. If they cannot, the uncertainty belongs in the task design. A model cannot be expected to produce a dependable distinction that the people responsible for the workflow have not settled.

  • State the visible or audible condition precisely.
  • Separate missing evidence from a negative observation.
  • Keep high-consequence judgments outside the first trial.

04 — Pipeline tradeoffCompare native media with the evidence a cascade preserves

A staged pipeline can transcribe speech, extract selected frames and pass those representations to a decision model. This has practical advantages: people can inspect intermediate text, reuse a transcript and identify where a failure occurred. It can also discard timing, non-speech sound or motion that the decision actually needs. The correct architecture follows that information requirement.

For a request whose category depends only on spoken words, a transcript-based route may be entirely adequate. For the product-click example, transcription alone is an incomplete representation. A native-media route can retain more of the relevant evidence, but it also requires a way for reviewers to locate and inspect that evidence when a decision is disputed. Richer input does not remove the need for an understandable review process.

Our transcript-verification guide is useful when text remains in the pipeline. Compare what is lost or preserved at each stage rather than assuming that fewer model calls automatically mean a better result. A one-call system can still require file preparation, retries and review; a multi-step system can still be reliable when its intermediate representations preserve everything the task requires.

Proposed architecture comparison. These are task-design considerations, not measured Clef Omni performance results.
QuestionTranscript or stills may sufficeNative media is worth testing
What did the customer request?Clear spoken wordingSpeech is mixed with relevant nonverbal evidence
Is the label visible?A suitable still imageVisibility changes during the clip
Did the action make a sound?A dedicated sound representationTiming between motion and audio matters
Why was this decision disputed?Inspectable intermediate textReviewer needs the original segment and context

05 — Evaluation inputsDesign media examples that expose the difference

Build pairs of examples where the words stay the same but the relevant sound or motion changes. A silent closing action and one with an audible click can test whether the model is using the intended evidence. Conversely, change irrelevant background details while keeping the target action the same. The decision should remain stable when the task meaning has not changed.

Include poor-quality inputs deliberately: background noise, an action partly off camera, a clipped beginning and a long interval before the relevant event. Record whether the request was technically accepted and whether the evidence was sufficient for the expected label. Do not score every uncertain result as failure if the correct operational response is to ask for a clearer recording.

Keep the examples synthetic or properly authorized for the evaluation. The point is to isolate task behavior, not to accumulate sensitive recordings. Use the same source material for both the native and staged routes, and document any transformations applied to it. If one route receives a manually selected ideal frame while the other receives the whole clip, the comparison has changed the evidence as well as the model.

Include a synchronization case. Use a clip where the relevant sound occurs before the visible action, then another where it occurs at the expected moment. Whether that timing changes the answer depends on the question, so define it explicitly. The case is valuable because a pipeline that independently summarizes audio and video may preserve both events while losing their relationship. A native route may preserve that relationship better, but the paired test is needed before making that claim about the actual endpoint.

A useful contrast pair

Keep the spoken instruction identical while changing the audible event the question asks about. This tests whether the decision follows the relevant media rather than a shortcut in the text.

06 — TraceabilityKeep review evidence connected to the result

A structured label is easy to store, but a reviewer needs the context that produced it. Preserve a reference to the source asset, the question definition, the model route and the result. If your workflow trims a clip or extracts frames, record which segment or frames were used. Otherwise, a later reviewer may inspect a different representation and be unable to explain the disagreement.

Do not invent an explanatory transcript if the model did not generate one. A separate explanation can be useful, but it is another model output that needs its own evidence check. For simple operational decisions, the better review interface may be the original clip, the precise question and the returned scores. The person can then judge the same observable condition without treating generated prose as proof.

This distinction matters when building an automated creative or support workflow. A reviewer should be able to accept, reject or request a clearer input, and the system should preserve that outcome. Our AI transformation service helps connect model outputs to these review and ownership steps so a new modality becomes useful work rather than an isolated API experiment.

Decide how long the review evidence needs to remain available for the workflow, and apply the organization’s actual retention rules. A stored decision whose source asset has disappeared may be impossible to investigate later. Conversely, keeping every recording indefinitely is not required merely because a model processed it. The operational design should state what is retained, who can inspect it and how a disputed result can be reviewed while the evidence remains available. This is a records-management choice, not a capability supplied automatically by a multimodal model.

  • Preserve the original asset reference and any transformations.
  • Keep the exact decision question with the result.
  • Record the reviewer’s outcome separately from the model’s score.

07 — Operational costPrice and time the whole media decision

Cloudflare’s announcement lists Clef Omni at $0.15 per million input tokens. Media-to-token accounting and serving conditions matter, so do not convert that number into a universal cost per minute of audio or video. Consult the current pricing rules for the selected endpoint and use its reported usage for representative requests. A raw token rate is not yet a media budget.

Compare the native route with the full staged route. The latter may include transcription, frame extraction and a decision request; the former may include upload preparation and a different media billing rule. Both can incur failed attempts and human review. The audio/video processing cost guide explains why the unit of work needs to remain explicit across these layers.

Measure elapsed time from usable input to usable decision under the same conditions. Include the slow cases and media failures rather than selecting only successful short clips. Cloudflare’s published latency examples can help set expectations for a trial, but they are not a service guarantee for your files, network or workflow. This guide does not claim that the proposed comparison has already been run.

Avoid false unit conversions

A price per million input tokens is not a price per million seconds, frames or files. Preserve the provider’s media accounting rules before converting a workload into a budget.

08 — Routing outcomeAdopt multimodality where it changes the decision

The strongest reason to use Clef Omni is that a task needs evidence a text-only representation would lose. That can justify a native-media trial even when the current process already classifies transcripts well. The weakest reason is simply that a model accepts more input types. Sending more information is not helpful when the task remains poorly defined or the additional media is irrelevant.

Write a routing rule based on the evidence requirement. Word-only intent routing may stay on an established text path, while a bounded sound-and-motion check uses the verified native endpoint. Keep an explicit review outcome for obscured or ambiguous recordings. This gives the team a maintainable reason for each route and a clear place to investigate errors.

Cloudflare’s release expands the choices available for structured multimodal decisions. Your workflow should still determine whether that capability belongs in production. Start with a small question, preserve the media and evaluate the complete process. A narrow, well-supported use is more valuable than a broad claim that a single model can understand everything in a recording.

The first rollout can simply flag candidate clips for a person. That keeps the observable decision useful while avoiding an early leap from classification to a consequential action. Expand only after the team has inspected the uncertain and disputed cases, because those cases reveal whether the media question is actually well framed.

  • Use native media when the task needs its extra evidence.
  • Keep endpoint support and billing conditions explicit.
  • Retain a review path for insufficient or ambiguous inputs.
Your next step

Choose a question whose answer needs the media

Build a paired test where sound or motion changes the correct label. Compare the native and staged routes using the same source assets and inspect both the decisions and the review burden.

Adopt the capability where it preserves useful evidence. Keep the decision scope narrow enough that a reviewer can explain what the recording supports.

Put the method to work

Build a workflow your team can verify

Digital Applied helps teams turn a promising AI capability into a clear operating process, with useful evaluations, review points and a practical path to production.

Workflow designPractical evaluationsClear ownership
Work with us

From trial to useful work

  • →Define the task and its acceptance criteria
  • →Connect the right information and tools
  • →Review failures before expanding access
FAQ · Practical decisions

Questions before you start

The capability discussed here is structured decision scoring over multimodal input. Cloudflare describes using the comprehension backbone and discarding speech-output components.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue exploring

AI Development

Open Decision Models Compared: Clef, Decider 2B and Jev

Cloudflare's Clef and Amazon's Decider 2B put decision models like closed Jev into open weights. Licence, size, context, latency and benchmarks in one table.

October 1, 2026 · 6 minRead
AI Development

Microsoft Decision 1 vs OpenAI Decisions API: A Guide

Compare Microsoft Decision 1 and OpenAI Decisions API by access, answer types, probability handling, evaluation design and the cost of a useful decision.

October 9, 2026 · 10 minRead
AI Development

Managed RAG Services Compared: Cloudflare, Google, AWS

Cloudflare AI Search is now generally available. How it compares with Google Agent Search, Bedrock Knowledge Bases and OpenAI file search on price and limits.

October 2, 2026 · 6 minRead
AI Development

Web Search APIs for AI Agents Compared: Price and Limits

Cloudflare added a Web Search API to AI Gateway on October 2. How it compares with search tools from OpenAI, Anthropic, Google and others on price and limits.

October 2, 2026 · 6 minRead
AI Development

Cloudflare Blocks AI Agents on Ad Pages: Which Bots Are Hit

From September 15, 2026 new ad-supported Cloudflare domains block AI agents on ad pages and refuse AI training by default. A 20-bot census of who is affected.

September 15, 2026 · 10 minRead
AI Development

OpenAI + Dell Codex: On-Premises Enterprise Agents

OpenAI and Dell partner to bring Codex to hybrid and on-premises environments via Dell AI Factory. What changes for enterprise coding workflows.

May 18, 2026 · 12 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source