Cloudflare Clef-omni turns multimodal routing into one API call
The open-weight model scores predefined decisions across text, images, audio and video. We would test it as a pipeline replacement, not trust it as an autonomous gate.
What shipped
Cloudflare released Clef-omni on October 9, an open-weight model that accepts text, images, audio and video in one request and returns probabilities for predefined options. Workers AI prices it at $0.15 per million input tokens, with free output and a 64,000-token context window.
The operator math
The interesting possibility is not another multimodal interface but replacing several sequential services with one call. Cloudflare reports median latency of about 130 milliseconds for text decisions and roughly 1.5 seconds for a 21-second video with sound, but teams should treat those as vendor benchmarks until reproduced on their own traffic.
What we would do
We would start with request routing, incident triage or output verification, running Clef-omni beside the existing system. Self-hosting reduces vendor dependence, but the model card says its BF16 backbone needs about 64 GB of GPU memory, so open weights do not automatically mean cheaper production.
Clef-omni matters because it attacks the least glamorous but often most expensive part of multimodal software: the chain of transcription, frame extraction, classification and output parsing that sits before a real action. If one schema-bound model can score text, images, audio and video together, the win is not a prettier demo; it is fewer services, fewer retries and a lower cost per completed triage or routing task. Virtual Arc would not rebuild a production control plane around it yet. Cloudflare's quality and latency numbers are vendor-run, calibration can drift on live traffic, and a wrong probability can still trigger the wrong tool. We would put it in shadow mode against one bounded workflow, compare end-to-end task cost and false-action rates with the current pipeline, and only promote it behind deterministic thresholds, audit logs and a human-review escape hatch. The Apache-2.0 weights make that pilot more attractive because success does not force permanent API lock-in.