Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
What happened
In fact, the story of Clef came together over the course of less than a week. This changes the paradigm for decision models, which have been largely text-only since the debut of Jev from TypeSafe.
Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality. In production, Clef-omni executes a quick prefill pass across the complete payload, scoring all modalities and valid parameter options simultaneously.
Media elements map straight into the unified sequence, where video and audio are synced with visual frames for joint processing. The result is a model that is resilient and designed to succeed regardless of schema variations, field ordering, and prompt structures.
Sources & evidence
- Cloudflare Blog Primary / official
Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash ↗
https://blog.cloudflare.com/clef-faster-cheaper-multimodal/