AI Inference: Meta Teams with Cerebras on Llama API

Sunnyvale, CA — Meta has teamed with Cerebras on AI inference in Meta’s new Llama API, combining Meta’s open-source Llama fashions with inference know-how from Cerebras.

Builders constructing on the Llama 4 Cerebras mannequin within the API can count on speeds as much as 18 instances sooner than conventional GPU-based options, based on Cerebras. “This acceleration unlocks a completely new technology of purposes which are unimaginable to construct on different know-how. Conversational low latency voice, interactive code technology, prompt multi-step reasoning, and real-time brokers — all of which require chaining a number of LLM calls — can now be accomplished in seconds moderately than minutes,” Cerebras stated.

By partnering with Meta to serve Llama fashions from Meta’s new API service, Cerebras positive factors publicity to an expanded developer viewers and deepens its enterprise and partnership with Meta and their unimaginable groups.

Since launching its inference options in 2024, Cerebras has delivered the world’s quickest Llama inference, serving billions of tokens via its personal AI infrastructure. The broad developer group now has direct entry to a strong, OpenAI-class various for constructing clever, real-time methods — backed by Cerebras velocity and scale.

“Cerebras is proud to make Llama API the quickest inference API on the earth,” stated Andrew Feldman, CEO and co-founder of Cerebras. “Builders constructing agentic and real-time apps want velocity. With Cerebras on Llama API, they’ll construct AI methods which are basically out of attain for main GPU-based inference clouds.”

Cerebras is the quickest AI inference answer as measured by third occasion benchmarking website Synthetic Evaluation, reaching over 2,600 token/s for Llama 4 Scout in comparison with ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.

Builders will be capable to entry to the quickest Llama 4 inference by deciding on Cerebras from the mannequin choices inside the Llama API. This streamlined expertise will make it straightforward to prototype, construct, and scale real-time AI purposes. To join early entry to the Llama API and to expertise Cerebras velocity as we speak, go to www.cerebras.ai/inference.

Source link

Optimizing DevOps for Large Enterprise Environments

Datavault AI to Deploy AI-Driven HPC for Biofuel R&D

Voltage Park Partners with VAST Data

The Future of Humanity with AI: A New Era of Possibilities. | by Melanie Lobrigo | Apr, 2025

How AI Is Transforming the SEO Landscape — and Why You Need to Adapt

Adding Training Noise To Improve Detections In Transformers

Meta Layoffs Begin: Inside Meta’s Rankings of Low Performers

What is ANOVA? Types of ANOVA and Their Applications | by Meriç Özcan | Feb, 2025

Most Popular

How do you teach an AI model to give therapy?

Sama Launches Agentic Capture for Multi-Modal Agentic AI

This Is the Most Underrated Leadership Skill in 2025

Our Picks

What Building an App Taught Me About Parenting — And Successful Startups

Making extra long AI videos with Hunyuan Image to Video and RIFLEx | by Guillaume Bieler | Mar, 2025

Ever Wondered What’s in a Neural Network Summary? Let’s Break It Down Together! | by Saketh Yalamanchili | May, 2025

AI Inference: Meta Teams with Cerebras on Llama API

Related Posts