OpenAI and Cerebras Confirm 750MW AI Inference Deployment Through 2028

OpenAI and Cerebras have announced a multi-year partnership to deploy 750MW of wafer-scale AI compute capacity to support ultra-low-latency inference. The rollout will occur in three 250MW phases, with completion scheduled by the end of 2028. This infrastructure is specifically designed to enhance real-time AI interactions, such as conversational assistants and coding tools, rather than model training. While an SEC filing notes early operational use for an OpenAI Codex Spark model, the companies have not yet disclosed specific product integration plans, pricing, or performance metrics. The agreement represents a significant long-term commitment to scaling AI responsiveness, though the immediate impact on individual customer experiences remains to be seen. Businesses are advised to evaluate their specific technical requirements rather than assuming this infrastructure expansion will automatically resolve all latency or cost challenges.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Your tool returned the rows. The model counted them wrong.
A recent investigation into AI agent frameworks reveals a critical reliability issue: LLMs often fail to accurately count items when provided with raw…
AI Update Overview: Opus 5.5, Fable 5.1, GPT-6, DeepSeek, and Grok
Over the past five weeks, the generative AI market has seen significant shifts with the release of new models, including Claude Fable 5.1, Opus 5.5, a…
Sean Parker, the former Napster co-founder and early Facebook executive, is spearheading a significant strategic pivot for Stability AI. After previou…



