Technologies
Back
Artificial Intelligence & Machine Learning

Qwen 2.5 72B and 32B models now available on Cerebras at 1500 tokens/s

Hacker News (YC)
Advertisement468 × 90
Qwen 2.5 72B and 32B models now available on Cerebras at 1500 tokens/s

Cerebras Systems has announced the availability of the Qwen 2.5 72B and 32B models on its inference platform. By leveraging the company's specialized Wafer-Scale Engine technology, these models achieve high-performance inference speeds of up to 1500 tokens per second. This integration aims to provide developers with access to state-of-the-art open-weights models while bypassing the latency bottlenecks often associated with traditional GPU-based cloud infrastructure. The Cerebras Inference API is designed to support high-throughput applications, making it suitable for real-time AI agents, complex data processing, and large-scale enterprise deployments. By optimizing the hardware-software stack specifically for transformer-based architectures, Cerebras continues to position itself as a high-speed alternative for organizations looking to scale their generative AI capabilities without compromising on speed or model complexity.

This is a summary. Read the full article at the original source:

Hacker News (YC)
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250