Cerebras Teams Up With Gimlet Labs For AI Inference Cloud - First Data Center Due This Year
Article excerpt
Highlighted: the sentence this signal was extracted from
Cerebras Systems Inc. ( CBRS ) is teaming up with the applied AI research and infrastructure startup Gimlet Labs to bring its wafer-scale AI systems into Gimlet Cloud. With this collaboration, the companies are targeting inference speeds of up to 3,000 tokens per second for agentic and real-time AI applications. The first Cerebras-powered Gimlet Cloud data center is expected to come online later this year. The collaboration will span infrastructure through developer application programming interfaces (APIs), as they work to make faster AI inference available at production scale. Inference is the process through which a trained AI model produces responses or other outputs. For applications such as AI agents, voice and video systems, and assistants, how quickly those outputs are generated can directly affect how responsive the application feels. Gimlet Cloud will combine Cerebras' Wafer Scale Engine with graphics processing units (GPUs) rather than relying on one type of processor for every part of the workload. Gimlet's technology divides AI inference into distinct phases and directs each phase to the computing hardware best suited to perform it. This approach combines Cerebras' fast token generation with the high throughput offered by GPUs. The companies plan to deliver up to 3,000 tokens per...
Keep reading with a free account
The rest of this article, and every signal for Cerebras, is in your free account.
