Cerebras is supplying its wafer-scale processors to run a new high-speed tier for OpenAI's GPT-5.6 Sol model, delivering 750 tokens per second in limited preview for API customers. The performance exceeds standard processing by 14 times, according to the announcement.
The Ultrafast tier uses Cerebras' WSE-3 architecture, which houses 44 gigabytes of on-chip SRAM in a single wafer. This memory density eliminates traditional bottlenecks that force processors to shuttle data between separate chips, allowing inference workloads to run at speeds that favor real-time applications like interactive agents and live translation services.
OpenAI disclosed the partnership as part of a broader push to offer inference speed as a product tier, much as cloud providers offer compute sizes or regions. The move segments its customer base along latency requirements: users running batch jobs on GPT-5.6 Sol pay standard rates, while those needing sub-second response times access the Ultrafast tier at a likely premium. Cerebras' previous partnerships with other model providers show the company positioning itself as the hardware layer for speed-sensitive workloads rather than competing on model architecture.
The 14-fold speedup is rare in AI inference. Typical performance gains from architectural changes land in the 2x to 5x range; Cerebras claims such large jumps because its on-chip memory model reduces the memory-access latency that normally dominates inference time. The company's wafers are designed to hold entire models in SRAM during inference, not just portions of them, flattening the access curve.

Limited preview status means the tier is not yet available to all OpenAI API customers. OpenAI typically restricts new inference modes to a small cohort first to measure real-world demand, cost structure, and thermal performance under load before wider rollout. The duration of preview phases for OpenAI products has ranged from weeks to months, and no timeline was announced.
Cerebras, which went public via SPAC in 2021, has struggled to demonstrate revenue scale from its wafer-scale approach. The Ultrafast tier partnership is the company's highest-profile application win to date. If the tier reaches production at scale, it would represent the first major U.S. cloud inference service built on Cerebras hardware rather than GPUs or custom TPUs.
The number to watch is the migration rate of OpenAI's existing API customers into the Ultrafast tier once it exits preview. Adoption near or above 10 percent of the customer base would reveal whether demand exists for latency-optimized inference among OpenAI's user base.