At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeno, including the first batch of benchmark results for the new system. Tested on SemiAnalysis' InferenceX benchmark, Jalapeno registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors. 'The bottom line is that the results show a very, very significant performance advance over state of the art,' said Richard Ho, OpenAI's head of hardware, in a press call. 'Jalapeno can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency.' Notably, that comparison is against an Nvidia Blackwell system ' but by the time Jalapeno reaches full deployment, the competition may have advanced significantly. Ho estimated that Jalapeno would deploy at the end of 2026 'in very small volumes,' with more significant deployment coming in...
learn more