Nvidia has announced that its Groq 3 LPX inference chip can process 3,400 tokens per second on the Gemma 4 31B model, reportedly achieving speeds four times faster than competitor Cerebras, according to The Decoder. This performance highlights Nvidia’s push in AI chip efficiency.

However, The Register notes that Nvidia’s impressive token throughput requires a minimum of 64 accelerators to reach this level, whereas Cerebras achieves similar results with only one or two accelerators. This contrast underscores differences in hardware scalability and deployment strategies between the two companies.

For Japanese markets, where AI and machine learning applications are rapidly growing, such developments in inference chip technology could influence investment decisions in semiconductor and AI-related equities, as well as impact FX flows linked to tech sector performance.