Tired of ads? Enjoy an ad-free experience by signing up.
👩‍🍳 How we use AI at Tech in Asia, thoughtfully and responsibly.
🧔‍♂️ A friendly human may check it before it goes live. More news here

Nvidia says Groq 3 LPX rack enters full production

Nvidia said on August 24 that its Groq 3 LPX rack is in full production, after its US$20 billion deal for Groq assets.

Nvidia said Nebius, a Netherlands-based cloud provider, will deploy the system later this year alongside Nvidia’s Vera central processing units and Rubin graphics processing units for low-latency AI inference.

Nvidia said each rack contains 256 Groq 3 chips manufactured by Samsung.

Citing an Artificial Analysis benchmark, Nvidia said the system delivers 3,400 tokens per second.

It added that OpenAI’s Ultrafast mode is listed at 750 tokens per second and runs on Cerebras.

Nvidia said demand is growing for chips that handle the decode stage of inference, which helps generate model outputs quickly.

The company said the Groq system would complement graphics processing units rather than replace them.

🔗 Source: CNBC

Recent Nvidia developments

Stay ahead in Asia’s tech landscape

You've reached your 2 free content limit for the month. Sign up for free to read the full story.

🏄 For casual readers / 👶 Free

Basic

US$0

Free forever

Get instant access to this article and more every month

0 premium content

Unlimited news briefs

5

5 articles

Ad-free reading experience

Just US$0 per day

⌛Sign up in 20s. No payment details needed.

📖 For learners / 👍 Starter

Lite

US$4.92/month

Billed annually at US$59/year

Get instant access to this article and more every month

4

4 premium content

Unlimited news briefs & articles

Ad-free reading experience

Just US$0.17 per day

Cancel anytime

Our subscriber community includes professionals from these companies:

Stay updated on the go with our mobile app.

Get latest insights with smoother, more personalized experience through TIA mobile app.