JS Wei (Jack) Sun

Cerebras hits 1,851 tok/s on Gemma 4, but the voice-agent bill is TTS not LLM

Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.

← Back to the issue

Sources

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI huggingface.co

References

AlphaSignal alphasignal.ai

Cerebras currently delivers the fastest inference for Gemma 4 31B, reaching a record 1,851 output tokens per second… approximately 35x faster than traditional GPU-based providers… time-to-first-token (TTFT) of roughly 1.5 seconds, even when the model’s native reasoning mode is enabled.

LiveKit pricing / voice agent cost breakdown livekit.com

Total estimated cost to run a voice agent using Gemma 4 31B is approximately $0.0735 per minute… the Cerebras LLM portion is remarkably cheap (~$0.0014/min), while the bulk of the cost is driven by high-quality TTS (e.g., Cartesia at $0.03/min) and the LiveKit session fee.

Jeff Geerling blog — Reachy Mini review jeffgeerling.com

While the hardware is impressive, replicating the ‘utterly trivial’ demos shown at major keynotes like CES is often more difficult than marketed… children quickly anthropomorphize the robot, leading them to share sensitive personal information with the AI, which is often processed by third-party cloud services like OpenAI.

AI Weekly — HN/community reaction roundup aiweekly.co

Cascaded pipelines require complex signaling logic to shut down the TTS immediately when new audio is detected… skeptics point out the lack of published head-to-head benchmarks against other production stacks, leaving ‘dramatically faster’ as a largely unverified marketing claim.

The Robot Report therobotreport.com

Hugging Face launched an ‘agentic toolkit’ in May 2026, allowing users to describe desired behaviors in plain English while an AI agent generates and ships the necessary code to the robot in under an hour.

r/singularity — Gemma 4 benchmarks discussion reddit.com

While the model excels in general intelligence and reasoning, some users on platforms like Reddit argue it may not retain the same depth of factual knowledge as much larger models… local testers observed performance degradation after 10k–20k tokens, particularly in precise code recall tasks.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare