← AI News

August 8, 2026 · CNBC

AMD Buys Taalas, Startup That Hardwires AI Model Weights Into Its Silicon

My take: The way Taalas works changes something fundamental about AI inference. Today, when a model processes your query, its weights (the parameters that define how it reasons) are loaded from memory on every single call. The HC1 chip permanently engraves them into the transistors themselves, eliminating that read entirely. In Taalas's own benchmarks: Llama 3.1 8B running at 16,960 tokens per second, 48 times faster than an Nvidia GPU on the same task.

The trade-off is real: each chip is locked to one model permanently. It cannot be reprogrammed. But for high-volume applications like customer service, large-scale document processing, or real-time translation, the economics of single-purpose chips start to make a lot of sense.

Worth noting: these figures come from Taalas's own internal tests before the acquisition, not from an independent third party. Read them as technology-demonstrator data until external validation arrives at production scale.

If you are building AI products, the relevant question is: when inference costs drop by several orders of magnitude, which use cases currently considered too expensive become part of the product roadmap?

Read at the source: CNBC ↗

Want to use these tools? See the unbiased reviews or back to the news.