In one week of August 2026, two headlines landed that looked unrelated: Anthropic started building a team to design its own chips, and AMD bought a startup that bakes an AI model straight into the silicon. They are the same story. Both are attacks on a problem most coverage never names — the memory wall, the point where the world's fastest chip sits idle, waiting for data to arrive. Hover or tap any underlined term.
On August 5, 2026, Anthropic publicly confirmed it is assembling an in-house team to design its own chips — posting silicon-engineer and hardware-architect roles at $320,000–$485,000, anchored by Clive Chan, who joined in June. The stated goal isn't to become a chipmaker; it's to co-design the hardware and Claude together, tailoring the chip to the model so each inference costs roughly half as much. Anthropic already runs on a mix of Google's TPUs, Amazon's Trainium, and Nvidia GPUs; now it wants a say in the transistors themselves.
A day later, on August 6, AMD announced it is buying Taalas, a Toronto startup founded in 2023. Taalas does something radical: instead of loading a model's weights from memory for every answer, it etches the model directly into the chip's transistors. Its first test chip, HC1, ran Meta's Llama 3.1 (8B) at a claimed ~17,000 tokens/second — which Taalas said in February is ~73× an Nvidia H200 at one-tenth the power. AMD plans to slot Taalas chips into its Helios rack systems (in mass production since July, with Microsoft, OpenAI, and Anthropic lined up for Q4).
For years the story was “we need faster chips.” We got them — almost too good. The processor can now do the math far faster than the system can hand it the numbers to work on. So the expensive chip sits there, idle, waiting for data to be carried over from memory. The waiting — not the calculating — is now the limit.
Why does AI hit this wall so hard? Because answering a prompt means reading billions of weights from memory for every single token of the reply. The math is cheap; the fetching is not. That's why the whole industry is suddenly obsessed with memory, and why these two August moves are really one move.
Most coverage asks “can Anthropic really build a chip?” or “is this bad for Nvidia?” Those are horse-race questions, and horse races are hard to call. The Lens question is the one that pays: if the bottleneck is memory, who owns the memory bottleneck?
| The layer | Who | Why they win either way |
|---|---|---|
| The memory makers | SK Hynix, Samsung, Micron | Just three companies make nearly all the world's HBM, and it's sold out through 2026. Every escape that “widens the hallway” buys more of their memory. When memory is the scarce thing, the people who make memory hold the cards. |
| The factory & the stacker | TSMC | Nvidia's GPUs, Google's TPUs, Anthropic's future chips, AMD's Taalas parts — TSMC makes them all, and its advanced packaging is what physically stitches the memory to the chip. It doesn't matter whose logo is on top. One factory sits under every path. |
| The moonshot 🌙 | New memory tiers (HBF) | Clearly speculative. SK Hynix and SanDisk are building High-Bandwidth Flash aimed straight at inference. If a genuinely new memory tier lands, it reshuffles the board — early, unproven, we'd only act on data. |
Dragonfly Lens maps the AI buildout as one connected chain — and finds the layer where the value actually pools. Plain English, every claim sourced and flagged. When we're wrong, we say so.
Join the Lens →What is the “memory wall”? It's the gap between how fast a chip can compute and how fast memory can feed it data. Processors have gotten enormously faster over the decades while memory bandwidth grew far less, so a modern AI chip spends much of its time idle, waiting for data. For AI inference — which reads billions of model weights for every word it generates — feeding the chip, not the math, is the real bottleneck.
Why is Anthropic designing its own chips? To co-design the hardware and Claude together so each answer costs less — Anthropic is targeting roughly a 50% cut in per-token inference cost. It confirmed the in-house chip team on August 5, 2026, with silicon roles paying $320K–$485K. It's a multi-year effort, and Anthropic still uses Google TPUs, Amazon Trainium, and Nvidia GPUs today.
What did AMD buy Taalas for? Taalas etches an AI model's weights directly into a chip's transistors, so the model never has to be fetched from memory — extremely fast and low-power for that one model, at the cost of flexibility. AMD (Aug 6, 2026) plans to place Taalas chips in its Helios systems. It's a bet on model-specific inference silicon.
Who benefits most from the memory-wall race? The layer underneath the chip-brand fight: the three HBM memory makers (SK Hynix, Samsung, Micron), whose memory is sold out through 2026, and the foundry that manufactures and stacks nearly every advanced AI chip (TSMC). Whoever wins the chip design race, the memory and the factory get paid.
Sources: Anthropic confirms in-house chip design team; $320K–$485K silicon roles; ~50% inference-cost target; Clive Chan; multi-chip (TPU/Trainium/Nvidia) strategy — TechCrunch, Data Center Dynamics, Forbes. AMD acquires Taalas; model-into-transistors; HC1 on TSMC 6nm ~17,000 tok/s on Llama 3.1 8B (~73× H200 at 1/10 power, Taalas's Feb claim); Helios integration — CNBC, SiliconANGLE. The memory wall; HBM4 ~1.5–2 TB/s per stack; SK Hynix/Samsung/Micron oligopoly (SK Hynix ~60%); HBM sold out through 2026; SK Hynix + SanDisk High-Bandwidth Flash for inference — EE Times, Introl, 24/7 Wall St.
Educational research, not personalized investment advice. Dragonfly Lens is not a registered investment advisor. Figures are as reported by the sources above and were accurate at publication; performance multiples cited are the companies' own claims and await independent verification. Company names illustrate a structural shift in the AI compute supply chain, not buy recommendations — verify against primary filings before acting. Past performance does not guarantee future results.