← All digests
evening

AI hardware makers redesign memory for inference

Summary

Across both software and hardware, the constraint is shifting from raw model capability to the economics and integration work required to use it well. Organizations are responding by routing tasks to cheaper sufficient models, paying for people who can translate domain problems into deployments, and redesigning memory packages around bandwidth, heat, and data movement. The common advantage is no longer simply having the strongest component, but co-designing the surrounding system so that component can operate productively.

🛠️ Tooling & Dev ExLlamaV3 Releases

1.4.3

ExLlamaV3 1.4.3 adds preliminary support for GLM 5.2 through GlmMoeDsaForCauslLM, while warning that support remains a work in progress. It introduces partial CPU-layer expert offloading with dynamic placement, replacing the earlier expert-cache approach, plus CPU cache offloading in tensor-parallel mode. The release also improves speculative decoding with automatically calibrated draft-confidence thresholds and adds experimental quantization optimization. Other changes include mid-stream text injection for reasoning-token budgets, more accurate autosplit allocation, and a reduced, stabilized VRAM footprint during DSA prefill.

Read the source →
🤖 Agents & Coding Chase AI

How to Build Animated Websites With AI (3-Step Cycle)

The article argues that avoiding generic AI-designed sites requires a repeatable loop: collect strong references, recreate their mechanics locally, then generate and compare deliberate variations rather than relying on a one-shot prompt. Its recommended workflow uses a reusable “taste vault” of screenshots and links, then a Claude Code skill to inspect a reference URL and build a close local reproduction that can serve as a scaffold. For animated heroes, it proposes generating a composition-aware image, converting it to a subtle Seedance 2.5 video through the Higgsfield MCP, and reserving visual space for copy instead of centering every subject. The author recommends showing a still rather than video on mobile and stresses that reference cloning should be transformed heavily with original copy, assets, and design choices.

Read the source →
💬 Opinion & Essays Nates Newsletter

Executive Briefing: The $350K Job Has Three Parts and You Already Do One

Forward-deployed-engineer openings are paying unusually high salaries—OpenAI lists $162,000–$280,000 plus equity and Handshake lists $250,000–$350,000—because the role combines discovery, implementation, and ownership after deployment. The piece argues that companies themselves have not settled on a single definition, reflected in widely varying pay bands and requirements. It says candidates should treat their existing engineering, operations, or industry expertise as an asset: the domain understanding needed to choose the right production problem is not easily acquired in a bootcamp. Citing Anthropic’s analysis of 400,000 Claude Code sessions, it notes that people outside software occupations performed within a few points of software engineers on tasks that produced code, then proposes a 30-day project to demonstrate the missing third of a candidate’s skill set.

Read the source →
🏢 Industry & Business Simon Willison

Anthropic’s best AI model struggles to attract users as cheaper tools thrive

Simon Willison highlights Financial Times figures saying Anthropic’s annualized July revenue reached $65 billion, up from $47 billion in May, and that it expects Q3 profitability under the same model used to call Q2 profitable. The report says Anthropic has 6,000 customers spending at least $100,000 annually, while OpenAI’s annualized revenue has passed $40 billion after rising 35% in the quarter to date following GPT 5.6’s July launch. But Ramp billing data suggests expensive flagship models are not automatically winning spend: Anthropic’s Fable 5 represented 8.0% of July model spend, versus 28.0% for Opus 4.8 and 8.3% for Sonnet 4.6. Opus 5, released July 24, reached only 3.5%, supporting the view that price is constraining adoption despite high capability.

Read the source →
💬 Opinion & Essays Simon Willison

Quoting Drew Breunig

Drew Breunig argues that before Fable, improving a coding harness or context-management strategy often seemed unnecessary because newer models would erase many workflow shortcomings at the same or lower price. Fable changed that calculus: it is highly capable but costly enough that teams must actively route work to cheaper models. The quote says Opus, GPT 5.6, K3, and even GLM are sufficient for most code, making model selection and task allocation practical engineering concerns rather than incidental details.

Read the source →
🖥️ Hardware & Infra ServeTheHome

d-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026

d-Matrix presented Raptor as a 3D-DRAM inference accelerator aimed at the widening gap between model/KV-cache demands and conventional memory bandwidth. It contrasts SRAM’s extreme bandwidth but tiny capacity with HBM’s larger capacity but roughly 20 TB/s practical package ceiling, proposing compute stacked directly atop DRAM for about 0.3–0.4 pJ/bit vertical I/O—roughly a tenth of HBM’s chip-level energy. The company says a 72-card system with 32 GB per card can hold a frontier-class model such as Kimi K3 at one-million-token context, and claims approximately 1,000 tokens per second per user for a three-trillion-parameter-class model. Its design addresses non-aligned DRAM bank accesses with stream blocking, lowers I/O toggling through stream flipping, and uses ECC, refresh techniques, and bank chaining to preserve performance and yield under heat; the presentation claims about 20× HBM bandwidth density and 13.5× lower power per GB/s, though the article notes that real-world validation and drawbacks remain unclear.

Read the source →
🖥️ Hardware & Infra ServeTheHome

SK hynix HBM Packaging at Hot Chips 2026

Read the source →
🖥️ Hardware & Infra ServeTheHome

Samsung Evolving HBM Base Die at Hot Chips 2026

Samsung’s roadmap turns HBM’s base die from a mostly passive PHY and routing layer into a logic-rich component that can offload work from the accelerator. It says HBM5 can exceed 60 GB and 6 TB/s per stack, driving the base die toward 4 nm logic beginning with HBM4; custom HBM can use that logic for memory controllers, repair SRAM, sensors, self-test, external-memory interfaces, and eventually processing elements. The move creates thermal challenges as PHY power density rises, for which Samsung proposes a Heat Path Block that it says lowers peak temperature by more than 35% when PHY coverage is high. Its longer-term zHBM concept vertically integrates the XPU and memory stack, eliminating conventional 2.5D interfaces; Samsung estimates that four zHBM stacks alongside a 1,200 W GPU could save about 100 W while increasing bandwidth.

Read the source →
🖥️ Hardware & Infra ServeTheHome

Micron Evolving Memory Architectures for AI at Hot Chips 2026

Micron argues that AI accelerators are increasingly memory-bound because accelerator compute has been growing roughly 3× every two years while 2.5D-attached HBM bandwidth has risen at less than 2× over the same period. HBM raises bandwidth and capacity through denser 3D stacks and more I/O, but the cost is substantial: Micron says providing the same capacity with HBM3E consumes about three times the silicon of DDR5. The company emphasizes that memory can occupy more than eight times the surface area of a GPU die in an HBM4 package, putting advanced packaging, reliability, and thermal engineering at the center of accelerator design. Its proposed path pairs faster I/O, larger interposers, memory-optimized die-to-die links, and co-packaged optics with liquid cooling, hybrid bonding, and materials changes to manage heat and mechanical stress.

Read the source →
🏢 Industry & Business YT Nate B Jones

OpenAI Pays $280,000 For This Job. You Don't Have To Be An Engineer.

The video describes forward-deployed engineers as the people who make general-purpose AI useful inside specific companies, bridging a difficult “last mile” between a leadership goal and a working deployment. It points to Anthropic’s effort to train tens of thousands of engineers for installations in industries such as banking, airlines, and insurance, saying only 86 had actually been trained, and cites OpenAI compensation of up to $280,000 base plus equity and Handshake openings at $300,000. Using an insurance-claims example, it explains that an executive target such as doubling processing speed is not yet a buildable specification because the real work spans messy documents, processes, and adoption constraints. The proposed route into the role is to identify which of discovery, building, and deployment ownership a person already has, then build concrete evidence of the missing capabilities over roughly a month.

Read the source →
#ai#digest