Nvidia licenses Poolside’s factory and hires its team
Summary
The day’s strongest thread is that AI capability is increasingly governed by operational economics: who can afford the compute, how reliably systems behave under real workloads, and whether a cheaper model actually produces accepted work. That same pressure is pushing organizations toward open-model routing and local optimization, while public resistance to the physical infrastructure behind frontier training is becoming a design constraint rather than a background concern.
Show HN: Huzzah – a novel approach to coding with AI
Huzzah is an experimental editor designed to replace long, impermanent chat prompts with persistent pseudocode files. A developer edits a .hz specification, and the tool captures the diff on save to regenerate only the affected source code through an LLM. The pitch is greater control, readability, and durability than iterative natural-language prompting, while retaining agent-assisted implementation. It is early-stage software, with source and setup instructions available for people willing to test the concept.
Read the source →Ladies and gentlemen I present to you Qwen3.8 27b 1bit brain damage quant
The post presents a 1-bit quantization of Qwen3.8-27B, framing it provocatively as a “brain damage” quant. No readable article body was extracted, so the raw material does not establish its method, quality, hardware requirements, or benchmarks. Treat the title as an announcement rather than evidence of practical performance.
Read the source →IsoExec: Unified Execution to Eliminate Trainer-Inference Mismatch in SkyRL
vLLM and SkyRL introduce IsoExec to address a subtle RL failure mode: rollout and training engines can assign different token probabilities to the same weights because their kernels, batching, and parallel reductions round floating-point operations differently. Its execution contract pins rounding-sensitive choices, while a unified model uses batch-invariant, bitwise-consistent kernels across training, prefill, and decoding. In an 8×H100 synchronous Qwen3.5-35B-A3B DAPO run, the authors report reducing rollout-versus-training log-probability divergence with a 25% overhead versus the SkyRL baseline. The work is motivated by prior reports in which small mismatch destabilized RL, caused heavy token clipping, and collapsed reward.
Read the source →not much happened today
This roundup’s central business signal is hybrid model routing: it cites AT&T routing 40% of employee AI usage to open models, targeting 60–70%, while reducing coding costs 56% for a reported 2% quality loss across 45 billion tokens per day. It also tracks pricing and access pressure around frontier models, including discounted GPT-5.6 Sol and complaints that intensive agent use can exhaust expensive subscription allowances quickly. On the local side, Qwen3.8-27B quantization and inference work points to a growing focus on making capable models faster and smaller, though commenters question benchmark comparability and hardware portability. The issue also groups new agent-product features, benchmarks, infrastructure work, and memory-oriented workflows as the other active fronts.
Read the source →Measuring benchmark optimization in speech recognition
Hugging Face researchers argue that high public ASR benchmark scores can reflect models learning benchmark-specific artifacts rather than becoming better transcribers. Across 11 open models, they found systems reproducing incorrect reference text even when audio contradicted it, filling in masked numbers, and switching spelling to match expected benchmark conventions. Their probes flagged possible reference errors in 40% of analyzed VoxPopuli clips, affecting roughly 3% of reference words; benchmark-optimized models reproduced erroneous references 18–30% of the time. The practical recommendation is to supplement public sets with held-out and newly collected audio, which weakened many of these effects.
Read the source →Grab my six-line handoff and cost scorecard, then find out whether a cheaper model actually saved you money.
The author argues that expensive coding-agent runs need cost controls because autonomous validation and repair can make a planned $20 overnight job exceed $300. They propose routing less critical work to GLM-5.3 through its $18-per-month coding plan while keeping Claude Code or Codex as the familiar interface, preserving files, permissions, hooks, MCP servers, tests, and review flow. The proposed handoff template compensates for the fact that a new model does not inherit the prior conversation, and the model should be assigned only work suitable for the cheaper queue. The key measurement is cost per accepted result: retries and cleanup can erase an apparent token-price saving.
Read the source →[AINews] Poolside gets $12B reverse-execuhire to NVIDIA; founders stay for $1B, employees go for $6B, Infraco scaling to 7GW neocloud
Nvidia has licensed Poolside’s “Model Factory” and hired 109 of its employees, which the article says represents the overwhelming majority of the technical team, in an unusual deal where founders remain rather than joining the buyer. Poolside says it missed a six-week window to raise $2 billion for a 40,000-GB300 cluster and concludes that frontier-scale training is now constrained not just by capital but by physical data-center capacity and contracted compute. The founders’ thesis is that human-level intelligence will become a low-margin open-source commodity, whereas AI systems paired with real-world experimental feedback will create the durable scientific-discovery moat. Poolside’s separate PIC infrastructure company is presented as pursuing a 7GW neocloud, while the company has not yet disclosed its revised vision.
Read the source →From Atari to EVE Online: Building on 15 Years of AI Research in Games
Google DeepMind says it is partnering with game studios to prototype new forms of AI gameplay. The extracted material does not provide technical details, named partners, release plans, or findings from those prototypes. The announced direction is therefore collaboration with game developers rather than a documented product launch.
Read the source →ChatGPT Apple Messages 💬, Anthropic’s meeting recorder 💼, Mistral Agentic Search 🔍
The extracted item is a brief pointer to DX’s analysis of quarterly changes in AI adoption, spending, and engineering output across more than 500 organizations. It promotes a discussion with DX’s Distinguished Scientist and Deputy CTO aimed at engineering leaders. The supplied body contains no supporting figures or details for the title’s ChatGPT, Anthropic, or Mistral product claims, so those claims cannot be evaluated from this item.
Read the source →9 AI Techniques You Probably Haven't Tried
The video surveys nine newer ways people are applying AI, framing the goal as practical techniques rather than a claim that users are broadly “doing AI wrong.” Examples named in the available transcript include Claude/design, Codex live voice mode, and a Grok feature that learns a workflow by watching a screen. It argues that constant product change makes it difficult to keep up, and that trust will come from concrete outcomes rather than glossy marketing; it cites Moderna and Merck’s successful Phase 3 personalized-cancer-vaccine trial as the kind of result that can shift opinion. The transcript available in the raw material is truncated, so this summary is limited to its supplied portion.
Read the source →The AI Backlash Is Getting Stupider But Also Smarter
The video argues that public opposition to AI data centers has become highly visible, from viral stunts to politicians changing course, but that the policy response is becoming more concrete. It contrasts a centrist governor’s strong executive order that makes data-center construction harder with a blanket moratorium, emphasizing that builders can meet specified criteria instead. It also points to OpenAI voluntarily pausing training as evidence that the backlash may produce more constructive limits rather than only obstruction. The transcript available in the raw material is truncated, so this summary is limited to its supplied portion.
Read the source →