DeepMind leaders leave to launch Discovery Loop
Summary
The day’s strongest theme is that AI advantage is shifting from a single model or a fixed engineering team toward systems that can direct, evaluate, and retain increasingly large amounts of machine work. Discovery Loop makes automated research the institutional bet; the agent-app and token-cost pieces make the same shift tangible for individual builders. That expansion also makes auditable methods, controllable context, and clear rules more important than raw access to models.
Show HN: Ex-Deloitte auditor open-sourced the whole SOC 2 method for your AI
Chiaro has published its SOC 2 readiness and examination methodology under CC BY 4.0: the controls, acceptable evidence, collection rules, and the rationale for judgment calls, while keeping its product code private. The release includes 528 synthetic calibration cases; 317 correct an AI that was too strict and 211 correct one that was too lenient, a deliberate attempt to make the system’s audit judgments challengeable without exposing client data. Its default is full-population testing rather than sampling, with independent reconciliation where possible and CPA confirmation before an AI-identified deviation becomes an exception. When sampling cannot be avoided, selection is seeded from a hash of the population, and changes to the method are to be published before use on an engagement.
Read the source →The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering
The piece argues that engineering capacity can no longer be modeled only as people, time, and talent: a single engineer can now run several agents in parallel on investigation, testing, prototyping, and documentation. It frames those agents as an elastic second workforce whose practical unit of capacity is token consumption rather than headcount. The warning is that organizations are beginning to optimize the readily dashboarded metric—tokens—in the same way they once optimized lines of code, which may be a poor proxy for useful engineering output.
Read the source →Google DeepMind reshuffle 🧠, Meta Muse Code 💻, Anthropic chip team 🧩
The available newsletter text is a sponsored introduction to Memoket Gem, a 0.4-ounce wrist-worn device that records conversations, extracts notes and tasks, syncs that context to Claude or other agents, and can place tasks on a calendar. It pitches the product as a way to make offline conversations available to AI without manual typing, with a $179 preorder and no subscription fee. The supplied article body does not include the newsletter’s reported items on DeepMind, Meta, or Anthropic.
Read the source →How To Make Claude Code Tokens 20x CHEAPER (& 4 More Usage Hacks)
The video argues that prompt caching is the most consequential lever for reducing Claude Code usage costs, more important than terse instructions such as “be brief.” It explains that input and output tokens accumulate in the context window across follow-up turns, so a short new request can still require the model to process the prior conversation; output tokens are described as roughly five times the price of input tokens. Its practical premise is that developers should understand caching and context behavior before trying to control spend with superficial prompt edits.
Read the source →I'm using a new agent app
Ben Tossell says he has abandoned t3 for “bb,” a cross-model desktop/mobile agent app that he found far easier to set up and more extensible. He values being able to switch among Claude, ChatGPT, Pi, Cursor, Factory, and other harnesses, and argues that agent workspaces should let users create plugins, task trackers, and other capabilities on demand. The broader claim is that a growing class of “builders,” not just conventional developers, will use agents to make small personal widgets and larger workflow tools; AI lowers the learning barrier that held back no-code products. The post also flags Airtable’s acquisition by Bending Spoons, DeepMind leadership changes and the Discovery Loop spinout, Meta’s Muse Code, and several agent infrastructure products.
Read the source →Why the Data Center Debate Has Little to Do with AI
The episode says the White House has discussed a new, tightly held voluntary framework for pre-release safety testing of frontier AI models with selected companies including OpenAI, Anthropic, and Google. Under the described proposal, eligible companies could submit models for up to 30 days of government testing before release, but neither the precise national-security threshold nor the identities and role of “trusted partners” with early access have been disclosed. It warns that the government and companies may not yet share a definition of which models qualify, leaving future release procedures unclear. The supplied transcript excerpt does not reach the episode’s main discussion of data centers.
Read the source →[AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM???
Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le are leaving DeepMind to found Discovery Loop, a public-benefit corporation focused on automating machine learning, science, and engineering research. Google is investing in the startup, while Radical Ventures and Khosla Ventures lead its seed round with Alphabet and other firms participating; the stated aim is automated discovery loops rather than another general-purpose model company. At DeepMind, Demis Hassabis becomes Chair of Google DeepMind and Alphabet Chief Scientist, stepping back from operations, while Koray Kavukcuoglu becomes SVP overseeing Gemini, frontier research, and product/development teams. The article treats the apparently amicable departures as a consequential governance and execution reset for Google, especially given other senior exits and the gap since the last Gemini Pro update.
Read the source →