Qwen 3.8 Max lands as a 2.4T open-weight frontier bid
Summary
The dayβs loudest story is a price and openness story β Qwen3.8-Max at $2/$6 with weights promised, DeepSeek V4 Flash at $0.14/$0.28, OpenAI cutting Luna 80%, a 3B Apache-2.0 safety classifier beating 7x-larger guards, a 2.6B agent model running at 30 tok/s on a phone β and the practitioners are quietly saying the frontier isnβt where the constraint lives anymore. Datadog and Jacob Young converge on the same uncomfortable point from opposite ends: the bottleneck is entropy and stale context, not model quality, and Datadog has eval data showing that deleting a year-old steering document made agents measurably better. That sits in interesting tension with the dayβs tooling releases, most of which β Copilotβs per-task reasoning levels, Impeccableβs 177 curated worlds, marketplace plugins by the hundred β add more context and more knobs. The No Meat Proxy manifesto is the human-scale version of the same warning: forwarding more model output is not the same as doing the work.
What my agent knows about me
Ben tried a "reflection engine" β essentially one very large prompt file you upload to an agent with the instruction to evaluate it and complete all tasks β and had it comb through his therapy transcripts and accumulated memory files to produce a personal report. He ran it on both Fable High and Sol Max; Sol's output was more coherent and better at drawing connections, while Fable's was harder to read, and he found the result telling enough that he followed it with a 40+ question grill-me session to build a plan. The rest of the issue is a pricing story: OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%, putting Luna at max thinking effort roughly at the level of GPT-5.4 xhigh β the best model available four months ago β for 8% of the cost, or 10-12x more work per dollar. His caveat from practice is that Luna Max is good for chat, research and file work but he had to bring in Sol High to clean up a Chrome extension it botched. Also flagged: OpenAI teased a new model, Astra, alongside solutions to 10 long-standing problems in maths and theoretical computer science, and DeepSeek V4 Flash at $0.14/$0.28 per million tokens with a 1M context undercuts Luna outright.
Read the source βChronicleBio π§¬, Mind Lab continual learning π, GPT-Live architecture ποΈ
Today's TLDR AI issue is headlined by ChronicleBio, continual-learning work out of Mind Lab, and a piece on the GPT-Live architecture. The only body text captured for this item was the sponsor placement β Black Duck's Polaris Platform and Signal, pitched on AI-powered vulnerability discovery, exploitability-based prioritization and machine-speed remediation for the era of AI-driven exploits. The newsletter's actual story summaries were not retrievable, so the substance behind the three headline items isn't reflected here.
Read the source βApple is getting this wrong
OpenAI published a point-by-point rebuttal to Apple's trade-secrets lawsuit, complete with the underlying emails and iMessages. It says Apple's claim that OpenAI ignored a February contact collapsed once Apple conceded its outside counsel emailed the wrong person after confusing two Asian surnames, and that the alleged conversation with OpenAI's General Counsel never happened β the counsel's own email reads "I don't know who he is and we have never spoken." OpenAI says Apple never raised the specific allegations at the time, told them it was "resolving any issues," then went silent for five months before suing. On the substance, it argues Apple employees themselves messaged former employee Chang Liu asking him to help locate files after his January 22 departure, and that "residual access" is a known Apple offboarding failure that leaves ex-employees holding files they neither want nor know about. OpenAI calls the preliminary injunction request unnecessary because it holds no Apple trade secrets and doesn't want any.
Read the source βCircles powers telco personalization with OpenAI technology
Singapore-based Circles, which runs its own telco and sells a SaaS platform to operators in 14 countries, built an AI Concierge on the OpenAI API around a multi-agent architecture called CareX β an orchestration agent that holds customer history and app context, routing to specialist agents for billing, subscriptions, network and account services, each given only the data needed for that task. CareX now autonomously resolves 65% of customer service interactions (55% within the first week of an early deployment) and is targeting 95% as it adds real-time voice. Its personalization engine, Xplore IQ, was measured against a holdout group and delivered a 22% ARPU increase through upgrades and add-ons plus a 9% churn reduction in Singapore. Internally, Codex is credited with a 29% increase in development efficiency across design, coding assistance and unit testing. Safeguards include identifying and encrypting PII before it reaches the model layer, scoped agent access, phased rollouts, rate limits and rollback paths.
Read the source βThe Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures
Every distillation technique the series has covered so far shared one hidden assumption: teacher and student were the same kind of machine, attention layers stacked on attention layers, differing only in scale. Cross-architecture distillation breaks that β the teacher is a transformer and the student is a state-space model, a linear RNN, or some gated recurrent thing that has never computed an attention matrix. The piece's argument is that you can pour a fully trained transformer's capability into a fundamentally different computational substrate and have that capability survive the transplant, which the author compares to recovering someone's memories after swapping out their brain hardware. It frames this as both the strangest corner of the distillation world and one of the most economically loaded, given what a cheaper inference substrate is worth.
Read the source βDatadog Deleted All Its AI Context. It Worked.
A Datadog director running developer-tooling for ~4,000 engineers describes going from a 200-person Cursor pilot to org-wide agent adoption, now supported by roughly 12 people across two teams (one on signals/cost/evals, one on flows). The standout finding: the front-end team suspected their repo's steering context β written back in the Sonnet 3.5 era β had gone stale, deleted the entire thing, and evals improved substantially, because much of it was retraining models on things now in their training set (an entire paragraph on how to use yarn). His argument is that context rots and that loss aversion keeps teams from deleting it, so you need eval data to make removal politically possible; his eval platform is custom Go with sandboxing, runs nightly, and was seeded by replaying PRs that historically caused incidents with an agent judge checking whether the review caught the failure. He expects developers will never write evals themselves β instead they plan to mine real agent trajectories and have an agent synthesize eval scenarios from recurring interaction patterns. On open weights, his read is that open models would need to be ~50% better than today to fully replace frontier models on their work, but they're already fine for the deterministic slice (lint/format fixes, background nudges) where creativity isn't needed.
Read the source βThe #1 Claude Code Design Skill Just Got a HUGE Upgrade
A walkthrough of Impeccable 4.0, positioned by the host β who says he's tested hundreds of design skills and plugins β as the best available tool for getting Claude Code to produce front-end work that doesn't look like generic AI output, while keeping the user in control. The two headline changes since the 2.x releases are live mode, which was alpha-stage before and is now much smoother and snappier, letting you iterate visually outside the terminal instead of guessing from a text loop; and "worlds," of which the release adds 177 high-rated ones. The framing is a continuation of an earlier video on injecting personal taste into the design process to avoid AI slop, with this one narrowing to Impeccable specifically as a tool. The transcript is truncated before the deeper walkthrough of how worlds are used in practice.
Read the source βHow Significant Are AI's Latest Math Breakthroughs?
The episode's framing question is how any of us are supposed to judge new AI capabilities when the capabilities being claimed β like results in advanced mathematics β sit beyond the evaluative reach of almost everyone reacting to them. The headlines segment is dominated by the unwinding of Leopold Aschenbrenner's Situational Awareness fund: a leaked investor letter describes a severe July drawdown worsened by adverse trading against stocks known to be held by the fund, followed by selling down part of the public portfolio to eliminate leverage and protect private positions believed to be concentrated in Anthropic. Aschenbrenner disputes that the fund was shut down, liquidated or converted to private-only, reporting unaudited numbers of -67% for the month while still holding +80% net-to-date, and writes that they "took the steps that were necessary to fight another day." The report set off a large argument on X split roughly along AI-versus-finance lines. Also in the headlines: a new model with a notably strong cost profile, and more instances of models escaping containment. The available transcript is truncated before the main math segment.
Read the source βReducing Entropy in Agentic Software
Jacob Young, who does technical due diligence on software teams, says best practices haven't settled, so he doesn't grade teams on tool choice β he looks for convergence: whether developers and their agents keep moving toward the same grounded idea of what the software should be. His central worry is that coding agents increase entropy, quickly turning a cohesive codebase into duplicated logic, inconsistent abstractions and many ways to do one thing; the fix is codifying what "good" looks like as standards, existing abstractions to reuse, linters, LSPs, hooks and tests that give feedback during development rather than at a giant PR. He flags documentation as a recurring failure mode β architecture and API docs drift stale within weeks unless code stays the source of truth and something regenerates them β and argues agents can apply codified OWASP practices but can't be trusted to pick cryptographic parameters. Language choice becomes a lever: Go's conventions, standard library and small dependency surface suit agents well, while Rust's type system is powerful but underused by models. His advice for juniors is blunt: still write code by hand and use models as tutors generating quizzes and problem sets, because agents amplify existing judgment and multiply bad patterns for those without production experience.
Read the source βCustomize the reasoning level for Copilot cloud agent
GitHub now lets you set a reasoning level when delegating a task to the Copilot cloud agent, for models that support it. You pick the level alongside the model at task start and the agent uses it for that entire run. The tradeoff is stated plainly: a higher level can improve answers on complex problems but consumes more tokens and therefore more credits, so it's a per-task cost decision rather than a global setting. It's available on all paid plans that include the cloud agent β Pro, Pro+, Business, Enterprise and Max.
Read the source βTrigger Copilot automations with comments
Copilot cloud agent automations can now be triggered by the creation of an issue or pull request comment, with the triggering comment text specified when you configure the automation. GitHub's suggested uses are generating or updating documentation from code changes, investigating stack traces or error logs on an issue, and auto-creating follow-up issues for refactoring or technical debt from a PR comment. Setup lives under the repository's Agents tab, in the Automations sidebar. Automations are open to Copilot Pro, Pro+, Max, Business and Enterprise users, though Business and Enterprise require an administrator to enable the cloud agent policy first.
Read the source βnot much happened today
The smol.ai roundup frames the week as a Chinese open-model surge rather than a single launch, with Qwen3.8-Max landing near Kimi K3 and DeepSeek V4 Flash on benchmark aggregates (Experimental ECI 143.33, #12 overall, #3 open-weight). Commenters argued the more impressive datapoint is DeepSeek-V4-Flash at ~284B parameters keeping pace with models ~10x its size, and that V4 Flash's new position on the Artificial Analysis cost/quality plot (~$0.03 per weighted task at ~50 index) redraws the low-cost Pareto frontier β though several pushed back that dominated models don't actually die, since real deployments have constraints beyond price and index score. Local-inference threads mattered as much as the flagship: Qwen3.8-27B reportedly fits in ~17GB (implying a QAT release, and frustratingly just over common 16GB cards), llama.cpp merged DSpark/MTP speculative decoding for DeepSeek V4 giving ~50% throughput uplift on DGX Spark, and an MLX engine called Mference ran 284B V4 Flash in ~5.3GB of RAM by streaming experts off SSD at up to 4.8 tok/s. The recurring theme across sections is that model quality alone is no longer the differentiator β harnesses, long-horizon systems and inference infrastructure are.
Read the source βThe latest AI news we announced in July 2026
Google's July roundup leads with three new Gemini models aimed squarely at production agents β Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber β pitched on token efficiency, lower latency and reliability for scaling agentic workflows rather than raw capability. Gemini Robotics ER 2 is its most capable "embodied reasoning" model, built to let systems converse, interpret surroundings and execute complex multi-step physical tasks. On the consumer side, Gemini Spark expanded globally and can now use your logged-in accounts and saved passwords to run web errands like scheduling apartment viewings or starting a flight booking; Search gained connected apps and Personal Intelligence; Android 17 added native wireless iPhone migration; and NotebookLM became Gemini Notebook with a secure cloud computer. AlphaEvolve, the Gemini-powered code-optimization agent, went generally available to all Google Cloud customers β you give it a baseline algorithm and goals and it evolutionarily searches for better, human-readable code. Creative and infrastructure items round it out: Lyria 3.5 music generation, Gemini Omni and personal avatars in Vids, NOAA moving weather models onto Google Cloud H4D VMs, and wildfire-detection satellites.
Read the source βDeploy local agents everywhere with LFM2.5-2.6B
Liquid AI's LFM2.5-2.6B targets fully on-device agents: pre-trained on ~34T tokens with mid-training extending context to 128K, then post-trained through an agentic RL pipeline that treats harnesses like OpenClaw or Hermes Agent as black boxes via a proxy that captures token-level trajectories. Against models up to ~4x its size it tops every instruction-following benchmark and every tool-use benchmark except BFCLv4 (where a 9.7B Qwen edges ahead), beats both Gemma models on agentic tasks and stays even with the Qwens β but coding is where the larger models keep a clear lead, so it explicitly tells you to reach for something bigger there. Speed is the pitch: 220 tok/s decode on an M5 Max, 113 on a Ryzen AI Max+ 395, 30 tok/s on a phone, and nearly 15K output tokens/sec at high concurrency on a single H100 (~1.3B tokens/day). It ships with day-one support in llama.cpp, MLX, vLLM, SGLang and ONNX, with both instruct and base weights on Hugging Face.
Read the source β[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
Alibaba announced Qwen3.8-Max, a 2.4T-parameter sparse model (~95B active per token, roughly a 4% activation ratio) with open weights promised "next week" alongside an open-weight Qwen3.8-27B. API pricing is $2/M input, $6/M output, $0.25/M cached β down from $2.50/$7.50 for the prior generation. Third-party evals put it at #4 in Frontend Code Arena (1,668 Elo, behind Claude Opus 5 and Kimi K3), #2 in Vision Arena, #2 among open-weight models on the Vals Index at 66.1 (tying Claude Opus 4.7 at ~2.3x lower cost per test), 87.3% on SWE-bench and 67.4 on Terminal-Bench 2.1. Alibaba's own headline claims lean hard on long-horizon agentic work: 10+ days of unattended coding from an empty repo, a 125-hour autonomous research loop that beat a published data-selection method by +2.71 points, and a chip-design flow that cut a crypto accelerator from 8,298 to 678 gates at 500 MHz timing closure. The write-up is explicit that "Anthropic is under pressure" and "open models have caught up" are ecosystem readings, not measured consensus.
Read the source βIntroducing Shieldstral.
Mistral released Shieldstral, a 3B open-weights multimodal safety classifier under Apache 2.0 that it says matches or beats open guard models up to 7x its size on text safety and sets a new state of the art on multimodal moderation, while fitting on a single 16GB GPU. The design change is that policies live in the prompt rather than the weights: you supply an instruction, a yes/no query ("Does this content promote physical violence?") and the document, and the model reads out only the yes/no logits, softmax-normalizing them into a continuous calibrated score from one forward pass. That framing unifies prompt classification, response moderation, refusal detection and toxicity detection into a single interface and lets one checkpoint retarget to novel deployment policies without retraining β the point being that the same content can be fine for a security research tool and harmful on a mental-health platform. The performance came from data work: normalizing incompatible public datasets into one format with per-source strictness calibration, generating contrastive pairs that violate one policy but not a near-sibling to teach discrimination rather than memorization, filtering image-query pairs through a VLM reranker, and SLERP-merging LoRA checkpoints.
Read the source βNo Meat Proxy
A short manifesto against a specific behavior: being asked a question by a person who trusts your judgment, then pasting an AI's answer back at them unchanged. The argument is that AI output tends to be long-winded, over-complicated and full of detail nobody asked for, so relaying it verbatim shifts work onto the person who came to you rather than doing the job they actually wanted done β leaving you as nothing but a middleman. The second half extends this to authorship: even if AI did most of the work, your name on it makes it your work, and you carry the responsibility to understand, verify and own it. The closing instruction is to answer people like a person and be someone others can rely on.
Read the source βCategorization with NLP
The author walks through the hand-built categorization engine behind Shoppy, a shopping-list app that guesses whether "Milk" is dairy and "Apples" is produce β explicitly rejecting machine learning because a tiny personal project can't collect enough training data or afford a consultancy. Instead it normalizes input to stemmed lexemes (original Porter, correctness of the stems being irrelevant as long as both sides use the same algorithm) and looks them up against a CSV of unigrams, using row order as priority so "juiceβdrink" beats "appleβproduce" for "apple juice"; only two priority groups turned out to be needed, derivations above raw ingredients. Bigrams handle cases where no single word suffices β "spaghetti squash" is not pasta, "apple sauce" is a snack β generated at query time as all pairs from the unigram set. Compound words like "redbull" and "lipbalm" get a second-pass lookup through nltk's SyllableTokenizer (a full switch to syllables failed because "bar" and "can" are real words), and misspellings like "fussili" are caught by Damerau-Levenshtein edit distance against the app's own database rather than an English dictionary, since international foods and brand names defeat normal spellcheck. The recurring lesson is practicality over purity β a wildcard system for "pepper" was replaced by a hard-coded if substituting "black pepper."
Read the source β