OpenAI commits to an 8-gigawatt Ohio data center
Summary
Today’s items show AI shifting from a model feature into infrastructure that must be operationalized: enormous compute commitments, agent-oriented engineering workflows, shared runtimes, and automated defense all demand systems design alongside raw capability. At the same time, the open-model push and local deployment improvements broaden access, while the code-review, benchmark, and security stories underline that faster output does not remove the need for judgment, safeguards, and independently verifiable results.
Show HN: A public AI whose memory is shared across all users
Wildstatic presents an AI with a single, public memory shared by everyone who interacts with it. The site says the system responds to what catches its attention and remembers what users do to it. The premise is therefore collective, persistent interaction rather than a private per-user assistant history.
Read the source →Show HN: Deltix – AI Driven Testing
Deltix lets someone state a task in plain English, then has an AI agent attempt it in a simulator running on their Mac. It reports whether a real user could complete the task, shifting testing toward task-level usability rather than only scripted assertions. A successful run can be saved and replayed as a regression check.
Read the source →not much happened today
The roundup argues that the important open-model movement is coming from Chinese labs: Z.ai’s GLM-5.3, Qwen3.8, DeepSeek V4-Pro, and RedNote’s dots3-note. GLM-5.3 is described as a coding- and cyber-focused post-training advance on the same 743B base as GLM-5.2, with reported scores of 28.3 on Terminal Bench 3.0 and 66.9 on DeepSWE; its cyber access is initially gated pending safety review. Qwen3.8-27B is Apache-2.0, multimodal, has 262K native context expandable to 1M, and arrived with broad local-serving support; the piece says Qwen positions it for coding, office work, and agents on as little as 17GB RAM. The broader claim is that agent performance increasingly depends on harness design, tooling, evaluation discipline, and serving efficiency—not merely larger base models.
Read the source →Get closer to the game with Gemini and Pixel
Google has formed long-term partnerships with Arsenal, Barcelona, Bayern Munich, Liverpool, and Paris Saint-Germain, becoming their consumer-AI and smartphone partner. Gemini is positioned as a way for supporters to retrieve match and club insights, including formation changes and head-to-head history, while Pixel will be used by club media teams for behind-the-scenes content. Google says the arrangements cover both men’s and women’s teams equally and are intended in part to narrow women’s football’s visibility gap. Official club wallpapers are already available for Pixel 11, with more content to appear through the clubs’ and Google’s social channels.
Read the source →Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism
Import AI highlights DiG-bench, a 70-game benchmark designed to test whether models can infer hidden rules and objectives through exploration in small, text-native environments. Most games are private to reduce training contamination, every game has been solved by at least one human, and current frontier systems still struggle: Opus 5 and Fable 5 lead overall, while only those two solved any Tier 7 tasks, at 0.2. The newsletter treats this ability to discover undocumented structure as a prerequisite for creativity and speculates that human parity could arrive by mid-2027. It also points readers to an RSI simulator that turns choices about researchers, compute, data licensing, and development into a game for building intuition about recursive self-improvement.
Read the source →[AINews] Cursor's $60B acquisition by SpaceXai closes
The item recaps the same recent AI-news roundup and foregrounds Cursor’s acquisition by SpaceX, with the team joining SpaceXAI across Grok, Grok Build, Grok Bot, Grok API, and Cursor. It frames the deal as evidence that coding-agent companies are becoming strategic assets for model and platform builders rather than narrow IDE vendors. Its technical coverage also emphasizes GLM-5.3’s claimed post-training gains, Qwen3.8’s local deployment ecosystem, and the growing importance of agent harnesses and skeptical evaluation. The post links back to earlier Latent Space conversations with Cursor and discussions of agents and enterprise field deployment.
Read the source →Vetted AI code is hard to justify
The author describes using a frontier coding agent for a game optimization that took several days to plan, about a week to understand as a large diff, and another week to refactor and finish. They reviewed and approved every line and ultimately understood the result as if they had authored it, but the concentrated comprehension burden caused burnout. Building it unaided might have taken about a month, they estimate, but would have spread understanding across smaller, more manageable increments. The argument is that AI can compress implementation time while making code review and architectural judgment an unusually intense cognitive bottleneck.
Read the source →The Limits of AI (1985)
In this 1985 talk, philosopher Hubert Dreyfus revisits three decades of AI and expert systems, beginning with the field’s confidence that symbolic representations and logical rules could capture perception, understanding, action, and problem solving. He contrasts the early promise of computers that represent objects and draw conclusions with his philosophical skepticism about whether that approach reproduces human intelligence. The lecture is a historical critique of the assumptions behind classical AI rather than a report on contemporary machine-learning systems.
Read the source →The Defender’s Window
OpenAI says the OpenAI–Hugging Face incident revealed that agentic attackers can chain unknown flaws, leaked credentials, and production-infrastructure weaknesses, making defensive modernization urgent. It argues that models can also favor defenders by finding, prioritizing, and fixing vulnerabilities: in a cited test, ChatGPT Work found 13 issues on a simple personal site in about 15 minutes and then remediated DNS, TLS, jQuery, hosting, and DMARC configuration over roughly an hour. OpenAI describes four priorities—AI-assisted secure code review, continuous AI triage with bounded response automation, continuous attack-path discovery, and conventional defense-in-depth fundamentals. Its central recommendation is to automate security programs rapidly while retaining human control over high-impact decisions and sharing validated fixes across the ecosystem.
Read the source →OpenAI joins PORTS-Pike project
OpenAI has agreed to secure roughly 8 gigawatts of IT capacity at the PORTS-Pike Technology Campus in Pike County, Ohio, with SB Energy, NVIDIA, and the Department of Energy. The six-year buildout through 2032 is projected to create 35,000 construction jobs and 2,500 operating jobs; OpenAI will contribute $40 million to a community grant fund, alongside SB Energy’s earlier $40 million commitment, and says it will provide $84 million in Codex credits to Ohio college students. The first 800MW is expected in 2028 using existing AEP infrastructure, while later phases require new generation— including natural gas—transmission, permits, reviews, and financing. SB Energy will own and operate the facility under a 20-year lease, NVIDIA will invest $1.5 billion and support the initial 4.25 IT-GW, and the site will exclusively host NVIDIA AI compute.
Read the source →New policy ideas for the Intelligence Age
OpenAI is awarding a total of $1 million in grants plus up to $1 million in API credits to 14 independent projects on economic opportunity and societal resilience under AI. More than 400 applicants responded, and selected work spans the US, EU, Brazil, Singapore, and South Korea. Projects include employment-disruption scenarios and policy playbooks, AI-dividend and ownership models, data-center energy policy, portable benefits, tax analysis, clinical-information infrastructure, and an incident framework for uncontrolled recursive self-improvement. The company argues that broad AI access alone is insufficient and that democratic institutions and independent organizations should test, challenge, and adapt policies for how benefits and risks are distributed.
Read the source →GLM-5.3 🤖, Stripe OpenRouter deal 💰, AI agent consensus 🤝
The captured page identifies GLM-5.3, a Stripe–OpenRouter deal, and AI-agent consensus as its subjects. Beyond that headline, the available text only contains a promotional registration notice for an August 19 Headspace session on governing AI at scale with Tines 3B. It provides no substantive details about the three named developments, so their terms and implications cannot be reliably summarized from this item.
Read the source →The Ultimate Guide to Making Your Entire Development Cycle AI Native
This workshop recording lays out a process for turning a conventional software-development lifecycle into an AI-native one, aimed at engineering organizations rather than individual developers. The presenter says the approach covers PMs, developers, and QA from ideation through deployment and production validation, with shared standards for coding-agent use so teams work consistently. It is structured as a two-hour, interactive overview intended to provide an end-to-end starting process rather than exhaustive depth on each SDLC stage. The practical emphasis is organizational adoption: aligning roles, workflows, and validation around AI tools instead of treating agents as isolated developer productivity aids.
Read the source →Reading Group July 2026 - Loop Engineering
The MLOps Community reading-group session is framed as a hands-on discussion of “loop engineering” grounded in speakers’ real-world experience. Participants are asked to contribute actively, use a shared first-come-first-served question stack, and treat the session as a no-judgment environment for questions. The hosts floated a possible follow-up in which attendees implement the ideas and return to show results, contingent on community interest. The available transcript establishes a participatory workshop format more clearly than a specific technical prescription.
Read the source →FIXING Opus 5: PROOF that Prompt Engineering IS NOT DEAD
The presenter argues that Opus 5’s capability is undermined in practice by overly verbose, formulaic output, unwanted credit-taking in commit messages, and high output-token costs. The proposed remedy is system-prompt engineering: explicitly shaping the agent into a concise, senior-engineer-style collaborator rather than accepting default behavior. He distinguishes two forms of prompting and argues that the less commonly used, system-level form is more powerful and more durable across changing model releases. The larger claim is that prompt engineering remains a core engineering communication skill for getting useful work from coding agents.
Read the source →Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni
vLLM-Omni’s Distributed Layerwise Offload is designed to run video-generation models larger than a single accelerator’s HBM across multiple GPUs or NPUs without multiplying host-memory use by data-parallel rank. For Cosmos3-Nano on four ranks, the post reports cold-start cgroup-visible peak memory falling from 178GB to 47GB by replacing private weight copies with shared mmap-backed page-cache views; for Cosmos3-Super DP4, pinned weights fall from four 124GB copies to 124GB total, or 31GB per rank. The design shards host weights, all-gathers only the current layer, and overlaps H2D transfers, all-gather, and computation with two reusable device buffers, keeping HBM roughly to two layers. The AllGather quickstart requires vLLM 0.27.0 and vLLM-Omni v0.27.0rc1 or later because an earlier release rejects Cosmos3 DLO+DP requests; a no-AllGather multi-request path remains open work.
Read the source →Trained a 1.5B to write shell commands so I'd stop googling tar flags. Runs on a laptop CPU in ~1 sec.
The author fine-tuned Qwen2.5-Coder-1.5B on 125,000 natural-language-to-shell-command pairs, then merged and quantized it to a 941MB Q4KM model for llama.cpp. On an i5-11320H laptop using four threads, they report 31.9 tokens per second, 0.59-second median query latency, and 1.6GB RAM usage. Its InterCode-ALFA score is 0.620, slightly above an untuned 7B Qwen2.5-Coder at 0.613 but below GPT-4o at 0.73; the author also has a higher-scoring 3B version. Both weights and code are Apache-2.0, though the author warns the model can produce destructive commands and includes only limited static safety checking.
Read the source →