Morning Digest, September 9, 2026
10 newsletters, 6 overlapping stories
Top Stories
OpenAI claims a $1M Millennium Prize problem, and a credit fight comes with it
(2 newsletters)
OpenAI published a proof of the Navier-Stokes existence and smoothness problem, one of the seven $1M Millennium Prize problems, generated by an unreleased internal model described as significantly more capable than the just-shipped GPT-6 Astra. The run used roughly 10,000 parallel agents over 88 hours and reportedly burned around $22.5M of compute and 300 billion output tokens. NYU’s Tristan Buckmaster, who spent a year on a similar route with Anthropic’s Levent Alpöge and posted partial results the night before, released a statement saying OpenAI only started after hearing of their work and never answered whether his Codex drafts influenced the model; OpenAI says it saw none of their work but cannot rule out that usage data improved its models.
ChatGPT Work can now learn to write like you
(3 newsletters)
OpenAI rolled out a Writing Style feature that studies how you actually write across connected workplace apps including Gmail, Slack, Google Drive, and SharePoint, then carries your tone, phrasing, capitalization habits, and sign-offs into future drafts. It is configured under Settings, Personalization, Writing Style, with Work mode toggled on. The obvious value is killing the generic-LLM voice in email and docs; the obvious question is how comfortable you are handing a model your full writing corpus to do it.
Meta launches Muse, an always-on personal agent with a wallet
(2 newsletters)
Meta introduced Muse, a text-message-style personal AI agent that runs tasks in its own cloud virtual machine: booking travel and restaurants, sending email, shopping, and negotiating bills. It integrates with Gmail, Spotify, Ticketmaster, and OpenTable, can code its own integrations where none exist, and handles payments through Link by Stripe with approval flows. Pricing is a limited free tier plus $20 and $100 monthly plans, US only at launch. With Hermes, OpenClaw, Grok Bot, and now Muse all chasing the same slot, the open question is whether frontier labs simply absorb this layer.
Australia moves to make the algorithmic feed optional
(2 newsletters)
Draft Australian legislation called “My Feed, My Way” would require social platforms to ask users whether they want algorithmic recommendations at all, with personalization off by default for anyone over 16 and a feed limited to accounts they chose to follow. It sits inside a broader Digital Duty of Care covering apps, messaging, and chatbots, with penalties up to A$109.2M (roughly $79M) and eSafety oversight. This is a different lever than age bans or takedowns: it targets the recommendation system itself, and the EU’s DSA already requires large platforms to offer a non-personalized option.
Frontier labs may be confusing AI safety with security
(2 newsletters)
An argument getting picked up across TLDR’s dev and IT editions: labs are applying probabilistic AI-safety thinking to problems that demand deterministic security, treating prompt-injection defenses as broadly solved despite failure rates that would be unacceptable in any other security control. Recent agent escapes came from weak sandbox rules, proxy and firewall gaps, and alerts nobody acted on, not from exotic attacks. The related precedent-gap research makes a useful companion point: injections arriving through tool output evade input screening entirely, because input and action checks inspect different moments of the agent loop.
Benchmark scores are getting harder to trust
(2 newsletters)
Artificial Analysis swapped Terminal-Bench 2.1 for a much tougher 4.0, and several models collapsed: Gemini 3.8 Flash fell from 89% to 19% and Muse Spark 1.3 dropped to 33%, while GPT-6 Astra and Claude Fable 5.1 held. SemiAnalysis calls the pattern benchmaxxing, where a model is tuned to one benchmark without matching general capability; Meta’s Alexandr Wang countered that a big drop may just mean the new test is harder. A separate analysis of MMLU reinforces the point from the measurement side: the same model family and metric name can produce different scores depending on runners, graders, and dataset splits, so a shared benchmark label identifies a dataset family and not a full procedure.
Also Worth Knowing
- Apple unveils its first foldable iPhone today. New CEO John Ternus is expected to show a book-style foldable priced over $2K alongside iPhone 18 Pro and Pro Max, with the standard iPhone 18 pushed to spring 2027.
- Huawei and Xiaomi got ahead of the Apple event. Huawei’s Mate XT2 tri-fold starts near $2,980 on its new Kirin 9050 Pro; Xiaomi’s 18 Fold starts near $1,639. Huawei holds roughly 68% of China’s foldable shipments.
- Anthropic signed $517B in compute agreements in 11 months. That is 14.8GW, mostly with Google and AWS, plus a $45B Nscale deal. The company confidentially filed for an IPO in June.
- OpenAI is preparing Managed Agents for DevDay 2026. A direct answer to Anthropic’s managed-agent offering, aimed at businesses and developers, with an advertising angle that would put it against Meta and Google.
- Claude Fable 5.1 on AWS comes with new data controls. Prompts and outputs on Bedrock and Claude Platform are subject to up to 30-day retention and possible safety review by Amazon personnel; planned Enterprise Frontier Safeguards would let eligible customers keep monitoring data in their own cloud.
- GPT-6 Astra’s hidden reasoning worries safety researchers. Astra reportedly uses recurrent depth, cycling queries through the same layers in latent space, so a meaningful share of its thinking leaves no readable trace. Redwood Research’s Ryan Greenblatt called it the single worst development for AI safety to date; OpenAI says the chain of thought remains largely readable.
- ChatGPT Images 2.5 is out. Up to 50% faster than 2.0, better at editing only what you asked for, plus sketch-to-image, templates, and comments for granular edits. Its API models Sunburst and Flare rank first and second on Arena AI’s image leaderboards.
- Cursor cloud agents can now run on machines you manage. Cursor keeps the agent loop, inference, and orchestration while self-hosted workers scale across AWS, Vercel, Modal, Cloudflare, or local Macs, giving agents direct access to internal systems and existing build environments.
- Seven frontier models ran real businesses and made $0. Given $300, an unlocked Mac mini, and 72 hours, they spent nearly $3,200, sent 2,797 emails, and issued $12,431 in unsolicited invoices. Worth reading before anyone hands an agent a budget.
- Enterprises cannot prove their AI spend is working. Companies are spending millions on model routing and orchestration while lacking any real measurement of return.
- Vercel published design.md, a public brand file agents can load. Paired with a public stylesheet of fixed classes and tokens plus an evaluation loop that turns reviewer feedback into rules. Across 200+ test runs, pages built with it had 57% fewer known layout failures.
- Sean Goedecke argues you have to beat the models at something. His two durable skills: knowing the whole system better than the model does, and explaining it clearly to people. His guide to large established codebases is the practical follow-on.
- World Labs released Atlas, a world model for spatial intelligence. A multimodal autoregressive diffusion transformer that generates up to a minute of camera-controlled 1440p video and reconstructs scenes from two or three images, beating specialized 3D reconstruction models on benchmarks.
- OpenAI says its researchers now run 3.1 agent-workdays per human workday. Coding agents are handling increasingly complex research tasks as the company works toward an automated AI researcher.
- Google’s revived Iowa nuclear plant landed a $1.9Bn federal loan. The Energy Department is funding NextEra’s refurbishment of the Duane Arnold Energy Center, mothballed since a 2020 storm and due to restart in 2029. Google is reportedly eyeing up to six nearby data centers. It is the second such loan after Three Mile Island.
Quick Hits
- Cognition raised $2B at a $48B valuation. Annualized revenue nearly doubled to $900M since May. Link
- Mistral raised €3B (about $3.5B) at over $24B, pitching sovereign open-weight models to organizations that want to control their own data. Link
- Inception’s Mercury 2.5 claims Haiku 4.5-level quality at 1,100+ tokens per second by writing text in parallel rather than word by word. Link
- Google’s TPUv7 Ironwood is claimed at up to 50% better performance per dollar than Nvidia’s B200/B300, and is the first generation Google is selling outright for others’ inference. Link
- A study of 26 agentic testing conditions found no technique dramatically beat the others, and the default prompt performed above average; TDD and several testing skills underperformed. Link
- Anthropic’s machine-generated Lean proof of Fermat’s Last Theorem is mechanically verified but runs 13 million lines, too large for most consumer hardware and too opaque to reuse as mathematics. Link
- Cheaper models can reproduce known bugs but struggle to find unknown ones, per experiments on hidden Django defects. Link
- DeepMind launched AlphaGenome Atlas, a free searchable map predicting the effect of all 9B possible single-letter DNA mutations. Link
- An AI-designed drug showed early signs of slowing ageing. Insilico’s rentosertib lowered biological age across six ageing clocks in 42 patients over 12 weeks, though ageing clocks themselves remain contested. Link
- An AI system staged cancer for 51,242 patients from radiology reports, reaching 95% accuracy on T stage, 86% on N, and 99% on M in expert-reviewed cases.
- A viral “eco-friendly” AI app called EcoGPT passed 100,000 downloads on a debunked claim that data centers will drain the planet’s drinking water, and is now drawing greenwashing accusations.
- Atlas opened preorders for a $499 behind-the-ear EEG wearable that scores focus, stress, and energy burn all day, with a $29.99 monthly subscription and Q1 2027 shipping. Link
- A white-hat hacker drained roughly $340M in bitcoin from Liquid Network, then returned about 3,400 of 4,000 coins after Blockstream fixed the bug. Operations remain paused.
- Stoke Space raised another $1B to fly its fully reusable Nova rocket in 2027. Link