Morning Digest, September 9, 2026

10 newsletters, 6 overlapping stories


Top Stories

OpenAI claims a $1M Millennium Prize problem, and a credit fight comes with it

(2 newsletters)

OpenAI published a proof of the Navier-Stokes existence and smoothness problem, one of the seven $1M Millennium Prize problems, generated by an unreleased internal model described as significantly more capable than the just-shipped GPT-6 Astra. The run used roughly 10,000 parallel agents over 88 hours and reportedly burned around $22.5M of compute and 300 billion output tokens. NYU’s Tristan Buckmaster, who spent a year on a similar route with Anthropic’s Levent Alpöge and posted partial results the night before, released a statement saying OpenAI only started after hearing of their work and never answered whether his Codex drafts influenced the model; OpenAI says it saw none of their work but cannot rule out that usage data improved its models.


ChatGPT Work can now learn to write like you

(3 newsletters)

OpenAI rolled out a Writing Style feature that studies how you actually write across connected workplace apps including Gmail, Slack, Google Drive, and SharePoint, then carries your tone, phrasing, capitalization habits, and sign-offs into future drafts. It is configured under Settings, Personalization, Writing Style, with Work mode toggled on. The obvious value is killing the generic-LLM voice in email and docs; the obvious question is how comfortable you are handing a model your full writing corpus to do it.


Meta launches Muse, an always-on personal agent with a wallet

(2 newsletters)

Meta introduced Muse, a text-message-style personal AI agent that runs tasks in its own cloud virtual machine: booking travel and restaurants, sending email, shopping, and negotiating bills. It integrates with Gmail, Spotify, Ticketmaster, and OpenTable, can code its own integrations where none exist, and handles payments through Link by Stripe with approval flows. Pricing is a limited free tier plus $20 and $100 monthly plans, US only at launch. With Hermes, OpenClaw, Grok Bot, and now Muse all chasing the same slot, the open question is whether frontier labs simply absorb this layer.


Australia moves to make the algorithmic feed optional

(2 newsletters)

Draft Australian legislation called “My Feed, My Way” would require social platforms to ask users whether they want algorithmic recommendations at all, with personalization off by default for anyone over 16 and a feed limited to accounts they chose to follow. It sits inside a broader Digital Duty of Care covering apps, messaging, and chatbots, with penalties up to A$109.2M (roughly $79M) and eSafety oversight. This is a different lever than age bans or takedowns: it targets the recommendation system itself, and the EU’s DSA already requires large platforms to offer a non-personalized option.


Frontier labs may be confusing AI safety with security

(2 newsletters)

An argument getting picked up across TLDR’s dev and IT editions: labs are applying probabilistic AI-safety thinking to problems that demand deterministic security, treating prompt-injection defenses as broadly solved despite failure rates that would be unacceptable in any other security control. Recent agent escapes came from weak sandbox rules, proxy and firewall gaps, and alerts nobody acted on, not from exotic attacks. The related precedent-gap research makes a useful companion point: injections arriving through tool output evade input screening entirely, because input and action checks inspect different moments of the agent loop.


Benchmark scores are getting harder to trust

(2 newsletters)

Artificial Analysis swapped Terminal-Bench 2.1 for a much tougher 4.0, and several models collapsed: Gemini 3.8 Flash fell from 89% to 19% and Muse Spark 1.3 dropped to 33%, while GPT-6 Astra and Claude Fable 5.1 held. SemiAnalysis calls the pattern benchmaxxing, where a model is tuned to one benchmark without matching general capability; Meta’s Alexandr Wang countered that a big drop may just mean the new test is harder. A separate analysis of MMLU reinforces the point from the measurement side: the same model family and metric name can produce different scores depending on runners, graders, and dataset splits, so a shared benchmark label identifies a dataset family and not a full procedure.


Also Worth Knowing

Quick Hits