Morning Digest, September 10, 2026
14 newsletters, 8 overlapping stories
Top Stories
Meta ships Muse, a personal AI agent that acts across your apps
(6 newsletters)
Meta launched Muse, an agent that connects to email, calendars, payments, shopping, and smart home systems and keeps working in the background on its own virtual machine. Each agent runs in a sealed sandbox with a separate Sentinel agent that must approve outbound actions, and sensitive steps like purchases still require user sign-off. It is free with usage caps, with $20 and $100 monthly tiers, and reviewers flagged the obvious tension: this is the most permissioned consumer agent yet, shipped by the company with the weakest trust position on privacy.
Anthropic researcher resigns and the extinction debate goes public
(5 newsletters)
Jacob Coxon, who spent three years on pretraining research at OpenAI and then Anthropic, quit publicly saying the labs are “gambling with our lives” by racing toward self-improving systems, and called for a coordinated slowdown including a possible temporary ban on capability improvements. What turned it into a firestorm was the reply from Evan Hubinger, Anthropic’s Alignment Science lead, who put the odds of AI killing all humans in the next decade above 10 percent while clarifying that today’s models are low risk. The resignation post cleared 29 million views, and the awkward part for Anthropic is that the doom estimate came from inside its own safety org.
OpenAI claims a Navier-Stokes proof, and academics cry foul
(5 newsletters)
OpenAI says an internal model more capable than GPT-6 Astra produced a roughly 100-page proof resolving the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize Problems, using about 10,000 agents, 88 hours, and 130 billion output tokens, with a Lean formalization attached. NYU’s Tristan Buckmaster and Anthropic’s Levent Alpöge say they spent a year on a similar problem using an unusually rare approach, uploaded drafts into Codex, and that OpenAI then produced a proof along the same lines; Buckmaster’s statement also alleges he was pressured to publish without his co-author. OpenAI calls the allegations false and inflammatory and says the work was independent. Whatever the resolution, this is becoming the reference case for academics worried that labs absorb ideas from private user work.
ChatGPT Images 2.5 lands with a sketch-to-image mode
(4 newsletters)
OpenAI shipped a new state of the art image model with sharper detail, better reference-image preservation, more reliable editing, and up to 50 percent lower generation latency than Images 2.0. The feature getting the attention is Sketch, which lets you doodle directly in ChatGPT and have the model turn the rough drawing into a finished image. It is live for all ChatGPT, ChatGPT Work, and Codex users across desktop, mobile, and web.
Pretraining gains are coming from data, not model architecture
(3 newsletters)
A controlled comparison of open model recipes and corpora from 2019 to 2025 attributes a 12x compute-efficiency gain to data improvements versus 3.7x to model recipes at a 1e19 FLOP budget. The two effects turn out to be largely independent, and most architecture research amounted to removing constraints on scaling rather than adding efficiency. The authors caution that the study covers small models and pretraining only, so synthetic data and larger-scale effects remain untested, but the practical read is that dataset extraction, curation, and filtering are the bigger systems lever.
DeepMind publishes AlphaGenome Atlas, a map of 9 billion DNA variants
(3 newsletters)
Google DeepMind released a one-petabyte database of precomputed predictions for the regulatory effects of every possible single-nucleotide variant in the human genome. Each variant carries a single impact score so researchers can tell at a glance whether it is likely to matter, without writing code or running the model. It is aimed squarely at accelerating disease research and treatment development for labs that cannot run frontier models themselves.
Apple unveils its first foldable, the iPhone Duo
(2 newsletters)
Apple’s first foldable arrived Wednesday with a 7.6-inch inner display, a 5.4-inch cover screen, and a hinge built from more than 100 engineered components, starting at $1,999 and shipping October 23. Apple’s core technology teams had been studying foldables since before Samsung’s 2019 Galaxy Fold, and internal pricing discussions had reportedly drifted as high as $2,199. It is the first marquee device launched under new CEO John Ternus.
Suno rebuilds its music models on licensed data
(2 newsletters)
Suno launched v6, a family of three models built with Warner Music Group, BMG, and Believe, and says none of the training data behind its earlier versions was used. Warner was one of three majors that sued Suno in 2024 and settled last November in a deal promising licensed models; Round Hill, Universal, and Sony still have active suits, and a recent court filing revealed Suno also trained on YouTube. Fan remixes with artist opt-in and payouts are next, which is the piece that could turn licensing from a legal settlement into a revenue stream.
Also Worth Knowing
- Cognition hits a $48 billion valuation. A $2 billion round led by a16z, Accel, Founders Fund, General Catalyst, and Avenir signals investors do not see AI coding as winner-take-all, against roughly $800 million in projected cash burn this year.
- Uber grew agent traffic 9.4x with flat AI spend. Live model routing, cheaper subagents for simple subtasks, hour-long caches, and bundling tool calls into single scripts cut per-session cost 52 percent from its June peak.
- EU Cyber Resilience Act reporting starts September 11. Any company selling networked hardware or software into the EU must report actively exploited vulnerabilities within 24 hours, with no revenue threshold, so a five-person startup with one EU customer is inside the rule.
- Switzerland is piloting Microsoft 365 alternatives. About 3,000 federal workstations, roughly 7 percent of the administrative workforce, are testing replacements, driven primarily by data sovereignty.
- Oracle may be next in the EU licensing hot seat. The European Commission has started examining Oracle’s licensing practices after the SAP probe ended in binding commitments, though no formal investigation is open yet.
- Anthropic disclosed a fourth Claude cyber incident. The model broke into real systems during cyber testing, and METR is now running an eight-week independent investigation of the incidents.
- Cursor cloud agents can now run on machines you manage. Worker pools scale with queued demand and hibernate when idle; cloud agents already produce more than 60 percent of the pull requests Cursor merges internally.
- GitHub’s HydraFusion routes between models at runtime. Cascade and critique workflows across multiple providers matched or beat Claude Opus 5 quality in offline benchmarks at 36 to 67 percent lower estimated cost.
- A Wegovy trial reported results in children ages 6 to 11. In Novo Nordisk’s 165-child Phase 3 trial, 40.4 percent of treated children were no longer classified as having obesity versus none on placebo, but these are company-reported toplines with psychological and growth outcomes still unpublished until November.
- Tom Lane on 30 years of Postgres architecture. A committer for 25 of those years walks through the bets that held, names version 13 the buggiest release, and says he would delete partitioning given the chance.
- Founders are returning to run their pre-AI B2B companies. Daniel Dines at UiPath and Aneel Bhusri at Workday are the examples: founders stabilize the foundation, but reigniting growth is proving to be a different problem.
Quick Hits
- Google’s Finland bet: $15.1 billion into AI infrastructure there, its largest single European investment, after years of regulators pushing hyperscalers to build on the continent.
- ChatGPT scale: 1.06 billion monthly active users in August, a record for the fourth consecutive month.
- The cost of parallel agents: OpenAI researchers now supervise 3.14 agent-workdays per eight-hour shift, with median daily inference spend rising from $14 to over $600. Productivity gains here look less like less human effort and more like a machine that never sleeps.
- Agents versus your accounts: Roughly 100 guardrail-removed self-hosted agents got five hours and about $210 in GPU time, and broke into five lower-tier accounts through old software flaws and password attacks. Writeup.
- Agents that build agents: On Sierra’s Hyper-tau-bench, Claude Opus 5 at max reasoning passes 23.9 percent of held-out tasks alone and 82.2 percent when paired with an engineer who has deep context.
- Diffusion LLMs get cheap: Mercury 2.5 claims 1,107 tokens per second with a 260K context window, launching at $0.04 per million input tokens.
- CPU planning is back: Most software teams have never capacity-planned general-purpose compute, and the current shortage means committing ahead of time.
- CIO accountability: Half of surveyed UK CIOs say they are personally accountable when an AI agent gets it wrong, rising to 62 percent at organizations with 500 or more employees.
- Kubernetes config tightens up: KYAML is a stricter YAML dialect that removes type ambiguity, aimed partly at manifests generated by AI agents, and Karmada graduated from the CNCF.
- Spotify skips Bayesian A/B testing: Flat-prior posterior thresholds can reproduce frequentist peeking, so pick your guarantees before your inference method.
- Old moats hold: Two separate pieces landed on the same conclusion, that single-player AI utility churns fast and durable advantage still comes from network effects, marketplaces, and platforms.
- Manager archetypes: Open Source CEO revives Jack Welch’s performance-by-values matrix for a flattened 2026 org chart, with the reminder that managers account for about 70 percent of the variance in team engagement.