Morning Digest, October 8, 2026

15 newsletters, 12 overlapping stories


Top Stories

OpenAI’s 722 AI-generated math manuscripts keep rippling

(5 newsletters)

The math drop flagged yesterday dominated this cycle: 722 manuscripts across 372 research families, produced by an unreleased internal model that spent about three hours of compute per result, with many proofs formalized in Lean and more formalizations promised. OpenAI itself cautions that unverified findings may contain errors, and outside mathematicians have not yet checked the work. The backlash is growing too: 28 mathematicians have signed an open letter accusing labs of using unpublished research without credit, and critics question whether the model reasoned creatively or finished proofs built on borrowed human ideas.

OpenAI’s Decisions API kicks off a “decision model” gold rush

(4 newsletters)

OpenAI’s Decisions API is now in public beta for all developers. It takes text or images and returns typed predicates, choices, or rubric scores about 10x faster than GPT-6 Luna via the Responses API, at $0.10 per million input tokens with no output or caching charges. It is a direct shot at TypeSafe’s Jev, which reached 13% of Vercel’s paid AI Gateway workflows within a day of launch. Early independent tests favor the incumbent: HiringCafe found Jev beat both Decisions and Gemini 3.1 Flash-Lite at job matching, and MotherDuck published five SQL analytics use cases for Jev.

The personal agent race heats up: Grok Bot now routes to Claude

(5 newsletters)

Elon Musk said SpaceXAI’s Grok Bot will hand tasks to rival models like Claude and Midjourney when they are likely to do better, which makes xAI an Anthropic customer and puts Opus 5.5 behind a competitor to OpenAI’s Dots and Meta’s Muse. Meta and Sierra are also developing the open Personal Agent Protocol for how personal agents authenticate and act on company sites, and Hark’s agent is now handling groceries, rides, and bills. Dharmesh Shah argues the real value is connecting context scattered across email, LinkedIn, texts, and notes, while Andrew Chen asks whether agents will ever have network effects if any agent can email any other.

Claude moves into Google Docs, Sheets, and Slides

(3 newsletters)

Claude for Google Workspace is in public beta on every paid plan, adding a sidebar (Extensions, then Claude) that reads and edits files in place. Anthropic also shipped new Docs, Sheets, and Slides connectors so users can create and edit Google files from inside the Claude app by dropping in a link. Relatedly, Google Docs can now open and edit Markdown files natively without conversion.

Mistral Large 4 preview lands

(3 newsletters)

Mistral’s 1-trillion-parameter multimodal model (49B active parameters, 160+ languages) is pitched as the top open-weight model outside China, claiming wins over DeepSeek V4 Pro and Qwen3.8-Max at agentic coding and trailing only Claude Opus 5 in a blind code-quality test. The preview API is live now, with self-hosted weights due by month’s end. Meanwhile, a separate report says DeepSeek V4.1 Flash now leads agentic coding benchmarks and DeepSeek is close to a $12B raise backed by Tencent and CATL.

Google ships EmbeddingGemma 2 and Nano Banana 2.1

(3 newsletters)

EmbeddingGemma 2 is a 740M-parameter, Apache 2.0 model that maps text, code, images, audio, and video into one embedding space on-device, enabling offline multimodal search and RAG over internal knowledge. Nano Banana 2.1, built on Gemini 3.6 Flash with a 1M-token context, improves visual design, mask-based editing, and subject consistency and is rolling out across Gemini, AI Studio, Search, Ads, Flow, and Stitch.

Nobody reads the code anymore, and verification has to catch up

(3 newsletters)

Addy Osmani argues teams that stopped reviewing diffs never replaced that scrutiny, so defects just move downstream; at Anthropic an automated reviewer runs on nearly every PR, engineers dispute under 1% of its findings, and substantive review coverage rose from 16% to 54%. Spotify reports no direct link between AI-authored code and incidents, but merged changes more than doubled while verification struggled to keep pace. The Pragmatic Engineer’s state of the industry piece goes further: hand-written code is nearly gone, reviews have become theater, and quality is slipping.

ChatGPT gets “Intelligent UI” with GPT-6

(2 newsletters)

OpenAI began rolling out Intelligent UI alongside a new GPT-6 model, letting ChatGPT answer with tappable buttons, task-specific calculators, interactive charts, maps, and editable graphs instead of walls of text. The pitch is easier learning of complex topics and a friendlier chatbot for mainstream users.

Cloudflare wraps Birthday Week with 46 launches

(2 newsletters)

Highlights include a new cf CLI, an open-source pipeline called Forge, plans to become a public certificate authority, post-quantum crypto prep, and a beta Monetization Gateway that lets creators charge AI agents for content via HTTP 402.

The DNS root key changes October 11

(2 newsletters)

Only the second root key-signing key rollover ever happens this Saturday. Most organizations need do nothing, but anyone running DNSSEC-validating resolvers should confirm they trust KSK-2024 or risk resolution failures.

Inference is becoming software’s most important market

(2 newsletters)

Tom Tunguz projects AI inference spending will hit $350B by 2027, about double the database market. That pushes infrastructure into COGS and gross margins below typical SaaS levels, forcing efficiency through smaller models, better harnesses, and smarter routing. Separately, Ramp data shows businesses are switching model providers at a record 8% monthly rate, a possible sign of commoditization.

Airbnb replays real production traffic to test databases

(2 newsletters)

Airbnb captures MySQL traffic at ProxySQL, rebuilds transaction order offline, and replays it for load tests and migrations. It caught regressions like a query jumping from 0.03 to 2.6 seconds and carried the fleet from MySQL 5.7 to 8.0 without a major incident.


Also Worth Knowing

Quick Hits