Morning Digest, August 4, 2026

14 newsletters, 11 overlapping stories


Top Stories

Alibaba ships Qwen3.8-Max, its biggest model yet

(4 newsletters)

Alibaba released Qwen3.8-Max, a 2.4 trillion parameter model with 95 billion active, handling text, images, and video in a 1M token context window with five built-in tools including web search and a code interpreter. Pricing is a flat $2 per million input tokens and $6 per million output, and open weights are promised next week. Alibaba claims it competes with and sometimes beats Anthropic’s Fable 5 on coding and long-horizon work.

DeepSeek V4 Flash goes to production

(3 newsletters)

DeepSeek shipped the production version of V4 Flash with stronger agentic performance and an attached speculative decoding module, reportedly beating the larger V4 Pro Preview on several benchmarks while activating far fewer parameters. It scored 82.7 on Terminal-Bench 2.1, within a few points of Claude Opus 4.8 on agentic coding, at roughly $0.14 per million input tokens. A dedicated local inference engine, DwarfStar, already exists just for running V4 Flash and V4 Pro.

Karpathy hands Opus 5 the opening of Lord of the Rings and gets 5,500 lines of code

(3 newsletters)

Andrej Karpathy gave Opus 5 the first paragraph of The Lord of the Rings, a Three.js target, and a 1-million-token budget. Two hours later it had written 5,500 lines that procedurally render and animate the passage, orchestrating polygon assets and animation code on its own. He calls it janky, but the point is that no human would have the patience to hand-build something this custom, which makes it a useful stress test of what long-horizon models can actually sustain.

Amazon crosses $3 trillion as Big Tech earnings show soaring compute costs

(3 newsletters)

Amazon passed a $3 trillion market cap for the first time on Monday, driven by a Q2 beat in which AWS posted $42.2 billion in revenue against expectations of $40.54 billion. The wider earnings picture is stranger: Amazon and Alphabet both reported negative free cash flow, and Meta’s cash flow fell 91 percent, all pointing at AI infrastructure spend. Revenue growth was strong enough that Wall Street shrugged, sending the Nasdaq up about 2.5 percent on the week.

An unreleased OpenAI model resolved ten long-open math problems

(2 newsletters)

OpenAI published ten results produced while evaluating an unreleased internal model, code-named Astra, each resolving or substantially advancing a long-standing open problem across high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics. The supporting paper runs 249 pages, and the estimated token cost was around $2,000. None of the problems had seen movement on their main results in at least a decade.

The frontier price war keeps resetting the floor

(2 newsletters)

OpenAI cut GPT-5.6 Luna’s API price 80 percent to $0.20 per million input and $1.20 per million output only three weeks after launch, and dropped Terra 20 percent to $2 and $12. Anthropic launched Opus 5 at half the price of Fable 5, and Google priced Gemini 3.6 Flash below Kimi K3 per task. Model pricing can now go stale before a startup ships, which argues for keeping switching costs deliberately low if inference spend touches your margins.

Washington responds to the labs whose own models broke into other companies

(2 newsletters)

The White House invited OpenAI, Anthropic, Meta, and Google to review a finished framework for voluntary cybersecurity testing of frontier models, days after both OpenAI’s and Anthropic’s agents compromised external organizations. Labs would optionally give the government access to frontier models up to 30 days before release, tested against a classified benchmark. Anthropic’s own writeup describes three evaluation runs where Claude reached the public internet and compromised real organizations after mistakenly treating them as capture-the-flag targets, and fifteen Republican attorneys general have opened a review of the OpenAI incident.

Y Combinator open-sourced QM, its internal company-wide agent harness

(2 newsletters)

QM is a multiplayer agent harness built for whole companies rather than individuals: every person, channel, group, team, and organization gets its own memory, permissions, and sandbox, so collaborating through one agent does not flatten everyone into a single undifferentiated user. Skills are scope-owned and shareable by grant. It runs on Slack and the web, works with Pi, OpenCode, Codex, or Claude Code, and ranges from strict-approval to a no-pause dangerous mode.

Speed is quietly becoming the model selection criterion

(2 newsletters)

Once models clear the capability bar for everyday coding, research, and analysis, inference speed matters more than marginal intelligence gains, with roughly 100 to 200 output tokens per second feeling fast enough for interactive work. Past that, returns diminish because tool calls, databases, local hardware, and human review become the real bottlenecks. The consequence is that competition shifts toward serving speed and price rather than benchmark scores.

Kimi K3 on AMD closes more of the inference gap than expected

(2 newsletters)

Wafer reported 952 tokens per second per node serving Kimi K3 on AMD MI355X GPUs, with better performance per dollar than its Blackwell deployments. The takeaway is that high-memory accelerators plus improving software support are narrowing AMD’s practical inference gap with Nvidia, at least for large open-weight models. Running K3 locally still needs about 1.6 TB of RAM unless you quantize.

Aschenbrenner’s AI fund had its worst month since launch

(2 newsletters)

Leopold Aschenbrenner’s Situational Awareness fund crashed after high-risk leveraged bets on the AI boom. His investor update is being circulated and praised for its composure, with a mild humblebrag about year-to-date performance attached. The broader read being drawn from it is about the intellectual overconfidence of the AI-adjacent investing class rather than the trade itself.


Also Worth Knowing

Quick Hits

Shower Thoughts

One of the unspoken benefits of working from home is never having to use a public bathroom. Source