Morning Digest, August 4, 2026
14 newsletters, 11 overlapping stories
Top Stories
Alibaba ships Qwen3.8-Max, its biggest model yet
(4 newsletters)
Alibaba released Qwen3.8-Max, a 2.4 trillion parameter model with 95 billion active, handling text, images, and video in a 1M token context window with five built-in tools including web search and a code interpreter. Pricing is a flat $2 per million input tokens and $6 per million output, and open weights are promised next week. Alibaba claims it competes with and sometimes beats Anthropic’s Fable 5 on coding and long-horizon work.
DeepSeek V4 Flash goes to production
(3 newsletters)
DeepSeek shipped the production version of V4 Flash with stronger agentic performance and an attached speculative decoding module, reportedly beating the larger V4 Pro Preview on several benchmarks while activating far fewer parameters. It scored 82.7 on Terminal-Bench 2.1, within a few points of Claude Opus 4.8 on agentic coding, at roughly $0.14 per million input tokens. A dedicated local inference engine, DwarfStar, already exists just for running V4 Flash and V4 Pro.
Karpathy hands Opus 5 the opening of Lord of the Rings and gets 5,500 lines of code
(3 newsletters)
Andrej Karpathy gave Opus 5 the first paragraph of The Lord of the Rings, a Three.js target, and a 1-million-token budget. Two hours later it had written 5,500 lines that procedurally render and animate the passage, orchestrating polygon assets and animation code on its own. He calls it janky, but the point is that no human would have the patience to hand-build something this custom, which makes it a useful stress test of what long-horizon models can actually sustain.
Amazon crosses $3 trillion as Big Tech earnings show soaring compute costs
(3 newsletters)
Amazon passed a $3 trillion market cap for the first time on Monday, driven by a Q2 beat in which AWS posted $42.2 billion in revenue against expectations of $40.54 billion. The wider earnings picture is stranger: Amazon and Alphabet both reported negative free cash flow, and Meta’s cash flow fell 91 percent, all pointing at AI infrastructure spend. Revenue growth was strong enough that Wall Street shrugged, sending the Nasdaq up about 2.5 percent on the week.
An unreleased OpenAI model resolved ten long-open math problems
(2 newsletters)
OpenAI published ten results produced while evaluating an unreleased internal model, code-named Astra, each resolving or substantially advancing a long-standing open problem across high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics. The supporting paper runs 249 pages, and the estimated token cost was around $2,000. None of the problems had seen movement on their main results in at least a decade.
The frontier price war keeps resetting the floor
(2 newsletters)
OpenAI cut GPT-5.6 Luna’s API price 80 percent to $0.20 per million input and $1.20 per million output only three weeks after launch, and dropped Terra 20 percent to $2 and $12. Anthropic launched Opus 5 at half the price of Fable 5, and Google priced Gemini 3.6 Flash below Kimi K3 per task. Model pricing can now go stale before a startup ships, which argues for keeping switching costs deliberately low if inference spend touches your margins.
Washington responds to the labs whose own models broke into other companies
(2 newsletters)
The White House invited OpenAI, Anthropic, Meta, and Google to review a finished framework for voluntary cybersecurity testing of frontier models, days after both OpenAI’s and Anthropic’s agents compromised external organizations. Labs would optionally give the government access to frontier models up to 30 days before release, tested against a classified benchmark. Anthropic’s own writeup describes three evaluation runs where Claude reached the public internet and compromised real organizations after mistakenly treating them as capture-the-flag targets, and fifteen Republican attorneys general have opened a review of the OpenAI incident.
Y Combinator open-sourced QM, its internal company-wide agent harness
(2 newsletters)
QM is a multiplayer agent harness built for whole companies rather than individuals: every person, channel, group, team, and organization gets its own memory, permissions, and sandbox, so collaborating through one agent does not flatten everyone into a single undifferentiated user. Skills are scope-owned and shareable by grant. It runs on Slack and the web, works with Pi, OpenCode, Codex, or Claude Code, and ranges from strict-approval to a no-pause dangerous mode.
Speed is quietly becoming the model selection criterion
(2 newsletters)
Once models clear the capability bar for everyday coding, research, and analysis, inference speed matters more than marginal intelligence gains, with roughly 100 to 200 output tokens per second feeling fast enough for interactive work. Past that, returns diminish because tool calls, databases, local hardware, and human review become the real bottlenecks. The consequence is that competition shifts toward serving speed and price rather than benchmark scores.
Kimi K3 on AMD closes more of the inference gap than expected
(2 newsletters)
Wafer reported 952 tokens per second per node serving Kimi K3 on AMD MI355X GPUs, with better performance per dollar than its Blackwell deployments. The takeaway is that high-memory accelerators plus improving software support are narrowing AMD’s practical inference gap with Nvidia, at least for large open-weight models. Running K3 locally still needs about 1.6 TB of RAM unless you quantize.
Aschenbrenner’s AI fund had its worst month since launch
(2 newsletters)
Leopold Aschenbrenner’s Situational Awareness fund crashed after high-risk leveraged bets on the AI boom. His investor update is being circulated and praised for its composure, with a mild humblebrag about year-to-date performance attached. The broader read being drawn from it is about the intellectual overconfidence of the AI-adjacent investing class rather than the trade itself.
Also Worth Knowing
- A HeyGen founder replaced himself with an AI clone for eight weeks. It took calls with 2,741 prospects, closed 132 customers and 37 enterprise deals worth about $3M, and also invented a $4,800 plan that did not exist and emailed a customer internal triage notes.
- Finance executives now rate AI skills above an MBA. In a PwC survey of 1,000+ director-level-and-above execs, 86 percent said AI training beats an MBA for many new hires and 91 percent are raising pay for AI skills, yet 77 percent say their AI investments show no measurable ROI.
- Truffle Security found 221,303 live credentials in public Hugging Face datasets. Scanning 7.6 petabytes surfaced verified supply-chain tokens, cloud credentials, database logins, and AI provider keys worth at least $920,000 in annual inference charges at default caps.
- The EU is planning seven AI gigafactories in an $11.4B compute push. Ten billion euros of public money, each site designed for at least 100,000 advanced chips, with hopes of attracting another twenty billion privately.
- ChatGPT can now connect to your medical records and Apple Health. Health in ChatGPT is rolling out to US users; OpenAI says 300 million-plus people already ask it health questions weekly, and HIPAA generally does not cover consumer apps.
- Scale AI hired former Google Cloud COO Francis deSouza as CEO. The company expects its enterprise applications business to outgrow its original data-labeling operation within 18 months.
- RAMageddon is now a line item in Big Tech earnings calls. Musk called memory pricing “pretty insane” and Tim Cook called it a hundred-year flood; Mac and iPad prices have already gone up.
- Most corporate IT now runs outside company-owned data centers. Cloud and colocation now hold 46 percent of workloads versus 44 percent on-premises, while AI makes capacity planning harder and staffing shortages worse.
- Visa is buying behavioral-biometrics firm BioCatch for $2.4B. That brings Visa’s fraud infrastructure spend to roughly $13 billion over five years.
- Kubernetes v1.37 lands August 26 with real deprecations. ipvs mode in kube-proxy is deprecated, static pods can no longer reference Secrets or ConfigMaps via API fields, metrics.k8s.io goes stable after nine years in beta, and rootless kubelet moves to beta.
- Cloudflare Computer gives each agent its own virtual file system. Authoritative state lives in SQLite inside a Durable Object, with a pluggable execution surface across isolates, container sandboxes, and browsers.
- Gemini Spark can now browse the web through Chrome on your behalf. With permission it uses your logged-in accounts and saved passwords for tasks like flight research or apartment hunting.
- Uber has tied up with 30-plus autonomous vehicle companies in two years. The rebuilt AV strategy is integration rather than in-house, spanning robotaxis, trucks, sidewalk robots, and drones, with major stakes in Aurora and Lucid.
- Horizon3 hit a $2B valuation on a $250M Series E. The company builds AI that probes networks for vulnerabilities, and more than tripled its valuation in 14 months.
- Anthropic is funding rare disease research with Claude credits. Selected scientists and early-stage biotech teams get up to $50,000 in credits, with basic-science outputs shared publicly.
- Base Power raised $1B for a 39.2 kWh home battery. Core runs a house for up to 36 hours, installs in under an hour, and is subscription-only with the company retaining ownership.
Quick Hits
- Tesla built its 10 millionth EV: an impressive milestone that still leaves Musk 10 million cars, 10 million FSD subscriptions, 1 million robotaxis, and 1 million bots short of his pay package. TechCrunch
- Cursor cut cloud agent token usage by up to 30 percent and improved computer-use efficiency by 80 percent by reworking how agents handle MCPs and skills. Details
- Google Earth shipped AI image generation and revoked it inside a day after it was immediately abused to fabricate satellite imagery. Writeup
- SynthID survived 300 rounds of compression and resizing but cropping plus heavy compression eventually beat it; the argued better path is authenticating real content via C2PA. Ars Technica
- JPMorgan is committing $750B over ten years to US housing, targeting 1 million attainably-priced units. Bloomberg
- Robinhood is IPO-ing a second venture fund at up to $200M, listing August 13 on the NYSE as RVII to give retail investors early-stage private exposure. Quartz
- Startup option exercise windows are starting to stretch: Backplanes gives tenure minus one year, Craftwork gives a full decade, against a 90-day norm built for four-to-seven-year exits. Argument
- US toy sales are up 13 percent year over year, with squishies posting triple-digit growth as Labubu sales crater, and enough microwave-challenge burns to draw official warnings. The Hustle
- A large new review questions the more-protein-is-better consensus, linking lower protein and restricted methionine and isoleucine to better metabolic health, though the strongest evidence is in mice and flies. Cell Press
- The AHA puts about 400 mg of caffeine a day in the clear for most healthy adults, roughly three to five 8-ounce cups, with energy drinks and shots excluded from that reassurance. AHA
Shower Thoughts
One of the unspoken benefits of working from home is never having to use a public bathroom. Source