Morning Digest, August 25, 2026
14 newsletters, 9 overlapping stories
Top Stories
The memory shortage is now everybody’s price increase
(3 newsletters)
Nvidia has warned customers it will raise prices by more than 15% on servers shipping next year, including Vera Rubin and Grace Blackwell systems, blaming the cost of memory chips. Amazon moved first and harder, hiking Echo, Fire TV, Kindle, and Eero prices by as much as 60% over the weekend, with Apple already having raised prices earlier this year. One analysis argues Nvidia is passing the cost through rather than absorbing it because HBM is a relatively small share of an accelerator’s final price, which protects gross profit dollars now but sets up a harder squeeze in FY28. Expect “due to rising memory costs” to become the standard explanation across hardware budgets for the next few quarters.
Enterprises are building their own models on open weights
(3 newsletters)
Thomson Reuters introduced its first homegrown model, built by retooling Alibaba’s open-source Qwen and training on decades of Westlaw and Reuters content, at a total cost of $40M over two years with the latest training run at $450k. Legal AI startup Harvey did the same thing from the other direction, post-training its new Tenet model on Moonshot’s Kimi K3 despite OpenAI having been an investor since 2022. The pattern shows up in the usage data too: open source went from 28% to 62% of token share at Vercel in two months, and one widely shared analysis calls this summer the tipping point for open weights on price. For any company sitting on a deep proprietary content library, $40M reads less like a moonshot and more like a year of API invoices.
Anthropic opens its strongest model to defenders without opening the model
(3 newsletters)
Claude Mythos 5, previously restricted to the Project Glasswing partnership program, now powers code scanning in Claude Security for enterprise customers. Teams point it at a GitHub repo and get findings tagged by type, confidence, and severity, plus suggested patches, all billed as standard token usage on an existing plan. The design is the interesting part: users receive what the model finds but cannot prompt it directly, so nobody can ask it to write an exploit. Every fix still requires manual approval.
The grid is becoming the real constraint on AI buildout
(3 newsletters)
PJM is proposing rules that would make some new large-load customers, AI data centers included, interruptible when electricity supply gets tight, which would make power availability a hard ceiling rather than a line item. Public sentiment is not helping: one widely circulated analysis puts opposition to local data center development at 75%. The workarounds are getting exotic in both directions, with the AI boom pulling investment into solid-state power transformers that convert grid AC directly to the DC data centers need, while SpaceX and Nvidia announced they will build orbital Starmind data centers around a slimmed-down Vera Rubin NVL72 rack, targeting first racks in space by late 2027. Orbital compute still costs roughly four times ground compute today, but a site with no neighbors has an appeal that ground buildouts are losing.
Hugging Face is testing the market at $13B
(2 newsletters)
Hugging Face has reportedly worked with a bank to gauge acquisition interest at a valuation of $13 billion or more, nearly triple its 2023 mark. No buyer has been named and no agreement has been reached. The price reflects the model hub, the developer ecosystem, and the infrastructure underneath rather than any single revenue line.
Your website’s next visitor is an agent
(2 newsletters)
Vercel shipped Is Agentic, a free tool that scores how easily AI agents can discover, access, and use a website, running more than 100 checks per scan with one-click fixes and a CLI. It lands against a backdrop where agents already drive over half of all web traffic per earlier Cloudflare research, and new Mintlify data across 20,000+ documentation sites shows 66% of docs traffic is now agents rather than human developers. Agent experience is quietly becoming a real channel to optimize for.
Your local LLM is probably misconfigured, not dumb
(2 newsletters)
Someone tested quantized versions of the same open model across hardware configurations and logged 44.8 terabytes of results. Same model, same question, and the answer came out right on one graphics card, wrong split across two, right again across four. One common compression setting silently broke the model’s ability to call outside tools while a slightly larger one worked fine. If a self-hosted feature flakes in production but sails through the demo, the culprit is more likely a quantization or sharding setting than the code.
Also Worth Knowing
- Apple laid off 200+ across Vision Pro and Siri. Part of a restructuring that scales back Vision Pro gaming and immersive video while reorganizing Siri teams around the new Siri AI platform.
- Anthropic published an AI-native SDLC playbook (2 newsletters). The argument is that teams need to redesign planning, review, and operations around agents rather than bolting assistants onto existing process.
- Meta hired AI researcher Luke Metz (2 newsletters). Closes a two-year loop from OpenAI to Thinking Machines, back to OpenAI, and now to Alexandr Wang’s team.
- Slack launched a vibe coding tool. Slack Code lets teams tag Claude Code, Devin, or Copilot directly in group chats, requires human approval for production merges, and archives the channel when work is approved.
- Uber built a software factory for agentic coding. More than 70% of Uber pull requests are now written by AI agents and code shipped per engineer has doubled in a year, on a platform built around an MCP gateway.
- Uber was fined $966M by the Dutch DPA for automatically suspending drivers without warning or human oversight. Second-largest GDPR penalty ever issued, and Uber will appeal.
- Chinese state-tied hacking groups doubled their attack volume using DeepSeek. TeamT5 attributes the choice to low cost and loose guardrails, not raw capability, which makes cheap downloadable models the harder security problem.
- DeepSeek released an experimental Flash Vision model that nearly matches Opus 4.8 on agent benchmarks and works against both OpenAI and Anthropic API shapes.
- Opus 5 overtook Fable 5 in corporate model spending within a month of launch, though the analysis notes cheaper models can need more attempts and more human review, so cost per successful task is the number that matters.
- Ramp open-sourced its internal model router as a standalone product, routing on cost, quality, latency, and availability. Ramp says it cut its own LLM costs by roughly 30%.
- The FDA authorized Aletta, the first standalone robotic blood-draw device. It uses near-infrared light and Doppler ultrasound to find a vein and declines to proceed if it cannot, with one phlebotomist able to supervise up to three units in outpatient settings.
- Tesla confirmed a Cybercab launch event in Austin on September 3. Invite-only for top Robotaxi riders, against an active unsupervised fleet estimated at only 20 to 30 vehicles.
- The DOJ is nearly a year into an antitrust probe of a16z focused on partners holding board seats at now-competing AI companies. If firms start declining seats, term sheets and information rights change for anyone raising.
- nVent Electric is buying Maverick Power for $1.75B to add power distribution to its data center offerings, and Nvidia took a minority stake in Cloverleaf Infrastructure to develop gigawatt-scale AI factories in the US.
- Porsche signed a five-year, $1.46B deal with TCS to integrate AI across customer, factory, and engineering operations.
- Uncle Bob replaced his single coding agent with a five-agent assembly line: specifier, coder, cleaner, hardener, and QA, each resetting after its step. The full hour-long pipeline produces work he says would take a human half a day.
- The new MCP roadmap is out with five priority areas, and Google’s Agent2Agent protocol has joined the Agentic AI Foundation alongside the Linux Foundation-backed MCP ecosystem.
- Cloudflare now writes AI bot policies into robots.txt automatically across Search, Agent, and Training categories, free on every plan.
- Shopify’s CEO built an open-source Git server over a weekend. Walgit is a single-binary Rust server that turns cheap object storage into a code host with no database, serving large repos as static files.
- Stratechery argues autonomy is a startup advantage. Incumbents adopt AI cautiously because they have more to lose from mistakes, which gives startups room to push agent autonomy harder.
- A sloppy interface is now a security liability. AI makes faking a polished UI trivial, so high-fidelity craft details that are expensive to replicate double as a defense against imitation attacks.
- Wild AI-related reliability incidents are coming. The dangerous case is not an agent failing to fix an incident but an agent attempting remediation in an unexpected way, leaving humans to debug the system and the agent together.
Quick Hits
- Alibaba is raising roughly $10B in a share sale to fund AI infrastructure, chips, data centers, and model development, days after reporting a 75% jump in capex.
- Perplexity is in talks to raise from Nvidia at a valuation as high as $30B, with revenue reportedly more than tripling to $750M+ this year. The Information
- Spirit Airlines’ corporate records are the hot lot in its liquidation, not its planes. Mercor, Google, and Micro1 bid up to $12.5M for 100M emails, 500M Teams messages, and 7.5B transaction records to train models on how a real enterprise actually worked. Inc.
- More than a third of web pages published since ChatGPT’s release show signs of AI-generated writing, per a Pew Research analysis. Pew
- ChatGPT cut its Reddit citations by 86% in a single week, from 3.8% of all search citations down to 0.5%. Promptwatch
- Harvard Business School cloned its faculty into a $699 eight-week AI bootcamp with 24/7 access to the AI professors, already piloted at 100+ universities. Inc.
- McDonald’s loyalty program built a 515-page dossier on one Wired writer, including how often he orders and a prediction that he will never leave. Wired
- The 2026 B2B GTM benchmark: inbound is the most-adopted motion at 23%, channels run LinkedIn 66% and SEO 53%, 47% of teams are testing AI discovery channels, and AI credit pricing models are forecast to grow 114% in twelve months. Growth Unhinged
- Apple’s first new Mac mini in almost two years is expected within days, driven partly by demand from people running AI models locally, ahead of a September 9 event covering a foldable phone and the iPhone 18 Pro line.
- AI founders report they are working harder than ever to keep pace with their own agents, with erratic sleep and an inability to stop that they themselves call unsustainable.
- A wearable sensor matched arterial-line blood pressure readings in a 28-patient ICU study, and separately a machine learning screen on stem-cell-derived brain tissue flagged nine repurposed compounds for a childhood dementia. Both are early and neither is clinic-ready.
Shower Thoughts
Just like our browsers auto-detecting foreign languages, they should give us an option to convert imperial units to metric.