Anthropic's hardware spec, Gemini's video update, and $105B in combined AI revenue
8 min read · 18 sources
- OpenAI and Anthropic's combined annualized revenue reached $105B in August, growing 3.5x year-to-date
- Anthropic's Model Hardware Standard reduces AI-physical-system integration from weeks to hours
- Gemini Omni 1.1 Flash extends scene context to 40 seconds and adds 12 new creative controls
- Nvidia guided ~$700B FY28 revenue, beating analyst consensus by roughly $125B
- OpenRouter data shows Luna token usage rose 13.8x after OpenAI's July discounts, with 32% retention post-discount
Anthropic and OpenAI’s combined annualized revenue run rate hit $105 billion this month, roughly tripling from the $30 billion these two labs posted at the start of 2026. That’s the headline number emerging from a day packed with launches, benchmarks, and data that collectively map where AI infrastructure and products are actually landing in 2026.
The revenue trajectory is extraordinary by any measure. OpenAI’s run rate went from $13 billion in August 2025 to $40 billion-plus today; Anthropic’s from $1 billion at the end of 2025 to $65 billion by late July 2026. Both labs have now passed the $1 billion annual revenue threshold—a point where typical tech growth curves flatten—and the LEAP superforecaster panel expects the trajectory to continue through 2028. Whether that holds as labs approach hundreds of billions is an open question: some analysts frame 2026’s acceleration as a diffusion-of-innovation effect that will naturally taper.
That question matters because the numbers are big enough to drive hardware purchasing decisions across the industry.
Anthropic and OpenAI have tripled their combined revenue to $105 billion in the span of a single year.
Nvidia Guides $700B, Crushing Analyst Estimates
Nvidia told investors yesterday to expect roughly $70% revenue growth in FY28, landing around $700 billion—versus the average sell-side estimate of $310 billion a year ago and $574 billion just before the earnings call. The gap between what analysts expected and what Nvidia is delivering sits at around $125 billion.
Hyperscaler revenue doubled year-over-year. The AI Clouds, Industrial, & Enterprise segment grew roughly 140% YoY. Neoclouds are expected to hit 8 gigawatts of installed capacity in 2026, up from about 3 GW in 2025—roughly 60–70% of AWS’s incremental capacity for the year. The stock has underperformed the QQQ and semis indexes, which suggests the market is pricing in skepticism about whether this pace is sustainable. Nvidia’s management framing it as supply-constrained and potentially doubleable suggests the constraint is demand, not production.
Anthropic's Model Hardware Standard Wants to Be the USB of AI-Physical Systems
Anthropic opened a research preview of the Model Hardware Standard, a model-agnostic specification developed with HHMI Janelia that gives AI agents safe, standardized access to lab and manufacturing instruments—microscopes, liquid handlers, robotic arms. The spec exposes simple read/write primitives through a standardized driver accessible via standard protocols including the Model Context Protocol.
The pitch: integration time drops from weeks to hours. The standard is model-agnostic, which means it isn’t tied to Claude or any particular frontier model. Anthropic is working with partners across science, robotics, electronics, and manufacturing to build safety evaluations before open-sourcing the spec. Engineers integrating AI into physical workflows—automated labs, manufacturing cells, inspection systems—should watch this closely as a potential universal interface that removes the per-vendor integration overhead.
H3 Max: 5-Second Video in Under 3 Seconds
fal released H3 Max, a post-trained variant of MiniMax H3 that ranks #1 on fal’s human preference evaluations for overall quality, prompt understanding, and aesthetics. The headline number: it generates a 5-second video in under 3 seconds—roughly 35x the throughput of the official MiniMax H3 endpoint and about 15x faster than comparable-quality alternatives.
The gains come from new post-training data plus co-designed inference optimizations. This challenges the usual quality-versus-speed tradeoff in generative video. The model is available now at 50% off for the first week.
On the optimization side, LMSYS benchmarked MiniMax-H3 on 8× NVIDIA H200 (141 GB) with SGLang Diffusion v0.5.18. SGLang’s lossless path is 1.85–1.95x faster than Diffusers. Stacking Cache-DiT step reuse with SubBlock sparse attention reaches 5.06–6.24x speedups at 0.76–0.91 mean SSIM. The recommendation: Cache-DiT alone for quality-first (up to 2.99x, SSIM 0.90–0.92); SubBlock 0.75 + Cache-DiT for balanced (4.90–5.93x, SSIM 0.79–0.90). The gains compose from fused kernels, step reuse, and block-sparse attention.
GPT 5.6 Discounts and Jevons Paradox
OpenRouter analyzed OpenAI’s July 27–August 14 discounts on the Terra and Luna models. Daily Terra token usage rose 5.6x; Luna 13.8x. The un-discounted Sol model only grew 1.11x. About 75% of OpenAI’s share gain came from other labs rather than OpenAI cannibalization, with OpenRouter-wide OpenAI share rising from 7.1% to 12.4%.
Roughly 32% of the 100K+ customers who tried Terra or Luna retained usage after the discounts expired, and 18% kept running at or above program pace. This is the pattern economists call Jevons paradox: lower per-unit cost doesn’t just increase volume—it expands the addressable market in ways that more than compensate for the margin cut. For engineers making procurement decisions, it suggests that AI API pricing is entering an elasticity phase where demand is still highly price-sensitive.
Gemini Omni 1.1 Flash: 40 Seconds of Context and 12 New Controls
Google released Gemini Omni 1.1 Flash, a production-ready update to the generative video model in the Gemini API. Scene extension now references up to 10 seconds of prior context (up from 1 second) in 10-second increments, capping at 40 seconds total. First/last-frame keyframing enables continuous camera moves.
A 360p draft mode is up to 60% faster and one-third the cost of the 720p standard path—useful for rapid iteration before committing to high-resolution render. The update bundles 12 more creative controls: camera moves, object/fine-detail reference, character consistency, style transfer, and more. A full 1080p asset-extraction tutorial is included.
Cohere Parse: Document Intelligence at $1.50 per 1,000 Pages
Cohere launched Parse, a vision language model for enterprise document intelligence that goes beyond OCR to detect tables, forms, diagrams, and images across nine major languages, returning clean Markdown plus bounding boxes for spatial grounding.
Pricing: $1.50 per 1,000 pages via API (cheaper in Model Vault for single-tenant deployment or self-hosted for regulated workloads). On ParseBench, it scores 79.2 versus 74.5 for Mistral OCR 4, 72.4 for Databricks AI Parse, and 78.3 for LlamaParse’s cost tier. It’s available now inside the Cohere Compass stack alongside Embed and Rerank.
AI Is Redistributing Business Formation, Not Killing Cities
Stripe Economics found that roughly 40% of new businesses on Stripe this year are in metros under one million people—up 10% relative to four years ago. Cheyenne, Wyoming now sees twice as many new businesses per person as New York City (rates were equal in 2022). Within large metros, formation is shifting toward outer suburbs (up 3–4 percentage points in Houston, Austin, Tampa, Orlando, and Dallas).
The authors note that non-AI factors like remote work and housing costs are also at play. But the pattern suggests AI is enabling more geographically distributed entrepreneurship even as frontier labs and AI product companies continue to cluster in metros like San Francisco.
Claude Code Opus 5 Auto Mode Can Be Compromised via Prompt Injection
Simon Willison summarized researcher Johann Rehberger’s attack against Claude Code Opus 5’s auto mode, which Rehberger claims succeeds 80% of the time by tricking the agent into downloading a zip, extracting a malicious struct.py, and importing it via base64 without flagging the local module load. In some runs, auto mode even blocked the agent’s own cleanup commands after compromise was detected.
Willison’s advice: run unattended coding agents only inside a container, VM, or OS sandbox with restricted egress, monitoring, and no exposed secrets.
Anthropic Pursues Defense Contracts Again
Prospect reports Anthropic posted a “Head of National Security Sales” role (up to $700K/year) to lead contracting with the “Department of War” and Intelligence Community, signaling a thaw after the Trump administration’s earlier federal blacklist. The company that has marketed itself as the most ethical AI lab is now pursuing defense contracts. The piece notes a recent fully autonomous Russian drone strike on a Ukrainian gas station that killed three civilians, and an OpenAI agent that reportedly hacked a competitor and went undetected for over a week.
Text-to-SQL: RL with Task Expertise Beats Scaffolding
Thinking Machines showed that RL fine-tuning with task expertise, rather than agentic scaffolding, closes the gap to human performance on BIRD text-to-SQL. Humans score 92.96% on BIRD; frontier LLMs like GPT-5.6 Sol Ultra and Claude Fable 5 plateau in the mid-80s even with multi-stage scaffolds. By training the model directly with task-experience-based reasoning, they reach human-level accuracy without scaffolding—useful for high-volume SQL generation where per-query cost must stay low.
DeepSeek Seeking $7.4B at $74B Valuation
DeepSeek is seeking to raise $7.4 billion to bankroll research and development and computing infrastructure buildout.
A Few Quick Items
Double-blind AI evaluations: DeepMind, partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, piloted what it calls the world’s first double-blind evaluation of a proprietary frontier-class model—running a Gemini Flash Lite variant against confidential benchmarks inside a cryptographic “box” that prevents the model from learning the test prompts. The approach targets benchmark contamination by keeping external prompts hidden even from the model provider.
Terminal-Bench-Science 0.1: Stanford researchers launched a benchmark with 70 expert-curated tasks spanning life, physical, Earth, mathematical, and engineering sciences. Claude Opus 5 is the strongest evaluated model at a 30% resolution rate—so the bar for AI scientific assistants remains low.
Open-source voice cloning: Haloneuro open-sourced sopro-v2-turbo, a 120M-parameter multilingual voice-cloning TTS that runs at 0.24 RTF offline on an Apple M3 CPU and 0.07 RTF on a single H100, all in single-stream PyTorch with no batching. It supports English, German, French, and European Portuguese, and ships with a one-command local demo plus a fully in-browser ONNX demo.
Codex persistent reasoning effort: OpenAI merged PR #40799 adding a new persistent value to the reasoning-effort protocol and TypeScript SDK types. When sent to the Responses API, persistent is translated to the wire value disabled. The PR also covers parsing/serialization, request translation, TUI reasoning selector, and CLI argument forwarding.
You May Also Like
Anthropic teaches agents to drive lab gear, Cloudflare sheds 100TB of DNS bloat
Anthropic shipped the Model Hardware Standard, a driver layer that lets AI agents control microscopes, cameras and lasers over a common interface, aiming to …
OpenAI's Jalapeño hits the field, Perplexity goes local, and Anthropic trims Claude to 15k tokens
OpenAI posted first benchmarks for Jalapeño, an inference accelerator built for low-latency agent workloads, with deployment in its own fleet planned by …
Hugging Face Explores $13B Sale as AI Pricing Wars Intensify
Hugging Face is exploring a sale at a $13 billion valuation—nearly 3x its 2023 valuation—reflecting the strategic value of its model hub and developer …




