BriefTechNews

US startups swap OpenAI for Chinese models; Salesforce and Anthropic go all-in

8 min read · 12 sources

TL;DR
  • Chinese models took all five top spots on OpenRouter by token volume in July 2026, with DeepSeek V4 Flash priced at $0.14 per million input tokens versus $5.00 for GPT-5.5
  • Salesforce expects ~$300M in Anthropic token spend in 2026, with Claude now the default across Atlas Reasoning Engine, Agentforce, and Slack, plus 37 Salesforce skills exposed inside Claude itself
  • 58 SaaS Capital Index constituents cluster at 10–30% growth, with valuation multiples compressing from 5.5x at 20–30% growth to 1.9x under 10%
  • A study of 642 YC companies from W21–P24 found priced-above-median startups reached Series A 25% of the time versus 8% for priced-below, and shut down at 6% versus 16%
  • Cohere Parse scores 79.2 on the ParseBench evaluation at $1.50 per 1,000 pages, ahead of LlamaParse Cost mode at 78.3 and Mistral OCR 4 at 74.5

US startups are quietly routing real workloads to Chinese open-weight models, and the gap is no longer a rounding error. In July 2026, Chinese models took all five top spots on OpenRouter by token volume, led by Xiaomi’s MiMo V2.5, with DeepSeek, Alibaba’s Qwen, and Moonshot’s Kimi filling the rest. Chinese models have carried over 60% of traffic, up from roughly 30% US dominance a year earlier. CNBC-cited data has Chinese models above 30% of weekly OpenRouter traffic since February 8, peaking at 46%, and DeepSeek topped Ramp’s June 2026 trending vendors list.

The driver is price. DeepSeek V4 Flash lists at $0.14 per million input tokens, against $5.00 for GPT-5.5 — a 35x gap on the input side that compounds across chat, embeddings, and long-context workloads. For a startup whose inference line is the single largest cost on the income statement, that is not a procurement decision, it is a pricing-defensibility decision.

The unresolved half is the boring one. Data residency, exfiltration risk, and US lawmaker pressure on Chinese-trained weights are real, and the piece is explicit that none of those questions are settled. Most teams adopting these models are doing so behind private deployments, not the public APIs, and the ones that aren’t are the ones that should expect an enterprise procurement fight later.

34 of 58 SaaS Capital Index names now grow between 10 and 30%, and Lemkin’s read is that most B2B CEOs are only ~40% of the way through genuinely rebuilding for AI.

Salesforce and Anthropic go all-in on Claudeforce

Source: thenextweb.com ↗

Salesforce and Anthropic formalised an expanded “Claudeforce” partnership that pulls Claude into the default reasoning slot across Salesforce’s agent stack: Atlas Reasoning Engine, Agentforce Vibes, Agentforce Coworker, Agent Builder, and Slack (powering Slackbot, Claude Tag, and Slack Code). Salesforce is expecting about $300M in Anthropic token spend this year, on top of an Anthropic equity stake now valued around $5B. Slackbot alone is credited with 8.1M annualised hours of internal productivity, doubling quarter over quarter.

The more interesting half is the reverse direction. “Salesforce in Claude” is a plugin that exposes 37 prebuilt sales skills — meeting prep, deal health, pipeline review — inside Anthropic’s own interface, with governed-write actions. For regulated workloads, Claude is reachable via Amazon Bedrock inside Salesforce’s Trust Boundary, which is the part that matters if you sell into a bank. The pattern is worth naming: the deal is not just model choice, it is a cross-vendor plugin architecture where the assistant you already live in becomes the surface for someone else’s enterprise software. That is how distribution is going to be settled for the next few years.

Most CEOs are maybe 40% of the way through rebuilding for AI

Source: saastr.com ↗

SaaStr’s Jason Lemkin is the newsletter’s top story today, and the case is built on public-comp data rather than vibes. Using the SaaS Capital Index as of June 30, 2026, 58 constituents cluster at 10–20% growth (23 companies) and 20–30% (11), with only 6 above 30% and the 60%+ cohort effectively gone. PitchBook’s Q2 2026 has median 2026 revenue growth at 13.2%. Multiples compress hard with growth: 5.5x at 20–30%, 3.1x at 10–20%, and 1.9x under 10%.

The argument is not that AI features are useless, it is that they are insufficient. Lemkin’s “40% of the way” framing means product architecture, pricing, and org design — the parts of the company that take years to redo. He holds up Fin’s four-year rebuild leading to a $3.6B exit as the reference shape, not the exception. For engineers, the operational read is that the bottleneck for most B2B software has moved up the stack from model access to product surface, and the teams that win the next cycle are the ones willing to re-platform rather than re-skin.

Lazyweb hit 50k agent signups with zero marketing spend

Source: read.first1000.co ↗

The most useful tactical writeup of the day is Lazyweb’s founder reflecting on reaching 50,000 agent signups with $0 marketing in a couple of months. The framing inverts the usual funnel: agents, not humans, are the primary user, and the tactics target agent friction specifically.

Three moves did most of the work. Auto-routing from agents.md to MCP doubled daily active users. Removing email from agent signup lifted activation by 50%. A one-line install command added another 20%. A 365k-view MCP launch post on X was the standout human-channel play, but the lesson is that machine-readable surfaces — agents.md, GitHub skills, .md blog versions — beat SEO or content marketing for products where an agent is the buyer. Token-efficient responses matter too: a 4k-token reply is not a UX choice, it is a cost-per-call line item.

AI agents should live for a day, not forever

Source: tomtunguz.com ↗

Tomasz Tunguz is arguing against long-lived agent sessions, and the data he cites is uncomfortable. When conversations are compressed to fit context windows, standing rules disappear in 30–59% of runs. Memory quality degrades as conversational turns grow across every frontier model tested. And a multi-year inbox/calendar write token is a hijack vector waiting on a single poisoned email.

His proposed architecture is a daily coordinator that resets at midnight, writes durable learnings to a preferences.md file, and dispatches stateless sub-agents (calendar_bot, email_bot, etc.) with roughly 30-minute lifespans and minimal tool scopes. The win is not just hygiene — it is that ephemeral agents cannot accumulate state an attacker can use, and tool-scope shrinking limits the blast radius when something does go wrong. For anyone building agent platforms, the alternative — long-lived sessions with persistent memory — is a set of compounding failure modes around attention, context rot, and persistent compromise.

Cohere Parse is a $1.50/1k-page document model aimed at LlamaParse and Mistral

Source: cohere.com ↗

Cohere launched Parse, a vision-language document intelligence model that turns complex multimodal files — tables, forms, diagrams, images — into structured Markdown across nine major world languages. On the company’s ParseBench evaluation it scored 79.2, against LlamaParse Cost mode (78.3), Databricks AI Parse (72.4), and Mistral OCR 4 (74.5). It also returns spatial bounding boxes for grounding, which matters for any RAG pipeline where you need to cite the source region rather than just the page.

Pricing is $1.50 per 1,000 pages via API, with cheaper private deployment through Model Vault or on-prem. It plugs into the Compass retrieval stack alongside Embed and Rerank, so the play is the full ingestion pipeline rather than a standalone OCR call. For teams currently routing through hyperscaler Document AI or stitching together open-source OCR with a layout model, the math is worth re-running.

Expensive YC startups reach Series A three times as often

Source: jaredheyman.medium.com ↗

A fund that buys into the top slice of YC batches just published the numbers, and if you raised at a cap in the last four years, somebody is running this against you. The study covered 642 companies from W21 through P24. Companies priced above their batch median reached Series A or beyond 25% of the time, against 8% for those priced below. They shut down at 6% versus 16%. The pattern holds inside each batch and among companies the fund passed on, so it is not a selection effect. Median caps rose from $15M in W21 to $40M in P26, which is the number to anchor on for anyone thinking about what their next round actually prices in.

AI is dispersing new business formation away from big cities

Source: stripeeconomics.com ↗

Stripe Economics is making a geographic argument: AI is filling the capability gaps that previously required co-located collaborators, and new business formation is following. About 40% of new Stripe businesses in 2026 are in metros under 1M people, up 10 percentage points versus four years ago. Cheyenne, Wyoming is forming twice as many new businesses per capita as NYC in 2026, against equal rates in 2022. The outer suburbs of Houston, Austin, Tampa, Orlando, and Dallas have gained 3–4 percentage points of new business share each.

The companies doing the actual AI research still cluster in the major metros, but the companies building on top of it increasingly do not. The implication for founders is that talent, customer, and infrastructure assumptions built around dense-urban clusters may underweight distributed and suburban opportunity, especially for sales-heavy and service-heavy businesses where the cost of living delta is the moat.

The rest: a billboard nobody could decode, and a small SaaS that is dying

Source: newsletter.notablecap.com ↗

Three shorter ones worth your time. Listen Labs needed to hire 100+ engineers as a Series A and posted a San Francisco billboard showing raw LLM token IDs — https:// {64659, 123310, 75584, 8138, 38271} — that decoded to a constrained-optimisation problem framed as a Berghain bouncer. Solving it earned an interview and a Berlin trip. The puzzle is essentially the job’s work in disguise: Listen Labs runs AI-moderated customer interviews at scale, so the filter is the function.

Swizec Teller argues focus and followthrough are the moat, using a 4-person team that consistently clears about 70% of a sprint (~30 points per person) while triaging ~50 new ideas per cycle. The core claim is that every sprint plan is overcommitted by default and shipping everything is physically impossible, so the discipline is choosing which stakeholders to disappoint. It is a useful counterweight to the “just add more AI productivity” line — the binding constraint is prioritisation, not output.

And the cautionary tale: Bank Statement Converter is down 24% on revenue and 12% on MRR since a February 2026 peak, with MRR growth turning negative in April (−$19,960 in July) and new subscribers collapsing from 191 in January to 45 in August 2026. Angus Cheng blames free-tier AI chatbots absorbing the common case, a better-funded competitor that appeared after he shared his stack publicly, and stopping “build in public” posting in September 2025. He is continuing to ship rather than shut it down. The post is a concrete case study of AI collapsing a small-tool SaaS’s pricing power and customer acquisition in under a year, which is the version of the AI-rebuild story that does not make the headline numbers.

Get the brief

Liked this one? The rest of today's stack — AI, crypto, fintech, infra — lands in your inbox tomorrow morning. Five minutes, no hype.

About Me Author

My name is

BriefTechNews

A daily digest of what actually moved in AI, tech, crypto and fintech, assembled and written with AI, and reviewed before it publishes. Read More

You May Also Like