BriefTechNews

The 71% AI Cost Cut Hiding in Your Harness

9 min read · 15 sources

TL;DR
  • A Berkeley study shows the right harness cuts AI inference cost by 71% without losing accuracy, turning a $131k naive spend into $37k.
  • Pulley, a $50M+ funded Carta rival, is shutting down December 8 and directing customers to Carta itself.
  • OpenAI is testing Sponsored Agents that let users ask questions of an ad before clicking through, moving selling into the ad.
  • Column's four new products let developers build global card, stablecoin, and multicurrency products with one bank instead of many.
  • Zapier's AutomationBench scores GPT-6 Astra at ~40% task completion, the new state of the art for knowledge work.

The right harness can cut the cost of the same AI result by 71% without losing accuracy. That’s not a model upgrade or a cheaper API deal - it’s the plumbing around the model, and it’s now the biggest lever on your unit economics. A Berkeley study frames it as the next-generation software opportunity, and a hypothetical company using a harness spends $37k on inference versus $131k for a naive one, yielding a 75% gross margin versus 38%.

That changes what you build and what you sell. If the harness is the margin, then the moat isn’t the model - it’s your ability to compress common workflows into deterministic code and reserve expensive models for only the steps that need them. Engineers should care because the design of the harness, not just the model, is now the primary lever on unit economics for AI products.

The Harness Margin Opportunity

Source: tomtunguz.com ↗

The Berkeley study makes a concrete claim: the right harness can cut the cost of the same result by 71% without a loss of accuracy. The author, Tom Tunguz, argues that better harnesses compress common workflows into deterministic code, reserving expensive models for only the steps that need them. That creates both a cost advantage and a defensible moat - the harness encodes deep customer understanding, a collection of relevant evals, and a factory for automating hill-climbing.

In the hypothetical example, a company using a harness spends $37k on inference versus $131k for a naive one, yielding a 75% gross margin versus 38%. For engineers, the takeaway is that the harness is not an afterthought - it’s the thing you iterate on. The model is a commodity; the harness is the product.

The New Kingmakers

Source: x.com ↗

Revenue scale no longer matters much for fundraising. What matters is how many top AI-native companies use and love your product. A venture investor’s argument is that these “kingmaker” customers are the hardest to win because they can build internally and have high taste, so their loyalty is the strongest signal to investors. The post claims $50M+ rounds have been raised on the back of a single paying AI-native customer.

The advice for early-stage founders: target companies like those on Forbes’ AI 50 list instead of large enterprises like CVS or Walmart. Engineers should care because product quality and adoption by demanding AI-native users, not enterprise salesmanship, is now the primary fundraising currency. Revenue can be bought, but loyalty from a kingmaker can’t be manufactured.

Pulley Shuts Down, Ceding Customers to Carta

Source: techcrunch.com ↗

Cap table management platform Pulley is shutting down, with its final day of operations on December 8, and has partnered with rival Carta to offload customers. Founded by Yin Wu in 2020, Pulley raised over $50 million from General Catalyst, Stripe, and Founders Fund. A former employee speculated that Pulley was competing with spreadsheets, which startups can now easily maintain with AI.

A separate analysis suggests revenue growth stalled well below venture-scale expectations - from ~$5-10M ARR at Series B to ~$20M four years later - making it unattractive to acquirers. The piece implies Pulley made a deal with Carta to transition customers, with Carta honoring all contract terms. Engineers should note that even well-funded vertical SaaS tools are vulnerable when the underlying job can be done with generic, AI-assisted tools. Slow growth, not product failure, can force a well-funded startup to shut down.

Inside 12 Months of AI Search Experiments

Source: growthunhinged.com ↗

HubSpot’s former marketing leader details how the company’s blog traffic fell sharply after Google’s late-2024 algorithm change, though some of the decline was intentional as they deprioritized low-converting keywords. The experiments found that adding an llms.txt file did nothing, while 92% of highly specific industry and use-case pages eventually earned citations. A page-speed overhaul increased AI crawler visits by 1,600%, and qualified leads from AI search rose 1,850% over the year.

HubSpot shifted resources from educational content (commoditized by AI Overviews and ChatGPT) to “influence” content like YouTube, podcasts, and newsletters. The piece frames AI as a new channel to win rather than just a threat. Engineers should care because it shows how AI search changes the economics of content and the strategic value of building unique, non-commoditizable distribution. Track whether the content brings qualified leads, though, because more crawler visits alone won’t tell you whether it’s working.

Why Would Someone Pay for Your App When AI Does It for Free?

Source: revenuecat.com ↗

The RevenueCat article contrasts Chegg, whose stock fell 50% and which pivoted away from homework help due to ChatGPT, with Duolingo, which grew Q2 2026 revenue 18% and DAU 23% to 58.7 million despite free AI tutors. The difference is that Chegg offers a one-shot lookup experience, while Duolingo uses gamification and engagement mechanics to keep users on task. The author argues apps survive by focusing on areas where an LLM chat cannot beat them, such as habit formation and engagement.

Your customer can ask ChatGPT for a lesson. Getting them to practice again tomorrow is a different problem. Does your app remember their progress, organize what comes next, or give them someone to practice with? Those are reasons to keep using it even when the answers are available elsewhere. Engineers should care because it identifies the product design traits that determine whether an app gets commoditized by AI or thrives alongside it.

Mercury Books Brings AI-Powered Accounting to Banking

Source: mercury.com ↗

Mercury Books is a double-entry accounting product that pulls directly from Mercury banking data, with automatic categorization, real-time reconciliation, and P&L, cash flow, and balance sheet reports. It includes “Command,” an AI assistant that answers questions about finances and can offload work like writing journal entries and organizing the chart of accounts. It links external bank accounts, cards, and payroll platforms, and is free for Mercury customers.

The product was co-designed alongside accountants from leading firms, and existing Mercury customers can get started in minutes. Engineers should care because it represents the trend of banking platforms expanding into full accounting stacks with embedded AI. The accounting stack is becoming a feature of the bank account.

Column Collapses Global Payments into Four Products

Source: column.com ↗

Column released four new products: interoperable stablecoins (USDC/USDT) with instant 24/7 conversion to USD and payment rails; full-stack card issuing with its own issuer processor for Mastercard and Visa; global banking for issuing USD or local currency accounts to verified customers worldwide; and multicurrency accounts with local payouts via SEPA Instant and other systems. The company claims these are built on its own financial primitives, so developers need only one bank instead of stitching together multiple vendors.

Engineers should care because this collapses the number of integrations needed to build a global financial product, enabling complex cross-border flows via a few API calls. USDC and USDT are now completely interoperable with the US dollar and every underlying US and international payment rail on Column. That’s a big deal for anyone building payments infrastructure.

Reimagining Advertising with AI

Source: openai.com ↗

A customer can now ask an ad questions before visiting your website. OpenAI is testing Sponsored Agents with select US advertisers, letting users talk with a business-sponsored agent before clicking through to buy. It’s also adding ChatGPT Ads integrations to HubSpot and Shopify. That moves some of the selling into the ad itself.

If you get access to the test, your product information and the agent’s answers will need as much attention as the creative and landing page. The agent becomes the new landing page, and the quality of your product data determines whether the conversation ends in a sale or a dead end.

The Most Important Product Decision Is What You Don't Build

Source: liamnugent.me ↗

The author’s argument is that the most important product decision is what you don’t build, using “document hubs” and “notifications centres” in consumer fintech as examples of features that become bloated, multi-stakeholder projects. These features exist to solve internal/compliance problems (like legacy PDFs and “persistent communications channels”) but end up as bad clones of Google Drive or Gmail. The suggestion is to show stakeholders the running costs over years, not just the build cost, to kill such projects.

Engineers should care because it’s a playbook for resisting scope creep and avoiding long-term maintenance burdens from features that don’t serve users. Be extremely choosy about what you add to a system and be militant about taking things out.

The Curation Problem

Source: runthebusiness.substack.com ↗

The article warns that product teams using AI to ship faster are falling into a trap of focusing on output metrics, leading to UX clutter and feature bloat. The best predictor of product quality is the number of disciplined, targeted iterations (reps) that compound over time, not the raw volume of shipped features. With AI, more people can prototype and ship, but users’ ability to absorb changes is saturated.

Engineers should care because it cautions against the “ship all the things” mindset and argues for curation and disciplined iteration as the path to product-market fit. AI accelerates product development, but teams risk prioritizing speed over meaningful outcomes, leading to feature bloat.

Maximizing Return-on-Feedback

Source: avivbenyosef.com ↗

The Feedback-Squeezing Protocol starts with the premise that mishandling feedback (dismissing, attacking, or ignoring it) dries up future input. The protocol includes asking: Is it true? Did you realize it too late (and what held the person back)? Were you expecting to hear it from someone else? Addressing the gap between when an issue was first noticed and when you heard about it can reveal cultural misconfigurations.

When you let it be known that someone gave you feedback, and they see how you accepted and thanked the person who gave it, you normalize the behavior and make it safer for others to do the same. Engineers in leadership should care because it’s a practical checklist for turning complaints into actionable improvements and preserving a feedback-rich culture.

No Code Is Code: Zapier's Headless Bet

Source: cognitiverevolution.ai ↗

Zapier CEO Wade Foster discusses how AI work has consolidated around “daily driver” interfaces like Cursor, Claude Code, and ChatGPT, making headless tools like Zapier MCP (which bring context and data into those harnesses) the winning strategy over forcing users onto a proprietary agent platform. He cites Zapier’s AutomationBench, which scores frontier models on ~600 realistic knowledge-work tasks, noting GPT-6 Astra is the new state of the art at ~40% task completion, while Gemini 3.7 performs well at a fraction of the cost.

Foster emphasizes that deterministic code remains essential for reliable systems. Engineers should care because it signals a market shift toward model-agnostic, headless integration layers and provides benchmark data on model cost-versus-capability tradeoffs. No code is code - the integration layer is the product.

Get the brief

Liked this one? The rest of today's stack — AI, crypto, fintech, infra — lands in your inbox tomorrow morning. Five minutes, no hype.

About Me Author

My name is

BriefTechNews

A daily digest of what actually moved in AI, tech, crypto and fintech, assembled and written with AI, and reviewed before it publishes. Read More
Tags

You May Also Like