BriefTechNews

Jev's 194x Speedup Breaks Vercel Adoption Records, Rewrites AI Gateway Economics

6 min read · 12 sources

TL;DR
  • Jev reached over twice as many paid teams on Vercel's AI Gateway within 24 hours than any prior model launch.
  • TypeSafe AI claims Jev is up to 194 times faster and 445 times cheaper than LLMs in workflow evaluations.
  • RatHat malware on Android captures on-screen credentials and 2FA codes, with infection chains as simple as a phishing email.
  • VMware has quietly scaled back its SmartNIC ambitions, signaling a shift in virtualization hardware strategy.
  • Jev is free on AI Gateway until September 25th, with routing, guardrails, and agent tool selection available.

A model nobody had heard of yesterday is the fastest-adopted thing Vercel’s AI Gateway has ever seen, and it’s not even an LLM. TypeSafe AI’s Jev, a probabilistic decision model, pulled in over twice as many paid teams in its first 24 hours than any prior launch on the platform. If you run agents or guardrails, this is the kind of speed-and-cost curve that makes you question why you’re burning tokens on a transformer for a routing decision.

Meanwhile, the backdoor that won’t die just laughed at your incident response runbook. GraphWorm doesn’t need a C2 domain because it lives inside a real OneDrive account, and revoking its OAuth token does nothing when it can rebuild its identity from a hardware hash. Today’s newsletter has the details on both, plus VMware quietly ditching SmartNICs, Fastly’s new AI firewall, and why your best engineer is the worst person to judge your AI pilot.

Jev is up to 194 times faster and 445 times cheaper than language models, and it just broke Vercel’s adoption record.

Jev Blows Past Every Prior Vercel AI Gateway Launch in 24 Hours

Source: vercel.com ↗

TypeSafe AI’s Jev model hit Vercel’s AI Gateway and became the fastest-adopted model in gateway history, reaching over twice as many paid teams as any prior launch within a day. Jev is a probabilistic decision model, not a language model, and TypeSafe reports it is up to 194 times faster and 445 times cheaper than LLMs in workflow evaluations. It is free on AI Gateway until September 25th. For engineers, this matters if you are using LLMs for routing, guardrails, or agent tool selection, where a purpose-built decision model can slash both latency and the per-call cost that makes agentic systems uneconomical at scale.

NVIDIA's DSX-Ready Program Tackles AI Factory Power and Cooling

Source: blogs.nvidia.com ↗

NVIDIA launched DSX-Ready AI Factories to help builders evaluate offerings with confidence and reduce integration risk. The program provides a system-level view of power, cooling, water, and grid constraints, which is the unglamorous bottleneck that kills more AI deployments than GPU availability ever will. If you are standing up a dense cluster, this gives you a reference framework for what the facility actually needs to survive a full rack of B300s running at peak.

GraphWorm: Revoking the Token Didn't Kill the Backdoor

Source: csoonline.com ↗

A reverse-engineered backdoor called GraphWorm uses a real OneDrive account via Microsoft Graph as a command-and-control dead drop, with no C2 domain or hardcoded address to block. It authenticates as an OAuth application, uses heartbeat and fingerprint files, and can remotely replace stolen credentials, all over TLS to graph.microsoft.com, making it invisible to network-layer defenses. Revoking the token doesn’t stop it because the implant builds an identifier by hashing the network adapter’s hardware address, enabling persistent re-authentication. Traditional incident response is dead on arrival here: you need to treat the identity plane as an object of compromise and hunt using cloud telemetry, not just rotate creds and call it a day.

Who Owns the Data? It's an Accountability Problem, Not a Technical One

Source: fromdata2ai.substack.com ↗

A deep dive on data ownership argues that ownership is fundamentally about accountability for business meaning, rules, and quality, not technical possession. It highlights common mistakes, such as assuming technical access equals ownership, and questions like “What does Customer Status = Active mean?” that reveal accountability gaps across application, database, and AI teams. If your AI team is training on a column that three different teams define three different ways, the model is the least of your problems.

VMware Quietly Walks Back Its SmartNIC Ambitions

Source: theregister.com ↗

The Register reports that VMware has quietly scaled back its SmartNIC ambitions, a potential change in virtualization hardware strategy. Details are sparse, but if you invested in a SmartNIC roadmap for your vSphere estate, this is the signal to reassess. The same article notes a reader hit with a surprise bill after Microsoft portals disagreed, Treasury chief Scott Bessent stating humans not, Google’s new $899+ thin-and-light laptops, and Russians posing as Signal support for phishing.

Fastly Launches AI Firewall and Runtime Control for Enterprise AI

Source: investors.fastly.com ↗

Fastly’s new AI Runtime Control routes model calls through a single endpoint across public and self-hosted providers, uses virtual keys to protect credentials, and offers real-time token spend visibility. This is aimed squarely at governing AI agent access to APIs and controlling costs as AI traffic explodes. If your agents are calling multiple providers, you no longer need to scatter API keys across every service.

Presidio Opens Enterprise AI Lab with Dell and Nvidia B300s

Source: presidio.com ↗

Presidio expanded its PATH lab to include Dell Technologies accelerated infrastructure, featuring Dell PowerEdge XE9780 systems with NVIDIA B300 GPUs for large-scale training and XE7745 servers with RTX PRO 6000 GPUs. Customers can benchmark performance and validate reference architectures before production deployment. If you want to test digital twins, RAG, or autonomous agents on real hardware without buying it first, this is the sandbox.

Atlassian's Synthetic Data Engine: Tens of Millions of Work Items, Zero Real Customers

Source: atlassian.com ↗

Atlassian built a scalable synthetic data engine to generate product-shaped, relationally correct datasets for Jira, Confluence, and ecosystem apps, without using real customer data. The system handles tens of millions of work items across varying numbers of spaces, addressing stress differences in system performance. If you have ever tried to load-test with sanitized production data and watched it fall apart on referential integrity, this is the architecture you wish you had.

Sumo Logic Adds More Control Over Observability Data Costs

Source: prnewswire.com ↗

Sumo Logic upgraded its Data Pipelines product so enterprises can filter, transform, and route telemetry before it is ingested. The goal is to reduce observability storage costs while preparing cleaner data for AI-driven operations. If your observability bill is climbing faster than your infrastructure, this is the lever to pull.

Your Best People Are the Worst Judges of Your AI Pilot

Source: hitesh.in ↗

A new essay argues that your best engineers are the worst judges of AI pilots because they unconsciously compensate for model limitations, a bias called subjective validation. It compares the effect to a psychic’s con, where the user supplies meaning to vague outputs, and notes that vendor-chosen demos flatter the tool. The core warning: pilots measure human intelligence, not model capability, so test with competent users who don’t have a stake in the outcome.

RatHat Malware Targets Android On-Screen Credentials

Source: cnet.com ↗

A new AI-powered malware called RatHat is invading the Android ecosystem, running in the background to capture on-screen information like usernames, passwords, and two-factor authentication codes. The infection chain is not necessarily more complex than a phishing email on Windows, according to Malwarebytes’ Sav Wheeler. On-prem AI “token factories” promise lower inference costs with Kaytus’s MotusAI upgrade, which integrates with vLLM and SGLang for dynamic scaling under peak traffic. If you need to cut inference costs while keeping agents on-prem, that is the thread to pull.

Get the brief

Liked this one? The rest of today's stack — AI, crypto, fintech, infra — lands in your inbox tomorrow morning. Five minutes, no hype.

About Me Author

My name is

BriefTechNews

A daily digest of what actually moved in AI, tech, crypto and fintech, assembled and written with AI, and reviewed before it publishes. Read More
Tags

You May Also Like