BriefTechNews

OpenAI's 700W Jalapeño, GPT-5.6 lands in GovCloud, and SaaS prepares for its AI reckoning

6 min read · 12 sources

TL;DR
  • OpenAI's Jalapeño AI chip is rated at 700W TDP but stayed at or below 550W across tested inference workloads
  • GPT-5.6 Terra and Luna are now generally available on Amazon Bedrock in AWS GovCloud with million-token context and 90% prompt caching
  • Cisco is adding Supermicro liquid- and air-cooled servers to its Secure AI Factory with NVIDIA for dense GPU deployments
  • Perplexity's new Portable Computer runs agent workloads locally on Nvidia DGX Spark and RTX machines with zero token costs
  • Oracle Exadata Database@AWS is now generally available, with licensing experts warning of higher compute and storage costs versus native AWS

OpenAI just showed its hand on custom silicon. The company’s first AI accelerator, codenamed Jalapeño and built with Broadcom, is rated at 700W TDP but held at or below 550W across the workloads OpenAI publicly tested, according to fresh details reported by Data Center Dynamics. That gap between the rated ceiling and the real-world envelope matters: inference economics live or die on the watts-per-token line, and a 21% headroom under load is exactly the kind of slack operators want before they bet capacity planning on it.

Below the silicon story, a quiet but important distribution move. Below the silicon story, a quiet but important distribution move. GPT-5.6 Terra and Luna are now generally available in AWS GovCloud on Bedrock (US-West and US-East). Terra is the balanced tier, Luna is the cheaper, faster one, and both ship with million-token context windows and a stated 90% discount on cached prompts. For federal-adjacent shops that have been blocked from Bedrock by accreditation, GovCloud is the only door in.

OpenAI’s 700W Jalapeño accelerator stayed at or below 550W in tested inference workloads, hinting at the power curve OpenAI will need to make economics work

Cisco pulls Supermicro into the NVIDIA AI Factory

Source: networkworld.com ↗

Cisco is bolting Supermicro onto its Secure AI Factory with NVIDIA, adding both liquid- and air-cooled rack-scale systems to a stack that previously leaned on Cisco’s UCS and HyperFlex lines. The pitch is one SKU for enterprises, neoclouds, and sovereign clouds that need to stand up dense GPU islands without running a multi-vendor integration project.

For an SRE evaluating this, the interesting question is whether the networking, security, and Nexus fabric side of the Cisco stack still pulls its weight once the compute line is someone else’s. Liquid cooling in particular is where Cisco’s rack-level experience has to coexist with Supermicro’s traditional server strengths, and integration risk tends to live in those seams.

McKinsey: enterprise AI is finally getting a line on ROI

Source: theregister.com ↗

McKinsey’s latest numbers, summarised by The Register, say the spend curve keeps climbing but the revenue line is starting to bend. The pattern McKinsey is now reporting is familiar to anyone who has shipped past the first wave of copilot pilots: returns show up when the AI is embedded in a redesigned workflow, not when it is stapled onto an existing SaaS screen. Companies measuring redesign-driven deployments are the ones finally showing measurable earnings impact.

The takeaway for engineering leaders: the next budget fight is going to be about which process gets torn down and rebuilt, not which model is bought.

SaaS isn't dead, but the seat-pricing era might be

Source: aicoding.leaflet.pub ↗

A long read from aicoding.leaflet.pub argues the SaaS business model is about to be hollowed out from the inside. Cheap AI-assisted generation makes it cheap to spin up a tailored internal app, so the value migrates away from the standardised front-end and the per-seat licence, and toward the substrate underneath: authoritative data, compliance posture, identity and access, integrations, network effects, and operational accountability for things that break.

For a vendor, that is a re-platforming problem. For a buyer, it is a chance to renegotiate the next renewal, because the “per user per month” line item is the easiest thing to attack once the application itself is trivially replaceable.

Perplexity and Nvidia ship a local agent with zero token costs

Source: venturebeat.com ↗

Perplexity’s new Portable Computer runs agent workloads locally on Nvidia DGX Spark and RTX-powered Linux boxes. Tasks default to on-device execution, with the option to escalate to cloud models when the local model runs out of capability or context. The headline is the absence of per-token billing for work that stays on the box, which changes the unit economics for any agent loop that would otherwise be chatty and expensive.

For an SRE, the operational shape is closer to deploying a workload than integrating an API: capacity planning, model warm-up, observability around local GPU contention, and a clean hand-off to cloud when the local path is saturated.

Oracle Exadata lands on AWS, but the licence math is the story

Source: theregister.com ↗

Oracle Exadata Database@AWS is now generally available, running Oracle’s high-performance database platform inside AWS datacenters. The Register’s read, based on licensing specialists, is that the pitch is real for shops already capped on Oracle processor licences, who can collapse their licence footprint by moving onto Exadata-managed hardware. The catch is that the AWS-side compute and storage line items are higher than equivalent native AWS services, and the capacity steps are coarser. Oracle’s own sales story does not always survive contact with a real TCO model, so run the numbers before you sign.

Supabase, Anthropic, and Okta lock down enterprise MCP auth

Source: supabase.com ↗

Supabase has generally released enterprise-managed authentication for its MCP server on Team and Enterprise plans, built with Anthropic and Okta. An admin authorises Supabase inside Claude once, and access is then governed by the org’s identity provider. Each employee keeps the projects, role, and permissions they already have in Supabase, so the MCP surface inherits the existing access model rather than opening a parallel one.

The interesting bit is that this is a concrete answer to the “who is the MCP server acting as?” question that has been hanging over every Claude-in-the-enterprise deployment since MCP went mainstream. Expect other database vendors to copy the pattern quickly.

Google puts AI in the migration planner

Source: networkworld.com ↗

Google is bolting AI-powered assessments onto its Cloud Migration Center, automating the early-phase scoping and business-case work that usually lives in a consultant’s slide deck. The output is meant to feed the deeper technical validation that comes later, not replace it. Useful as a first-pass tool, dangerous if anyone treats the AI’s TCO sketch as a final number.

Open-weight models are eating enterprise AI from the bottom up

Source: constellationr.com ↗

Constellation Research is leaning on usage data from Vercel’s AI Gateway to argue that open-weight models have moved from developer curiosity to enterprise production traffic. The drivers are familiar: cost, the ability to mix models inside a single pipeline, the growing availability of managed inference, and a fit-for-purpose case for agentic workloads where fine-tuning and tool use matter more than peak benchmark scores. Notably, this is happening without enterprises having to run the weights themselves: the open-weight advantage is showing up in managed services first.

IETF 126: AI agents, IPv6, and the protocol fight underneath

Source: blog.apnic.net ↗

APNIC’s write-up of IETF 126 flags the throughline: every working group from DNS to IPv6 is now being asked to accommodate large-scale AI agent traffic. The boring but consequential work is in identity, naming, and addressing for entities that are not humans, and in capacity planning for traffic patterns that do not look like web browsing.

Apple pushes local inference on the new Mac Studio and Mac mini

Source: arstechnica.com ↗

Apple’s new Mac Studio and Mac mini are the clearest signal yet that Apple sees on-device AI as a primary use case, not a marketing bullet. The Perplexity/Nvidia story above is the same bet from a different angle: inference that does not need to leave the machine is the only way the unit economics of agentic workloads ever close, and the silicon on both sides of the Wintel/macOS line is being reshaped around that assumption.

Get the brief

Liked this one? The rest of today's stack — AI, crypto, fintech, infra — lands in your inbox tomorrow morning. Five minutes, no hype.

About Me Author

My name is

BriefTechNews

A daily digest of what actually moved in AI, tech, crypto and fintech, assembled and written with AI, and reviewed before it publishes. Read More

You May Also Like