Cloudflare’s “cf” CLI opens 3,000‑plus API ops; Modal shatters Kubernetes scaling limits
6 min read · 12 sources
- Cloudflare launches “cf” CLI beta exposing >3,000 API operations with JSON output
- Modal’s new sandbox scheduler creates 1 M concurrent sandboxes in <60 seconds, eliminating etcd
- Pulumi’s GeoDeploy uses TypeSafe AI’s Jev model to auto‑select optimal multi‑cloud regions
- Uber’s context‑aware retry system prevented an estimated 9.5 M unnecessary retries during a recent outage
- Bolna cuts analytics costs ~6× by migrating 15 M daily call minutes from BigQuery to ClickHouse
Autonomous agents are forcing infrastructure teams to abandon standard human-first interfaces. Wrangler capped out at roughly 280 commands, yet 48% of its active traffic last week came from non-human agents trying to navigate edge deployments. When half your tooling consumers are LLMs, human ergonomics like interactive terminal prompts and paginated text tables become immediate operational bottlenecks.
Cloudflare responded today by opening the public beta for its new agent-oriented CLI, while teams like Modal and Uber are dismantling older distributed primitives to withstand agent-scale container creation and microservice cascades.
Here is what shipped across platforms, architectures, and runtimes today.
Modal’s redesign lets a million sandboxes start in under a minute, wiping out the etcd coordination bottleneck.
Cloudflare replaces Wrangler with agent-native cf
Cloudflare unveiled the open beta of cf, an agentic CLI explicitly built to expose all 3,000+ operations across the unified Cloudflare API pipeline. While Wrangler maxed out at around 280 commands - and still saw 48% of its traffic driven by autonomous agents last week - cf treats automated systems as first-class citizens. It outputs structured JSON by default, reads configuration from a typed cloudflare.config.ts backed by TypeScript LSP, and adopts Vite for local dev execution.
The tool was generated entirely from Cloudflare’s internal Forge pipeline, eliminating the API drift that plagued Wrangler’s hand-maintained command structures. Cloudflare has committed to maintaining Wrangler for 18 months following the end of the cf beta. For platform teams managing edge workloads, the shift means you can retire custom internal wrappers and hand direct, type-safe API control to runtime agents without parsing human-formatted stdout.
Modal replaces Kubernetes and etcd to spin up 1M sandboxes
Platform engineers at Modal redesigned their sandbox architecture to break through standard container orchestration ceilings, achieving 1,000,000 concurrent sandboxes launched in under 60 seconds with sub-500ms startup times. The team removed Kubernetes and its centralized state storage entirely. Under high-churn workloads with tens of thousands of container creations per second, etcd’s consensus protocols routinely collapse under $O(\text{nodes}) + O(\text{containers})$ pressure.
Modal replaced this with horizontally scalable scheduling instances that dispatch jobs directly to workers over point-to-point RPCs. Worker state is centralized solely into a high-throughput Redis stream, which Modal validated up to 100,000 active nodes. Decoupling the scheduling logic from state reconciliation proves that hyper-churn execution environments must split the coordination and compute planes rather than relying on monolithic control loops.
Uber halts cascading retry storms at 9.5 million requests
Uncoordinated retry policies inside deep microservice call graphs remain one of the cleanest ways to take down an entire platform during degraded states. Uber detailed its error-ownership model embedded directly into its shared service-mesh infrastructure. In standard deployments, intermediary services cannot determine whether an error originated in an immediate peer or five hops down-chain, leading static retry policies to amplify requests exponentially toward failing leaf dependencies.
Uber solved this by propagating upstream failure metadata along the call path. When an downstream dependency drops, parent nodes identify that the failure was simply passed through rather than generated locally, disarming automated retries at the perimeter. In a recent production incident, the system halted an estimated 9.5 million unnecessary requests and collapsed the maximum blast radius of downstream retry storms from 25 network hops to three.
Bolna cuts analytics costs 6x moving from BigQuery to ClickHouse
Processing 15 million daily call minutes exposed fundamental ingest latency bottlenecks for voice AI platform Bolna, prompting a migration from BigQuery to ClickHouse Cloud via ClickPipes. BigQuery’s micro-batching and heavy MERGE query costs introduced continuous 5 to 15-minute dashboard delays.
The move required tuning Postgres logical replication to prevent corrupted change data capture pipelines. Bolna enabled REPLICA IDENTITY FULL to preserve TOASTed JSONB values that had previously vanished from downstream replication streams. They also shifted analytics rollups into materialized views and refactored the application tier to log only terminal state events rather than intermediate state transitions. The architectural migration dropped post-call analytics availability down to 1 - 2 minutes while slashing infrastructure spend by roughly 6x.
Self-driving cloud deployments via Pulumi and Jev
Pulumi announced an automated multi-cloud provisioning project dubbed GeoDeploy, powered by TypeSafe AI’s Jev model. Rather than leaving cross-cloud scheduling entirely to probabilistic LLM outputs, GeoDeploy draws a hard boundary: deterministic TypeScript handles balance sheets, unit costs, and API parsing, while Jev processes latency thresholds, regional budgets, and data-sovereignty constraints.
Jev evaluates real-time telemetry to determine whether workloads should route to managed Kubernetes clusters across AWS, Azure, or GCP. It passes those selections back to Pulumi code to execute deterministic plans. The pattern bridges the gap between unpredictable AI recommendations and the idempotent guarantees required by enterprise operations teams.
Tree-based document retrieval and hardened agent runtimes
Vector embeddings frequently lose context across lengthy, structured documents containing tables and cross-references. Vectify AI launched PageIndex, an open-source, vectorless retrieval framework that replaces flat vector databases and token chunking with hierarchical tree indexes. Built around a tree-generation engine called PageIndex Flash, the system uses an LLM to browse nested indices similarly to a human reader, recording a 98.7% accuracy score on the FinanceBench benchmark.
On the execution side, NVIDIA published OpenShell 0.1.x, an Apache-2.0 sandboxing runtime designed to protect hosts running autonomous coding agents. OpenShell deploys kernel-level isolation to monitor and intercept syscalls, network requests, and local disk writes. The runtime introduces an advisor-and-prover loop that mathematically verifies security policies before executing human-approved changes, and restricts secret injection only to destination endpoints validated by policy.
Database memory traps and cloud-native mocking blindspots
ClickHouse engineers published a breakdown explaining why Postgres nodes crash under unconstrained analytical queries. Postgres assigns memory allocations via work_mem (defaulting to 4MB) on a per-node, per-worker basis rather than enforcing a global per-query limit. In complex queries containing multiple parallel joins, sorts, or recursive CTEs with non-spillable executor hash tables, actual memory consumption rapidly outpaces the configured threshold, triggering the Linux OOM killer. The vulnerability is especially acute when non-deterministic SQL generated by internal text-to-SQL agents hits unmonitored replica nodes.
Meanwhile, an analysis on dependency mocking pitfalls in cloud-native systems detailed why tools like WireMock cause false-positive test runs. When upstream microservices deploy independently, static JSON stubs silently diverge from real production schemas. The post argues that next-generation integration testing requires deployment-event awareness and dynamic mocking powered by eBPF passive network capture to sync stub behaviors against actual network traffic in real time.
Infrastructure updates: Terraform Google 8.0, Lakebase Search, and AgentCore security
HashiCorp released version 8.0 of the Google Cloud Terraform provider. The major release alters default load-balancing behavior, purges deprecated resource types - including legacy Notebooks and IAP OAuth Admin endpoints - and converts a broad set of list attributes into sets to prevent unordered drift diffs in execution plans.
In the AI security domain, Harness added runtime observability for the Amazon Bedrock AgentCore Gateway. The platform monitors agent execution paths end-to-end - linking LLM prompt choices to dynamic tool calls and database lookups - to intercept complex prompt injections that look benign during isolated API calls.
Finally, Databricks announced the general availability of Lakebase Search for Lakebase Postgres across AWS and Azure. Exposing two extensions - lakebase_vector and lakebase_text - the serverless engine eliminates external vector databases for teams running hybrid full-text (BM25) and vector retrieval. Benchmark tests run on VectorDBBench 100M show the extensions achieving a P99 latency of 71ms at 97% recall, delivering double the throughput of standard pgvector deployments at a quarter of the infrastructure cost.
You May Also Like
Alibaba's OpenCodeReview, Bun's Rust Rewrite, and the Week's Biggest DevOps Shifts
Alibaba open-sourced OpenCodeReview, an Apache-2.0 AI code review CLI that matches Claude Code's precision using roughly one-ninth the tokens, though …
GKE Pod Snapshots Slash AI Cold Starts, and More Agent Infrastructure News
GKE's new Pod snapshots cut AI inference startup by up to 89%, loading a 70B parameter model in 37 seconds, a direct hit on the cold-start problem that forces …
Kubernetes finally tells you which PVCs are dead weight
Kubernetes v1.37 promotes the PersistentVolumeClaimUnusedSinceTime feature gate to Beta, so the control plane now flags PVCs no running pod references - no more …




