CoreWeave announced on September 30 that NVIDIA Vera Rubin NVL72 systems are in limited availability on CoreWeave Cloud. Cognition, the applied AI lab behind Devin, is the first customer running production workloads on the platform.
On SWE-2 inference workloads drawn from FrontierCode tasks, NVIDIA Vera Rubin NVL72 delivered up to 4.8 times higher total token throughput per GPU than a GB200 NVL72 baseline at matched interactivity. The workload was generated by sampling a subset of FrontierCode tasks and having AI agents solve them.
Higher token throughput does not equal higher task success rate or greater intelligence.
NVIDIA Vera Rubin has moved beyond roadmap slides. Production agent tasks now run on the system.
Cognition’s engineers performed the measurements themselves.
All figures come from Cognition operating on CoreWeave infrastructure, and independent third-party verification has not yet appeared.
Agentic Workloads Place Different Demands on Hardware
Most generative systems focus on training a model and then serving tokens. Agentic systems add many extra steps. Agentic systems can also be connected to post-training loops where production traces and successful or failed outcomes inform later evaluations and model improvements.
Software-engineering agents like Cognition’s Devin demonstrate the pattern clearly, repeatedly scanning large repositories and chaining dozens of actions. Latency at any single step multiplies across the whole sequence. Higher token throughput does not improve task success rates or intelligence on its own. Faster throughput can shorten wait times between successive reasoning steps.
CoreWeave points out the resulting pressure falls on GPUs, CPUs, networking fabric, and large numbers of isolated execution environments simultaneously.
NVIDIA Vera Rubin CPU Addresses the Sandbox Constraint
NVIDIA designed Vera as the first CPU built for AI agents. CoreWeave will bring the processor to its cloud. In CoreWeave testing, Vera produced more than 3× faster agent-sandbox startup than an unnamed alternative x86 CPU. On Terminal-Bench, the company recorded a 1.7x performance gain across every passing task.
Software-engineering agents repeatedly process large repositories, execute code, and iterate through long task chains, placing simultaneous pressure on GPUs, CPUs, networking, and thousands of isolated environments.
CoreWeave first brought up and validated Vera Rubin NVL72 in June; the September announcement marks the move into limited customer production, with Cognition as the first production user.

CoreWeave’s standalone rack-scale Nvidia Vera CPU configuration packs 128 CPUs and 11,264 cores; this should not be confused with a Vera Rubin NVL72 rack, which combines 36 Vera CPUs with 72 Rubin GPUs. According to NVIDIA and CoreWeave, the rack has enough CPU cores for more than 11,000 concurrent one-core environments.
CoreWeave Sandboxes keep each environment hardware-isolated so the environments can run alongside the training jobs they support. Spectrum-X Ethernet switches and BlueField-4 DPUs are designed to provide secure, high-performance, low-latency connectivity between the nodes. The standalone Vera CPU offering is not yet generally available on CoreWeave; the company says it is coming soon.
This is separate from the Vera CPUs already integrated inside the Vera Rubin NVL72 system.
CoreWeave Forge Links Production Signals to Improvement
Hardware forms only one part of the announcement. CoreWeave also launched Forge, a connected environment designed to keep production behavior tied to the next improvement cycle.
Forge covers the AI loop – run, observe, curate, improve, and evaluate – combining Weights & Biases Models, OpenPipe expertise, marimo and CoreWeave services.
New or expanded services include Agent Lens, which converts production traces into concrete insights but is not generally available. ARIA is now generally available for analyzing experiment runs and proposing follow-up tests. Sandboxes themselves are generally available.
Canva, Capital One and MasterClass rank among the first companies already building inside Forge. The environment stays open across different models, frameworks, and even other clouds.
| Metric | Reported Result | Comparison Baseline | Attribution |
| SWE-2 inference token throughput | Up to 4.8× per GPU at matched interactivity | GB200 NVL72 | Cognition on CoreWeave |
| RL output-token throughput | 3.8× per GPU at matched interactivity (this is a Cognition benchmark, not proof that all RL training on Rubin is 3.8× faster.) | GB200 NVL72 | Cognition on CoreWeave |
| Agent sandbox startup time | More than 3× faster | Prior CoreWeave setup | CoreWeave testing |
| Terminal-Bench performance across passing tasks | 1.7× | Comparison baseline not specified in NVIDIA’s announcement | CoreWeave testing |
Context Across Generations of Hardware
CoreWeave and NVIDIA have co-engineered infrastructure for nearly a decade. Older GPU generations continue to serve customer workloads while newer systems arrive. Ian Buck of NVIDIA noted that CoreWeave’s original V100 GPUs still run production jobs almost ten years after Volta launched.

Nvidia Vera Rubin NVL72 extends the same durable approach into agentic production.
The new systems’ capacity runs through CoreWeave Kubernetes Service, SUNK, Mission Control, Sandboxes, and CoreWeave Inference. CoreWeave says Cognition brought production workloads online within days of rack handover.
Practical Implications Remain Open
Whether the measured throughput and sandbox gains produce lower operating costs, shorter iteration cycles, or higher-quality software depends entirely on the specific agent and the surrounding engineering process.
CoreWeave has not published Nvidia Vera Rubin capacity pricing or a date for broader hardware availability; Forge, separately, is offered in Free, Pro, and Enterprise editions.
The longer question concerns how infrastructure must evolve once AI systems perform extended sequences of real work rather than isolated generations.