ChatGPT Launches 1st Chip, Aims To Challenge Nvidia

Sam Altman Says: “We Made A Chip And It Is Fast”

OpenAI has taken a major step into AI hardware with the unveiling of its first custom-designed inference chip, called Jalapeno.

CEO Sam Altman announced the development with a simple message: “we made a chip and it is fast.”

The announcement marks an important shift for OpenAI, which has historically depended heavily on Nvidia and other hardware suppliers to power its rapidly expanding AI services.

Jalapeño Is Designed For AI Inference

Unlike chips primarily designed for training AI models, Jalapeño has been built specifically for inference — the process of running trained models to generate responses.

This is particularly important for services such as ChatGPT, where millions of users continuously send requests and expect responses within seconds.

OpenAI has designed the chip’s computing resources, memory, networking and software together, with the goal of reducing delays caused by moving data between different parts of an AI system.

OpenAI Claims Major Efficiency Gains

According to OpenAI’s benchmark results, Jalapeño delivered between 1.5 and 1.9 times more AI work per watt at peak throughput compared with the systems it was tested against.

The company also reported 1.7 to 3.6 times lower end-to-end latency across the tested AI models.

For highly interactive workloads, OpenAI reported performance improvements ranging from 2.1 to 4.1 times.

These figures could be particularly significant for AI agents, where a single task may require an AI model to make multiple requests in rapid succession.

Nvidia Is Still A Major Part Of The Picture

The announcement should not be interpreted as OpenAI immediately abandoning Nvidia.

OpenAI said it will continue deploying accelerators from Nvidia and other hardware partners for both training and inference.

Jalapeño is therefore intended to supplement external hardware rather than replace it overnight.

The strategy could give OpenAI greater control over its computing infrastructure while allowing the company to use different chips depending on the workload.

The Chip Was Tested On Three AI Models

OpenAI tested Jalapeño using three large AI models: GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.

The tests used InferenceX, a public benchmark developed by SemiAnalysis that measures the process of serving AI requests.

On Kimi K2.5 1T, OpenAI reported around 1.5 times higher peak performance per watt and approximately 3.4 times lower end-to-end latency than the comparison system.

Jalapeño Uses 700 Watts

The chip is rated for 700 wattsalthough OpenAI said sustained measured power remained at or below 550 watts during the workloads tested.

Power efficiency is becoming increasingly important as AI companies build enormous data centres containing thousands of accelerators.

A chip capable of delivering more useful AI work without consuming proportionally more electricity could reduce infrastructure costs and allow data centres to serve more requests.

OpenAI Used AI To Help Design The Chip

One of the more unusual aspects of Jalapeño’s development is that AI itself was used during the chip-design process.

OpenAI said AI-assisted design helped the team reach tapeout in just nine months.

The company also used its coding tools and AI models to optimise parts of the chip and improve implementations of certain AI workloads.

This creates a striking feedback loop: an AI company used AI to help design hardware that will ultimately run AI models.

The Full-Stack Approach

OpenAI’s strategy goes beyond designing a processor.

The company has worked on the chip, memory, networking, systems and software as an integrated platform.

That approach allows engineers to optimise the entire path taken by an AI request rather than focusing solely on raw processing power.

For inference, where data movement and communication delays can become major bottlenecks, this could provide an important advantage.

Why Inference Speed Matters

AI systems are increasingly being used for interactive applications and autonomous agents.

A chatbot needs to respond quickly, but an AI agent may need to perform dozens of model calls to complete a complicated task.

If each request takes slightly less time, the cumulative improvement can become substantial.

This is why OpenAI is focusing heavily on latency alongside overall computing throughput.

OpenAI Plans Deployment By The End Of 2026

OpenAI plans to begin deploying Jalapeño inside its own computing infrastructure by the end of 2026.

The company has already started working on future generations.

According to OpenAI, Gen 2 is deep in development while Gen 3 is already taking shape.

That suggests Jalapeño is not being treated as a one-off experiment but as the beginning of a longer-term custom silicon programme.

This Could Reduce OpenAI’s Dependence On Nvidia

Building custom chips could eventually give OpenAI greater control over one of the most expensive parts of running AI services.

Nvidia’s accelerators have become the dominant hardware platform for generative AI, but they are also expensive and in extremely high demand.

If OpenAI can develop chips optimised specifically for its workloads, it could potentially improve efficiency and reduce the amount of external hardware required for certain applications.

However, the scale of any eventual savings remains uncertain.

The Benchmark Claims Need Context

OpenAI’s performance figures are significant, but they should be viewed in context.

The company selected the models and configurations used for its comparisons, and Jalapeño has not yet undergone broad independent testing in real-world production environments.

The chip’s performance against specific Nvidia systems therefore does not automatically mean it will outperform Nvidia across every workload.

The real test will come when Jalapeño is deployed at scale.

OpenAI Is Building More Of Its Own Infrastructure

The chip announcement reflects a broader change in the AI industry.

Major AI companies increasingly want control over the hardware, networking and software required to run their models.

Custom accelerators can be designed around specific workloads and potentially offer better efficiency than general-purpose solutions.

For OpenAI, that could become increasingly important as demand for ChatGPT and AI agents continues to grow.

A New Chapter In The Nvidia Relationship

The arrival of Jalapeño does not necessarily mean the end of OpenAI’s relationship with Nvidia.

Instead, OpenAI appears to be pursuing a multi-chip strategy.

Nvidia hardware can continue handling large-scale training and other demanding workloads, while custom OpenAI silicon can be optimised for specific inference tasks.

That approach could reduce dependence on any single hardware supplier without requiring OpenAI to abandon Nvidia altogether.

The Bigger Goal Is Cheaper AI

Ultimately, faster hardware is only part of the objective.

If OpenAI can process more AI work using the same amount of electricity and infrastructure, it could potentially reduce the cost of serving users.

Lower infrastructure costs could eventually support cheaper AI services, more usage or larger and more capable AI systems.

That makes efficiency just as important as raw speed.

Jalapeño Could Be The Beginning, Not The Destination

OpenAI’s first custom chip represents a significant expansion of the company’s ambitions.

It is moving beyond developing AI models and applications and increasingly taking control of the infrastructure underneath them.

Jalapeño’s early benchmark results suggest substantial improvements in efficiency and response times, although independent real-world testing will ultimately determine how meaningful those claims are.

With second- and third-generation chips already in development, OpenAI appears to be building a long-term hardware business around its own AI ecosystem.

OpenAI Wants To Control More Of The AI Stack

The Jalapeño project shows how quickly the AI competition is expanding beyond models.

The next battleground is increasingly the infrastructure that powers them — from chips and memory to networking and software.

OpenAI still expects to use Nvidia and other accelerators, but its own silicon gives the company another option.

If Jalapeño performs as promised at production scale, OpenAI could gain greater control over performance, efficiency and eventually the cost of running its AI services.

Summary

OpenAI has unveiled Jalapeño, its first custom AI inference chip, with the company claiming up to 1.9 times more AI work per watt and up to 3.6 times lower latency than comparison systems. The chip was developed with an integrated approach covering computing, memory, networking and software. OpenAI plans deployment by the end of 2026, while Gen 2 and Gen 3 are already in development.

Image Source


Leave a Comment