As AI systems evolve from training to everyday deployment, the focus is shifting toward more efficient, cost-effective inference hardware , exemplified by OpenAI’s new Jalapeño chip , driving a surge in semiconductor demand and reshaping industry strategies.
For much of the past three years, the AI arms race has been defined by model training: larger clusters, heavier compute bills and a relentless push for scale. That focus is now widening. The more commercially important question is increasingly how cheaply and efficiently systems can be run once they are built.
Gartner expects spending on AI inference to reach $23.3 billion in 2026, ahead of the $19 billion forecast for training. The shift reflects a basic change in workload. Agentic AI systems do not just answer one prompt; they can complete long chains of actions, multiplying the number of model calls and making running costs far more significant than in a simple chatbot setup.
That is already reshaping hardware strategy. Semiconductor firms and AI developers are designing chips with inference in mind, where latency, energy use and cost per request matter more than peak training throughput. OpenAI’s newly disclosed Jalapeño accelerator is the clearest sign of that move. According to published benchmark claims, the chip is built specifically for inference, not model training, and was developed with Broadcom.
The numbers attached to Jalapeño are striking. Tom’s Hardware reported that OpenAI says the chip uses about 700W, compared with Nvidia systems drawing roughly 1,200W to 1,400W, while delivering better throughput per kilowatt and lower latency in early tests. The same reports say the design includes six HBM4 memory stacks and is intended to support large language model serving at scale.
This is happening against a broader backdrop of booming semiconductor demand. Gartner has projected that global semiconductor revenue could nearly double in 2026 to $1.6 trillion, with memory components taking a far larger share of the market as AI infrastructure expands. That implies a simple but important conclusion: as AI moves from development into everyday use, the cost of serving models may become a bigger competitive issue than the cost of training them.
Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.





