
The economics of running AI just shifted, and most roadmaps have not caught up yet.
The economics of running AI just shifted, and most roadmaps have not caught up yet.
OpenAI and Broadcom unveiled Jalapeño this week, OpenAI's first custom inference chip,
purpose-built for large language model workloads and aiming for roughly half the cost per token of
current GPUs. Built in nine months, manufactured by TSMC, and designed to attack the memory
bottleneck that has quietly capped efficiency for years.
The performance numbers are self-reported and still need independent verification, so I would not
bank a budget on the exact figure. The direction, though, is the part that matters. Inference, not
training, is where most production AI cost actually lives. Every chatbot reply, every agent step, every
document parsed runs on inference. When the unit cost of that work falls, the math behind a whole
category of features changes with it.
For leaders, this is a planning signal more than a procurement one. Use cases that look marginal at
today's prices, high-volume summarization, always-on agents, real-time personalization, move into
positive territory when inference gets cheaper. The teams that win will be the ones who have
already mapped which ideas are blocked purely by cost, so they can move the moment the savings
land rather than starting the analysis then.
For builders, it is a reminder to keep your inference layer swappable. Abstract the model provider,
measure cost per task as a first-class metric, and design so that a cheaper backend is a config
change, not a rewrite. Custom silicon from the frontier labs means more competition on price, and
you want to be positioned to capture it.
Which of your roadmap ideas are sitting in the "too expensive to run at scale" pile right now, and
what would you ship first if inference cost halved?
Source:
Source: https://openai.com/index/openai-broadcom-jalapeno-inference-chip/
VAUGHAN ASHE
Technology leadership for FinTech and regulated industries.
Services
Technology Strategy
Digital Transformation
Cloud Architecture
AI Enablement
Technology Due Diligence
© 2026 Vaughan Ashe. All rights reserved.