AI & Tech Daily
OpenAI’s Custom Chip Signals a New Phase in the AI Infrastructure Race
OpenAI publishes the first benchmark results for its Jalapeño inference chip, promising lower latency and more work per watt while leaving production economics unproven. Intel outlines processors for data-centre and edge agents, and Cisco adds Supermicro rack systems to its NVIDIA infrastructure stack. OpenAI disrupts a covert influence campaign that used ChatGPT mainly for amplification, while Mistral packages multi-step search for enterprise agents. ARIA draws a line around predominantly AI-generated music, Microsoft shifts enterprise roadmap planning to a continuous feed, and OpenAI gives eligible administrators a conversational interface for managing ChatGPT Work and Codex.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
OpenAI’s Jalapeño Chip
OpenAI says its first custom inference chip can serve large AI models with much lower latency and more work from each watt. The catch is that every published result still comes from a pre-production laboratory.
That tension is why this development deserves a closer look. On 25 August, OpenAI published its first benchmark results for Jalapeño, a chip designed specifically for inference: the work involved in running a trained model and producing answers. OpenAI plans to deploy it in its own infrastructure by the end of 2026, once the chip completes production qualification.
The company reports between 1.5 and 1.9 times more AI work per watt at peak throughput than the comparison systems it tested. It also reports end-to-end latency between 1.7 and 3.6 times lower. The workloads included GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, so this wasn’t confined to one small model shaped around the new hardware.
Those are striking ratios, particularly for a company serving models at enormous scale. Faster inference can make an assistant feel more responsive. Better energy efficiency can lower the cost of producing each answer, or let a data centre serve more requests within the same power envelope. At OpenAI’s scale, even a modest improvement could affect both capacity planning and service economics.
There is some outside scrutiny, but not an independent production test. SemiAnalysis says it ran the InferenceX benchmarks alongside OpenAI engineers in an OpenAI laboratory. That gives the results more specialist examination than a chart produced solely by a marketing team, but the environment and hardware remained under OpenAI’s control.
Several unanswered questions are at least as important as the benchmark numbers. OpenAI hasn’t disclosed the planned deployment scale, the unit cost, manufacturing yield or expected operational savings. Production reliability hasn’t been demonstrated, and the comparison will keep moving as other chipmakers ship newer hardware. A design that looks excellent in qualification can encounter very different constraints once thousands of accelerators are running continuously across real services.
Jalapeño also targets inference, not model training. OpenAI says it will continue using NVIDIA and other suppliers, so this isn’t a clean break from merchant accelerators. It’s another option in the serving stack: one that could reduce exposure to a single supplier, give OpenAI more control over optimisation, and reserve general-purpose accelerators for workloads where they remain the better fit.
My reading is that the strategic value may exceed the headline benchmark victory. Owning part of the silicon roadmap lets OpenAI tune models, compilers, memory behaviour and data-centre operations together. It also gives the company leverage when negotiating for external capacity. But infrastructure buyers shouldn’t treat the reported ratios as a durable competitive lead until there’s evidence from production-scale deployment. For now, Jalapeño is a credible signal that OpenAI wants deeper control of inference economics, not proof that it has solved them.
Intel’s Three-Layer Agent Hardware
That brings us out of OpenAI’s lab and into the wider hardware contest, where the announcements are broader but less measured.
At Hot Chips 2026 on 24 August, Intel outlined three planned architectures for agentic workloads across data centres and edge devices. Diamond Rapids Xeon processors are aimed at orchestration and CPU-heavy work. Crescent Island is an inference GPU. Wildcat Lake brings local AI processing and a new chiplet connection into client systems.
Diamond Rapids is specified with as many as 256 cores, 1.28 gigabytes of last-level cache, 16 memory channels and 128 lanes of PCIe Gen 6 or CXL 3. Those lanes matter when a server needs to connect accelerators, memory and networking components without creating new bottlenecks. A large cache and memory footprint can also help the coordination, retrieval and data movement surrounding an agent, even when a GPU performs the main model computation.
Crescent Island is the more direct inference play. Intel says it will support up to 480 gigabytes of LPDDR5X memory in a 350-watt PCIe card that can be air-cooled. That combination could appeal to operators who want substantial model capacity without redesigning a facility around liquid cooling.
Wildcat Lake, meanwhile, is described as Intel’s first processor using UCIe, an open interconnect for linking chiplets within a package. Its neural processor is rated at up to 17 TOPS, a measure of theoretical operations per second. That number alone says little about how a complete agent will perform, but it shows Intel positioning local processing as one piece of a system that may span device and data centre.
Intel didn’t provide shipment dates, prices or comparative performance results. Software maturity, real power efficiency and product availability remain open questions. My practical assessment for infrastructure planners is simple: the air-cooled GPU and open chiplet approach could widen deployment choices, but these specifications aren’t solid enough for a capacity commitment. Treat them as an architecture signal until Intel supplies measured workloads and firm delivery windows.
Cisco and Supermicro Build the Rack
Silicon is only useful once somebody turns it into racks that can be powered, cooled, connected and operated without months of custom integration.
Cisco announced on 25 August that its Secure AI Factory with NVIDIA will expand through a partnership with Supermicro. The plan adds validated, rack-scale Supermicro systems in both air- and liquid-cooled configurations, alongside Cisco networking and management components.
The proposed stack combines Supermicro compute with Cisco Silicon One networking at the front end and NVIDIA Spectrum-X Ethernet at the back end. Cisco Nexus One provides the management layer. The companies are pitching configurations for enterprises, neocloud operators and sovereign-cloud deployments, including designs that comply with NVIDIA’s Cloud Partner requirements.
Availability is expected to begin in October 2026. Cisco hasn’t disclosed pricing, independently measured performance or efficiency results, and delivery capacity is still unknown. Everything published so far is a set of vendor claims rather than evidence from operating customer environments.
There is still a useful consequence. Buying a validated combination of compute, networking, cooling and management can remove a considerable amount of integration work. Organisations that lack hyperscaler engineering teams may get to a functioning AI cluster more quickly and with clearer support boundaries.
The trade-off is concentration. Standardising on a combined Cisco, NVIDIA and Supermicro design could make later component changes harder, particularly when operational tools and network assumptions become embedded around it. My advice for buyers evaluating the stack would be to price the exit as carefully as the installation. Faster deployment has real value, but so does knowing whether a future accelerator, network or management platform can be introduced without rebuilding the whole rack.
Influence Through Amplification
The next story shifts from infrastructure control to information control, and the most revealing detail is what the model apparently didn’t do.
OpenAI said on 25 August that it had banned a cluster of ChatGPT accounts it assessed as very likely originating in Russia. According to the company, the accounts were used in a covert influence campaign that amplified material from an organisation presenting itself as a policy institute.
OpenAI says the operators used VPNs, prompted mainly in Russian, and distributed material across several social and publishing platforms. In a sample of 36 institute articles published from September 2025 to May 2026, OpenAI found that 34 had been copied to other sites. Many copies carried misleading attribution.
The models were used mainly to support amplification rather than to write the institute’s core articles. That distinction is important. Public discussion of AI-enabled influence often focuses on whether a paragraph sounds synthetic. In this case, the more consequential activity was building volume around existing material: repackaging it, generating comments or social posts, and making a small operation appear more distributed than it was.
OpenAI described the campaign as elaborate but limited in audience reach. Its assessment of the origin, scale and impact hasn’t been independently adjudicated, so the attribution needs to remain attached to the company making it. The disruption demonstrates what OpenAI observed on its own service; it doesn’t establish the campaign’s complete activity across the internet.
For security and communications teams, the lesson is that prose detection is too narrow. An individual post may contain little or no model-written text and still belong to an AI-assisted influence operation. Repeated identities, synchronised distribution, copied material, false attribution and unusual account relationships can be stronger signals than the wording of any single message.
My view is that provenance and network analysis are becoming essential defensive tools. Organisations monitoring manipulation will need to ask where content first appeared, how rapidly it spread, which accounts moved together and whether the stated publisher matches the real source. Synthetic writing may be one clue, but treating it as the defining feature risks missing the operational pattern around it.
Mistral’s Multi-Step Search
There’s a related engineering problem behind many useful agents: finding the right evidence is often harder than generating the final answer.
Mistral released Agentic Search on 20 August through its Search Toolkit and development libraries. Instead of sending one query to an index and accepting the first batch of results, a model can search, open documents, navigate, read and grep indexed material across several steps. The system is available through Mistral Studio and Vibe, and Mistral says it can run in its cloud or on premises.
That multi-step loop is designed for questions whose evidence is scattered across files or buried inside long documents. An agent can inspect an initial result, notice that a figure depends on another section, open that section and search again. Developers have been building similar tool chains themselves; Mistral is turning the pattern into a supported retrieval layer.
The company reports a large improvement in selected test configurations. On FinanceBench, accuracy rose from 26.7 per cent to 86 per cent. On OfficeQA Pro, it increased from 6.3 per cent to 51.9 per cent. Mistral also says that using navigation reduced FinanceBench’s 90th-percentile latency from 255 seconds to 154 seconds and cut token use by as much as one-third.
Those are vendor benchmarks, not independently replicated results, and they vary by model and configuration. The OfficeQA score is also a useful reality check: moving from 6.3 to 51.9 per cent is substantial, but it still leaves nearly half the questions unresolved in that tested setup. More tool use doesn’t make hard documents reliable by default.
Customers also need an existing index, along with ingestion, ranking and evaluation work. Agentic navigation can’t recover information that was parsed badly, omitted from the index or made inaccessible by permissions. For simple, high-volume lookups, one-shot retrieval may remain faster and cheaper than allowing a model to explore.
For developers, the immediate gain is less plumbing. A packaged search-and-navigation loop can shorten the path to a capable research or document-analysis agent. My caution is to evaluate the entire retrieval system on real organisational data, including failures and cost, rather than judging it by the model’s final prose. Search tools can improve the journey to an answer; they don’t certify that the destination is correct.
ARIA Draws a Human Line
A different kind of boundary is appearing in Australian music, where eligibility now depends on who actually performed the recording.
On 25 August, the Australian Recording Industry Association adopted rules excluding wholly or predominantly AI-generated recordings from its charts and awards. Substantially human-created music can remain eligible when AI is used as an assistive tool.
Under ARIA’s definition, an eligible AI-assisted recording remains substantially human-made. Humans perform the lead vocal or primary instruments, and any AI service involved has to be used lawfully and with authorisation. The rule applies from the chart dated 31 August 2026, which is scheduled for publication on 28 August.
This doesn’t prevent streaming services from carrying predominantly generated recordings. It governs recognition through ARIA’s charts and awards. Artists, labels and distributors seeking that recognition will therefore need to distinguish assistance from generated performance and be prepared to show that the tools were used legitimately.
That sounds straightforward at the extremes. A human band using software to clean up a recording is quite different from a system generating the principal vocal and instrumental performance. Mixed cases will be harder: a song may involve human composition, generated instrumentation, edited synthetic vocals and extensive post-production, with no single step deciding the outcome.
ARIA hasn’t publicly described a comprehensive technical verification process. That leaves uncertainty around disclosure, evidence and enforcement when contributions are difficult to separate.
My assessment is that the rule establishes a visible preference for human performance without trying to ban AI-assisted production. Its credibility will depend on whether comparable cases receive comparable treatment. If the boundary rests largely on voluntary descriptions, the most conscientious artists could face more scrutiny than creators who reveal less.
Microsoft Ends Release Waves
Enterprise software planning is changing too, although this shift is about the timetable customers see rather than the code they receive.
Microsoft announced on 25 August that roadmap material for Dynamics 365, Power Platform and Dataverse will move into its continuously updated AI at Work roadmap from September 2026. The existing release-wave 1 and release-wave 2 publication model will be retired.
Entries in the new roadmap move through three states: In Development, Rolling Out and Launched. Customers can filter the information or consume it through CSV, RSS and Microsoft’s Release Communications MCP Server. The change affects how planned work is published and tracked; it doesn’t by itself alter the underlying software deployment cadence.
Microsoft also warns that roadmap dates are estimates and may change. Existing personalised views in the release planner won’t transfer automatically, so teams relying on saved selections will need to rebuild their monitoring around the new system.
Twice-yearly release waves gave change-management teams two obvious planning events. A continuous roadmap can provide fresher information and reduce the lag between a product decision and customer awareness, but it removes that convenient rhythm. Someone has to watch the feed, assess changes and route relevant items to security, training, support and business owners throughout the year.
My take is that this can reduce surprises only for organisations that build continuous review into their governance. CSV, RSS and an MCP interface make automation possible, but an automated alert still needs ownership and a decision process behind it. The publication system may become more current; whether forecasts become more accurate remains unproven.
What Changes for You
One practical release could remove a lot of console-hopping for the people who administer workplace AI.
OpenAI released an Admin plugin on 25 August for eligible ChatGPT Work and Codex environments. Authorised administrators can use a conversation to review activity and credit use, manage members and groups, inspect effective permissions and model access, and handle usage limits or spending requests.
The plugin can also configure recurring checks and route approval requests into Slack or Microsoft Teams. OpenAI says it respects existing roles and permissions, and broader-impact actions can be presented for review before execution. Access requires an eligible workspace, administrative permissions and plugin enablement through the web or desktop client.
For administrators, the immediate change is that routine questions and operations can happen through one interface. Instead of visiting several management views to learn why a person lacks model access or where credits are being consumed, an authorised operator can ask directly and move into the relevant action.
The limitation is just as concrete: this is a privileged natural-language control surface for permissions and spending. OpenAI hasn’t published independent reliability measurements or error rates for its administrative actions. Organisations using it will still need tightly designed roles, meaningful review for high-impact changes and audit records that show what was requested and what actually happened.
Conversational administration may make everyday work quicker, especially for teams already coordinating approvals in Slack or Teams. It doesn’t make the underlying authority less sensitive. The useful standard is whether the plugin preserves clear human accountability while reducing operational friction, not whether it can make an access change in fewer clicks.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- openai.com/index/jalapeno-first-results
- newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
- newsroom.intel.com/de/client-computing/intel-outlines-architectures-for-agentic-ai-at-hot-chips-2026
- newsroom.cisco.com/c/r/newsroom/en/us/a/y2026/m08/cisco-secure-ai-factory-nvidia-rack-scale.html
- openai.com/index/disrupting-malicious-uses-of-ai-influence-campaign-russia
- mistral.ai/news/agentic-search
- aria.com.au/chart-ai-definition
- apnews.com/article/9bfb0c91166ae4405a6df1a3c4891687
- publicnow.com/view/123E504D673FC3762FC77D2A2E353AD2F7B9FF14
- learn.microsoft.com/en-us/copilot/release-plan/2025wave1/copilot-finance
- openai.com/index/introducing-admin-plugin