AI & Tech Daily
NVIDIA’s New Inference Stack Puts Speed, Cost and Lock-In in Play
NVIDIA’s Groq 3 LPX system is entering full production, splitting long-context processing from latency-sensitive token generation in a tightly integrated rack-scale stack. Thomson Reuters is taking a different route to control by deploying its own specialised professional model. Australia has passed a charge-and-offset scheme for large search and social platforms, while reported progress on TSMC’s A16 process points to backside power delivery as a growing source of chip gains. The US Treasury has formed a quantum-readiness task force for finance, Google and Verizon are extending Gemini into operational systems, and organisations running Zimbra face an actively exploited command-injection flaw. For developers, GitHub Copilot’s Teams preview can now turn a shared conversation into coding-agent work, with existing repository controls still applying.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
NVIDIA Splits the Inference Job
An AI agent can be clever and still feel painfully slow. NVIDIA is now putting specialised token-generation hardware into production to attack the delays that accumulate every time an agent takes another step.
This is worth spending some time on because the system separates two parts of inference that are often bundled together. NVIDIA said on August 24 that its Groq 3 LPX rack-scale system is in full production as an extension to Vera Rubin NVL72. Rubin GPUs handle the heavy context processing, while Groq’s LPUs concentrate on generating output tokens with low latency.
The first stage is commonly called prefill: the system processes the prompt and its supporting context. Decode comes after that, producing the response token by token. Using different processors for those stages gives an operator room to tune each one for its particular workload instead of asking one type of accelerator to handle both in the same way.
That distinction becomes important when an agent is working with a long prompt, consulting tools and repeatedly deciding what to do next. Processing a large body of context is computationally demanding, but once that work is done, the user also cares about how quickly the system emits each new token. In a multi-step workflow, a modest delay can recur across planning, tool calls, intermediate results and the final response. Faster decoding can therefore improve the experience by more than the speed of one isolated answer suggests.
An LPX rack can contain 256 LP30 accelerators. NVIDIA named Nebius as the first AI-cloud adopter, with access planned for later in 2026, and said CoreWeave is deploying the associated Spectrum-X Multiplane networking. Those details show that NVIDIA is selling a coordinated compute, networking and serving architecture, not merely another accelerator card.
The headline performance number needs careful handling. NVIDIA reports that an LPX system produced 3,431 output tokens per second for Gemma 4 31B after receiving 100,000 input tokens in an Artificial Analysis test. That is striking for that particular model and configuration. It was also measured on a system hosted in NVIDIA’s own data centres. It doesn’t establish what customers will pay, how much power a rack consumes, or how the system performs across other models and real production traffic. A production service also has to manage concurrent requests and changing context lengths, conditions that one configuration-specific result can’t settle.
Nebius access is still planned rather than generally available, so developers can’t yet assume this capacity will appear as an ordinary cloud option. Independent production benchmarks, comparable energy measurements and broad pricing are also missing.
The larger shift is that inference is becoming its own competitive engineering layer. Model intelligence still counts, but so do decoding latency, network design, utilisation and the software that schedules work across different processors. For AI-cloud operators, a specialised decode tier could make long-running agents more responsive and potentially improve the economics of serving them.
My read is that those gains will come with a sharper infrastructure choice. An operator adopting the full arrangement may benefit from components designed to work together, while becoming more dependent on one vendor’s accelerators, interconnects and serving stack. Before treating the token figure as a buying case, organisations need workload-specific evidence covering latency, throughput, power, utilisation and cost. NVIDIA has now made the hardware production-bound; the unresolved question is whether its advantage survives outside the demonstration configuration.
Thomson Reuters Builds Its Own Model
That’s the big infrastructure bet. A different kind of control is appearing inside organisations that possess data valuable enough to shape a model themselves.
Thomson Reuters launched a proprietary language model called Thomson on August 24. It begins with Snowdon, an open-weight foundation model from Imperial College London, then adds continued pre-training using Thomson Reuters material from legal, tax, accounting and news collections. Selected CoCounsel workflows are due to start using it in production this month.
The company says less than 10 per cent of its content has been used in the continued pre-training so far. That figure is notable in both directions. Thomson Reuters has considerable unused material available for future training, but it also suggests that the present model’s results come from a limited portion of the company’s holdings rather than the whole collection. More data won’t automatically guarantee better performance; selection and evaluation still matter.
Thomson Reuters says it spent US$40 million on talent and compute and fully owns and controls the resulting model. That ownership changes the operating equation. A business using a rented frontier-model API depends on another company’s prices, release schedule, model changes and service boundaries. Running an internally controlled model creates its own costs and technical burden, but it gives the owner more influence over training data, evaluation and deployment timing. For legal and professional work, where factual reliability is especially important, that control can be strategically useful.
Thomson Reuters reports a factuality score of 0.83 for its Westlaw and Practical Law research system, compared with 0.65 and 0.68 for two unnamed frontier systems using open-web access. Those are company-run comparisons. The frontier systems aren’t identified, and a full technical report plus independent evaluation are still pending. Model size and external access terms also haven’t been published.
The practical development isn’t that a specialised model has conclusively beaten frontier systems. It’s that a large information business now considers its data and workflow knowledge sufficient to justify owning part of the model layer. Selected CoCounsel tasks provide a bounded place to test that choice in production before any wider legal or tax integration.
For organisations with deep proprietary collections, a narrower model can now be worth evaluating alongside frontier APIs. The trade is frontier flexibility for greater control over costs, data use and releases. The evidence still needs to be tested against the exact professional tasks where errors carry consequences.
Australia’s New Platform Charge
Control over valuable information also sits behind Australia’s latest attempt to reshape the relationship between digital platforms and journalism.
Parliament passed the News Bargaining Incentive and News Journalism Payments scheme on August 20. The framework creates a charge-and-offset system for large search and social-media platforms, aiming to encourage commercial agreements with Australian news businesses.
A service comes within scope when it belongs to a corporate group earning more than A$250 million in relevant Australian digital-advertising revenue. The charge is 2.5 per cent of the applicable advertising-revenue base. A platform can reduce or eliminate that liability through eligible spending involving at least eight Australian news-business corporate groups.
That structure differs from simply ordering a platform to carry particular news or pay a fixed amount to each publisher. It gives covered businesses a financial choice: enter qualifying arrangements across a minimum spread of news groups, or face the statutory charge. The framework is also intended to remove withdrawing news as a simple way to avoid bargaining.
The eight-group requirement creates a measure of breadth, but it doesn’t prescribe equal payments or reveal how much each participating business might receive. A platform could satisfy the spread requirement through agreements of very different value. That makes the eventual deal structure at least as important as the number of publishers involved.
Passage doesn’t tell us how the market will respond. The number and value of agreements remain unknown, as does the eventual distribution of funds across large and small publishers. Platforms may build different combinations of commercial deals, and qualifying expenditure doesn’t by itself reveal whether a payment will support additional reporting or mostly reinforce existing operations.
For covered platforms, this is now a material commercial and regulatory calculation rather than a voluntary public-relations exercise. For Australian news businesses, it creates another mechanism for bargaining over the value that journalism contributes to digital services.
My assessment is that the law’s real test will be distribution. Securing payments would meet only part of the policy goal if most of the benefit remains concentrated among the largest publishers. The requirement to involve at least eight corporate groups creates some breadth, but sustainable funding across the sector will depend on the deals platforms choose and how the payment scheme operates in practice.
A16 Puts Power Behind the Chip
At chip scale, the next gain may come from changing how electricity reaches the transistors, not merely packing in more of them.
Reporting based on Taiwanese industry sources says TSMC has completed development and validation of its A16 process and is targeting volume production in the fourth quarter of 2026. TSMC hasn’t publicly confirmed that validation milestone or the more precise quarter. Its official position is that A16 is scheduled to become production-ready in the second half of 2026.
A16 uses TSMC’s Super Power Rail backside power-delivery system. Conventional chip layouts route power and signals through the front side, where they compete for limited space. Moving power delivery to the back separates it from signal wiring, which can reduce routing congestion and improve the delivery of power to dense, high-performance designs.
TSMC projects that, compared with its N2P process, A16 can provide 8 to 10 per cent higher speed at the same voltage, 15 to 20 per cent lower power at the same speed, and up to 1.10 times the chip density. The speed and power figures represent different choices a chip designer might make: push performance while holding voltage steady, or preserve performance while cutting energy use. They aren’t benefits that can simply be added together.
These are company projections, not measurements from shipping customer products. TSMC is positioning the process mainly for high-performance computing, where dense power delivery and complicated signal routing are serious design constraints.
For AI-chip designers, the potential gain is broader than a smaller process label. Better power delivery could support faster designs or reduce energy use, both of which matter in accelerators deployed by the rack. Product schedules will still depend on yield, available foundry capacity and each customer’s qualification work.
The useful conclusion is that advanced-chip competition is increasingly about the full physical design. Transistors remain central, but backside power, packaging, memory and interconnects now determine how much usable performance reaches a system. The reported validation is encouraging, though customers need TSMC-confirmed ramp details before building a fourth-quarter assumption into firm plans.
Finance Starts Preparing for Quantum
Some infrastructure changes arrive with firm product dates. Quantum risk is harder: the timing is uncertain, but the migration work can still take years.
The US Treasury launched a public-private Quantum-Readiness Task Force on August 24 to coordinate the financial sector’s transition towards post-quantum cryptography. That means encryption and signature systems designed to withstand attacks from future quantum computers capable of breaking widely used public-key methods.
The task force has three workstreams. One covers alignment across the sector and the cryptographic transition itself. A second focuses on the readiness of suppliers and other third parties. The third examines digital assets and risks associated with emerging technology. Its remit includes identifying important dependencies, improving cryptographic agility, supporting interoperability and dealing with implementation problems that cross organisational boundaries.
The announcement doesn’t create a new compliance deadline, and it isn’t evidence that a cryptographically relevant quantum computer is about to arrive. Membership details, milestones, reporting expectations and enforceable sector deadlines haven’t been announced either.
What it does create is a coordination point for a sector with unusually complicated dependencies. A bank may control part of its own cryptography while relying on payment networks, cloud services, software vendors, identity providers and hardware suppliers for the rest. Replacing an algorithm in one application is manageable; changing it across interconnected institutions while preserving compatibility and auditability is much harder.
A useful cryptographic inventory therefore needs more than a list of software packages. Institutions need to understand which business services depend on each cryptographic component and which outside supplier controls the upgrade path. That dependency mapping is where a sector-wide task force may add value, because one organisation’s migration can be held back by another organisation’s interface or timetable.
For financial institutions and their technology suppliers, the sensible interpretation is to improve crypto-agility now: know which algorithms are in use, where keys and certificates live, and which vendors could delay a migration. That isn’t a prediction about the arrival date of a capable quantum machine. It’s recognition that uncertainty about the deadline makes a poorly mapped dependency chain more dangerous, not less.
Gemini Moves into Verizon’s Operations
A large enterprise deployment gives us another view of the stack question, this time inside day-to-day telecommunications operations.
Google Cloud and Verizon announced on August 24 that they’re expanding Gemini Enterprise, Google’s data stack and agent tooling across customer service, network operations, marketing, security and employee workflows.
The companies say their existing Gemini-based contact-centre platform handles the majority of Verizon’s inbound consumer calls and chats each month. They didn’t provide absolute volumes, independently audited performance results or a breakdown of what the system handles without human intervention. Verizon also plans an autonomous network-intelligence framework intended to predict and resolve anomalies before customers notice them.
That network plan could be consequential. Customer support errors are visible and frustrating, but an automated decision inside network operations can affect service reliability at much larger scale. The partners haven’t disclosed the boundaries for autonomous action, the deployment timetable or the value of the agreement, so it’s too early to judge how much operational control will actually move to the system.
There’s also a difference between broad deployment scope and demonstrated results. The announcement names many business functions, but the strongest current scale claim concerns the existing contact-centre platform, and even that claim comes from the two partners without absolute figures. The planned network framework remains just that: planned.
Verizon’s scope nevertheless shows generative AI moving beyond a collection of employee assistants. It is being positioned across customer-facing systems, internal work and infrastructure operations under one strategic partnership.
For enterprise buyers, my takeaway is that choosing an AI platform increasingly resembles choosing operational infrastructure. Model quality is only one part of the decision. Observability, approval controls, data governance, failure recovery and concentration in one provider become more important as the same stack reaches more workflows. The possible automation gains rise with that reach, but so does the cost of a shared failure or poorly governed action.
Zimbra Flaw Is Already Being Exploited
There’s also a security item where the uncertainty is no reason to wait.
CISA added CVE-2026-73570 to its Known Exploited Vulnerabilities catalogue on August 21 after observing exploitation. The flaw affects Zimbra’s SNMP notification processing when that feature is enabled and allows unauthenticated command injection. In plain terms, an attacker doesn’t need a valid account to exploit the vulnerable component and execute commands on the server.
Zimbra fixed the issue in Collaboration Suite 10.1.20 and strongly recommends upgrading. US civilian federal agencies were required to remediate it by August 24. CISA’s deadline directly governs those agencies, but the confirmed exploitation is the important signal for every organisation operating an affected server.
Administrators first need to know whether the affected SNMP notification feature is enabled across their Zimbra estate. That configuration check helps define exposure, but it doesn’t replace the upgrade on an affected deployment or the search for signs that exploitation occurred before remediation.
Public exposure counts can’t tell us how many reachable systems remain vulnerable, and they don’t establish how many have already been compromised. The identity, scale and objectives of the attackers also haven’t been publicly established. Those gaps should shape the investigation, not delay it.
An upgrade closes the known vulnerability, but it can’t remove commands an attacker may already have run or undo persistence placed on a system before patching. Administrators need to examine recent logs, filesystem changes and other indicators of compromise alongside the upgrade.
I’d treat this as incident response rather than ordinary patch management. Confirmed active exploitation changes the working question from whether a vulnerable server might eventually be targeted to whether someone may already have reached it. For organisations using the affected Zimbra configuration, version 10.1.20 is the immediate floor, and post-upgrade review is part of the remediation rather than an optional extra.
What Changes for You
For development teams, one new preview could remove a surprisingly awkward handoff between discussing work and starting it.
GitHub Copilot is now available inside Microsoft Teams as a public preview. After an organisation installs the GitHub app, authorised users can mention it in a Teams channel, group chat, meeting chat or one-to-one conversation and ask the coding agent to act on the surrounding discussion.
Supported work includes implementing features and fixes, expanding tests or documentation, and creating or updating pull requests. The practical difference is that a developer doesn’t have to manually copy a settled requirement into a separate coding environment and reconstruct why the team made each choice. The agent can receive the conversation where the requirement emerged and turn it into work that enters the normal review process.
Existing GitHub permissions, repository policies, branch protections and required reviews continue to apply. That’s an important boundary: a chat request doesn’t grant new repository authority or bypass a protected branch. It also means usefulness depends on the organisation’s current GitHub entitlements, app installation and access configuration.
Repository permission and conversation sensitivity are separate questions. A user may be authorised to start work against a repository while the surrounding chat contains customer details, security discussion or internal material that wasn’t written with an agent in mind. Teams need to decide which conversations are appropriate inputs, not rely on source-control permissions to answer that privacy question for them.
The other limitation is maturity. This is a public preview, not a generally available release, and Microsoft hasn’t given a final licensing model or general-availability date. Behaviour and access conditions may still change before a full release.
Used carefully, the integration can reduce context loss between product discussion and a reviewable pull request. It doesn’t remove engineering judgement: the resulting code still needs tests, human review and the same merge controls as work started anywhere else.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion
- storagereview.com/news/nvidia-groq-3-lpx-enters-full-production-3400-tokens-per-second-at-100k-context-256-lp30s-per-rack
- thomsonreuters.com/en-us/posts/innovation/how-we-built-thomson
- thomsonreuters.com/en-us/posts/innovation/the-future-of-ai-is-knowing-how-to-use-the-intelligence-available-to-you
- techcommunity.microsoft.com/blog/microsoftteamsblog/turn-conversations-into-code-with-github-copilot-in-microsoft-teams/4548305
- minister.infrastructure.gov.au/wells/media-release/support-australian-journalism-passes-parliament
- aph.gov.au/Parliamentary_Business/Bills_Legislation/bd/bd2627/27bd011
- biz.chosun.com/en/en-it/2026/08/20/7B5OSUVJI5HM7OVEYI6TLEC7LM
- investor.tsmc.com/sites/ir/shareholders-meeting/2026-06-04/2026AGM_Agenda_wmn.pdf
- home.treasury.gov/news/press-releases/sb0615
- googlecloudpresscorner.com/2026-08-24-Google-Cloud-Announces-Strategic-Partnership-with-Verizon-to-Scale-Enterprise-AI
- cisa.gov/news-events/alerts/2026/08/21/cisa-adds-one-known-exploited-vulnerability-catalog
- blog.zimbra.com/2026/07/patch-release-update-zimbra-10-1-20