AI & Tech Daily
NVIDIA Opens the Router as AI’s Power and Oversight Costs Mount
NVIDIA releases an open specialist model and a routing library designed to divide agent workflows across multiple AI models. Jesse examines why orchestration, evaluation and cost control are becoming as important as model selection. Also: Google makes visible AI watermarks optional while opening local provenance checks; a forecast tests the economics of gas-powered data centres; US courts prepare to count spyware-enabled wiretaps; Flock adds controls to licence-plate surveillance; autonomous trucks begin supervised California highway testing; Apple proposes commissions on purchases through external links; and Gemini 3.7 Flash reaches GitHub Copilot.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
The Multi-Model Agent Stack
The next big saving in AI may come from giving most of the work to a smaller model — then calling the expensive one only when the job genuinely demands it.
That’s why NVIDIA’s latest release is worth spending some time on. On August 11, the company released Nemotron 3.5 Lightning, a thirty-billion-parameter mixture-of-experts model aimed at high-volume specialist work. Alongside it came NeMo Switchyard, an open-source library that routes individual steps in an AI agent’s workflow across different models.
A mixture-of-experts model activates only selected parts of its network for a given request. That can reduce the computation needed for each token compared with running every parameter every time. Nemotron 3.5 Lightning can be obtained through Hugging Face, ModelScope, OpenRouter, NVIDIA’s build services and cloud partners. It’s designed to run locally, on premises, at the edge or in cloud infrastructure.
The more consequential piece may be Switchyard. Instead of choosing one model for an entire application, a developer can create a workflow in which a quick classification goes to a small open model, a difficult reasoning step goes to a frontier model, and a specialised task goes somewhere else again. NVIDIA says the router can be tuned around quality, latency or cost without requiring the calling application to be rewritten.
That opens a more useful question than which single model ranks first on a benchmark. Which model is good enough for each particular step, and how confidently can the system recognise when it needs something stronger?
For working AI teams, the immediate opportunity is fairly concrete. They can test Nemotron as a specialist and Switchyard as the control layer, while reserving expensive models for the parts of a task where their added capability earns its price. Local and on-premises deployment also gives organisations more options when data handling, network latency or infrastructure ownership matters.
There are important limits. The performance and cost numbers in NVIDIA’s announcement come from NVIDIA or its partners, so they haven’t been independently verified. Real savings will vary with the workload, hardware, model mix, routing decisions and the accuracy loss an organisation can tolerate. A router that sends a hard task to a cheap but inadequate model can save money while quietly making the product worse.
It also creates a new operational surface. Teams need evaluations for individual models, evaluations for the routing policy, traces showing which path a request took, and fallback behaviour when a provider is slow or unavailable. Updating one model can change the balance of the whole system. Cost accounting becomes more detailed, and debugging becomes less obvious because two apparently similar requests may travel through different paths.
My read is that agent development is moving from model selection towards systems engineering. The valuable skill won’t only be knowing which model is strongest. It’ll be designing a measured portfolio of models and proving that the router makes sensible decisions. If this approach works, capable agents may become cheaper and easier to deploy across mixed infrastructure. But the complexity hasn’t disappeared; it has moved into orchestration, evaluation and monitoring.
Watermarks Move Below the Surface
That shift towards invisible infrastructure has a close cousin in AI media, where provenance is becoming less visible to people and more dependent on software.
Google began rolling out a setting on August 14 that allows users to remove visible watermarks from supported images, videos and songs generated through Gemini and Flow. The affected output includes material from Nano Banana, Omni and Lyria. Google says invisible SynthID marks and C2PA metadata remain in use, and support in Search is described as coming soon.
The obvious benefit is cleaner output. A creator can use generated media in a professional layout without a visible mark competing with the design. The trade-off is that an ordinary viewer loses an immediate clue about where the material came from. Detecting its provenance increasingly depends on whether the viewing platform preserves and inspects machine-readable signals.
Google is also opening part of that inspection layer. It released Credentio on August 13, an open-source C++ library that validates C2PA Content Credentials locally. C2PA is a technical standard for attaching signed information about a media file’s origin and editing history. Credentio supports versions 2.2 and 2.4 of the specification, and it can perform the check without uploading the file to Google or another validation service.
That local design is useful for privacy-sensitive applications, internal media systems and processing pipelines that can’t send every asset to a third party. A newsroom, creative platform or enterprise archive could incorporate validation into its own infrastructure. The present limitation is important, though: Credentio validates credentials but doesn’t generate or embed them.
There’s also no independent evidence in these sources showing how reliably SynthID or C2PA information survives common edits, screenshots or platform recompression. The user setting is still rolling out as well.
For developers, local verification is a meaningful addition because it makes provenance checks easier to integrate without creating another external data transfer. For the public, the picture is more mixed. My assessment is that cleaner creative output is reasonable, but removing the obvious signal places more responsibility on platforms and tools to expose the remaining evidence. Provenance that exists but is never checked offers much less transparency than its technical presence suggests.
The Gas Bet Behind AI Power
Now to the physical cost of all that computation, because a shortcut around grid delays can bring a very different kind of exposure.
Energy research firm Noreva has produced a long-range forecast warning that growing data-centre demand, slower gas-supply growth and expanding liquefied natural gas exports could push natural-gas prices above ten US dollars per million British thermal units at some American hubs. TechCrunch reported the forecast on August 14.
That figure isn’t a current price. The reported range at relevant hubs is roughly two to four dollars fifty per million BTUs, and current futures markets aren’t predicting a comparable surge in the near term. Noreva’s number is a scenario based on assumptions about supply, pipelines, exports and demand. Its timing and regional severity remain uncertain.
Even so, it lands at an awkward moment for AI infrastructure planning. Amazon, Google, Meta and Microsoft have announced gigawatt-scale natural-gas power projects associated with data-centre expansion. Generating power close to a facility can help a company move ahead when a grid connection is delayed or insufficient. It can also give an operator more control over how capacity is built.
But the arrangement replaces one bottleneck with another risk. If fuel prices rise sharply, the cost of electricity from those projects rises with them. Operators may absorb that expense, pass it through to customers, hedge against it, or lean more heavily on public grids. Greater data-centre demand can also interact with local energy markets, raising questions about who bears the cost when industrial and household users depend on the same fuel and infrastructure.
The useful takeaway for hyperscalers and investors isn’t that ten-dollar gas is inevitable. It’s that privately arranged generation doesn’t make energy economics predictable. A plan that looks attractive under today’s fuel prices can become much less compelling under a plausible stress case.
In my view, this forecast turns an abstract concern into a scenario worth putting through the financial model. AI companies assessing gas projects now have a reason to test higher fuel costs, regional pipeline constraints, hedging expenses and the possibility of returning to the grid at difficult moments. Bringing generation in-house may ease a connection problem, but it also brings commodity exposure and community energy consequences much closer to the centre of AI strategy.
A Delayed Count of Government Spyware
The next development is a transparency gain, although the timetable shows just how slowly oversight can follow surveillance technology.
The Administrative Office of the U.S. Courts confirmed on August 14 that it plans to add a spyware and hacking category to the judiciary’s 2028 Wiretap Report. That report will be published in 2029. The category is intended to count court-authorised uses of hacking tools that intercept communications in real time.
At present, there’s no public recurring baseline showing how often American authorities use those tools for wiretapping. A dedicated category should give lawmakers, researchers and the public their first official count, allowing changes over time to be tracked instead of inferred from scattered cases.
The scope is narrower than the phrase government spyware might suggest. The count will cover real-time interception authorised under wiretap procedures. It won’t include remote searches of devices used to extract stored files, photographs or location information. A hacking operation could therefore be highly intrusive and still sit outside this particular category if its purpose is retrieving stored data rather than listening to live communications.
The judiciary says its forms and reporting procedures need to be updated before collection can begin. The final presentation hasn’t been settled, and it’s unclear how many deployments will fall inside the new definition rather than the excluded device-search category.
For the public, the benefit is a recurring official measure where none currently exists. The weakness is that disclosure will arrive well after the surveillance occurs, with the first figures not published until 2029.
My assessment is that the new category improves accountability, but only within the boundary it draws. A count can reveal frequency and trends; it can’t by itself establish whether particular tools were proportionate, secure or properly controlled. The long delay and exclusion of stored-data searches also mean lawmakers will still be working with an incomplete view. Oversight is gaining a useful instrument, just several years behind the capability it is trying to observe.
Flock Adds Surveillance Guardrails
Staying with oversight, a private surveillance vendor is tightening its own controls — and asking the public to trust measures whose effectiveness hasn’t yet been demonstrated.
Flock announced on August 13 that automated licence-plate data will have a seven-day default and recommended retention period. The company is also introducing offence-based sharing controls, mandatory case codes and mandatory adoption of its Audit Assistance misuse-detection tool for law-enforcement customers.
Shorter default retention reduces the amount of historical location information kept as a matter of routine. Offence-based sharing is intended to limit searches according to the type of investigation, while case codes create a record connecting a search with an identified matter. Flock says Evidence Mode, which will allow selected investigation data to be preserved, is expected in the coming weeks.
The qualifications matter. Existing customers keep the retention periods that were previously approved, so seven days won’t instantly become a universal limit. Local agencies will continue to make important choices about retention. Case codes and Audit Assistance are due to become mandatory for law-enforcement customers by the end of 2026 rather than immediately.
Audit Assistance looks for activity matching defined patterns of abnormal behaviour, and automatic suspension can be applied when those criteria are triggered. Flock says the system doesn’t use machine learning or AI. Pattern-based detection can be perfectly legitimate, but the label tells us little about performance. Flock hasn’t published false-positive or false-negative rates, nor results from an independent audit showing how reliably the tool detects misuse.
There’s also an accountability question in what happens after an alert. Local agencies initially investigate activity inside their own organisations, while the vendor operates the system that raised the concern. That may identify obvious misuse, but it isn’t equivalent to independent scrutiny.
For police customers, these changes add real friction around access and create better records for later review. For people whose movements appear in licence-plate databases, shorter retention and stronger auditing can reduce some privacy risk.
I’d treat the announcement as progress, not proof. Controls such as case codes, limited sharing and shorter defaults are concrete improvements when they’re consistently enforced. The harder question is whether Audit Assistance catches abuse without flooding agencies with bad alerts or missing sophisticated misuse. Independent testing and published performance would make that judgement possible. Until then, the guardrail exists, but its strength remains an unverified vendor claim.
Autonomous Trucks Reach California Highways
A smaller but tangible shift is happening on California roads, where autonomous trucking has moved from policy debate into supervised testing.
Aurora Innovation and Kodiak AI received permits to test autonomous heavy-duty trucks on public roads in California. Kodiak began operating a small test fleet around Mountain View following the permits reported on August 14.
California’s updated rules created a testing path for autonomous vehicles weighing more than ten thousand pounds. These permits don’t authorise driverless commercial freight service. A human safety operator is required, and the trucks are prohibited from operating on most roads with posted speed limits of twenty-five miles per hour or lower.
That restricted format gives developers access to roads near their California engineering teams. Until now, companies doing this work have concentrated much of their road testing in states such as Texas. Testing closer to engineering bases can make iteration and technical investigation easier, while exposing the systems to a different regulatory and road environment.
The political dispute is far from settled. Teamsters California has sued the state, arguing that officials didn’t adequately consider safety and employment effects when developing the rules. That legal challenge remains unresolved.
For autonomous-truck companies, the permits lower a practical barrier to testing, but they don’t establish that the technology is ready for broad unsupervised deployment. The fleets remain small, geographically constrained and monitored by people in the cab.
The most useful consequence, in my view, is that some claims can now be tested against observable California operations. Regulators and the public can examine safety events, disengagements and operating limits instead of relying entirely on projections. Labour effects will take longer to assess because supervised trials don’t reproduce a commercial driverless network. This is an opening for evidence gathering, not a verdict on the final shape of trucking.
Apple’s External-Payment Toll
For app developers, the route around Apple’s checkout may still come with an Apple charge attached.
On August 13, Apple filed a proposed US commission structure that would charge standard apps fifteen per cent on purchases made after a user follows an external payment link from an iOS app. Small-business developers would face a proposed five per cent rate, while specified partner programs would carry a ten per cent rate. Subscription renewals would also attract a proposed ten per cent commission.
These aren’t final rules. The filing followed the US Supreme Court’s refusal to pause lower-court proceedings in Apple’s litigation with Epic Games, and the court hasn’t approved Apple’s proposal. The rates and implementation details may change.
The proposal nevertheless shows Apple’s preferred interpretation of external-payment freedom. Developers could direct users to another checkout system, but qualifying transactions would still generate a platform fee. That could reduce the commission compared with some App Store arrangements while preserving Apple’s claim on economic activity originating inside an iOS app.
For a developer deciding whether to build an external flow, the calculation would include more than the headline rate. Payment processing, customer support and conversion friction sit alongside Apple’s proposed commission. A lower percentage doesn’t automatically make the alternative route cheaper or better for users.
My read is that the central competition dispute remains intact in a modified form. External links can change who operates the checkout and may give developers more control over customer relationships, yet they don’t necessarily eliminate the platform toll. Until the court responds, developers can model the proposed rates, but they can’t safely treat them as settled economics.
What Changes for You
One practical option is arriving inside a tool many developers already use, although access and cost controls differ by account.
GitHub began adding Google’s Gemini 3.7 Flash to the Copilot model picker on August 13. The gradual rollout covers Copilot Pro, Pro Plus, Max, Business and Enterprise plans. Supported surfaces include Visual Studio Code, Visual Studio, JetBrains, Xcode and Eclipse, along with Copilot CLI, the Copilot cloud agent and the Copilot app.
For individual developers on an eligible plan, the change means another fast coding-model choice can appear without adopting a new editor, command-line workflow or agent service. Teams can compare it with their existing Copilot models on actual repository tasks rather than moving code into a separate product.
Business and Enterprise accounts have an extra gate: administrators need to enable the Gemini 3.7 Flash Preview policy. Availability is still rolling out, so an eligible account may not see the model immediately. Usage is charged at the provider’s list pricing through GitHub’s usage-based billing, which makes cost tracking relevant when developers or cloud agents can switch among a growing set of models.
GitHub says its early testing found improved coding and agentic performance, but those qualitative claims haven’t been independently substantiated in the announcement. The sensible comparison is therefore local: measure completion quality, review burden, latency and total spend on the work your team actually performs.
The useful change is choice without workflow disruption. The limitation is that a larger model menu can make spending and evaluation harder to govern. For organisations, the administrative preview policy provides a controlled way to test before broad adoption. My judgement is that model selection inside coding tools is becoming ordinary infrastructure; the advantage will come from disciplined evaluation, not from enabling every new option by default.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx
- developers.googleblog.com/introducing-credentio-open-source-c-library-for-c2pa-content-credentials-from-google
- techcrunch.com/2026/08/14/google-will-now-allow-users-to-remove-visible-watermark-from-its-ai-generations
- github.blog/changelog/2026-08-13-gemini-3-7-flash-is-now-available-in-github-copilot
- techcrunch.com/2026/08/14/hyperscalers-might-regret-embracing-natural-gas-if-new-forecast-proves-correct
- techcrunch.com/2026/08/14/us-courts-will-start-publishing-how-often-the-government-uses-spyware
- flocksafety.com/blog/flock-guardrails-address-lpr-privacy-concerns-and-police-transparency
- techcrunch.com/2026/08/13/flock-says-its-new-tool-will-help-identify-police-abuse-but-hasnt-explained-how-it-works
- techcrunch.com/2026/08/14/self-driving-trucks-are-officially-testing-on-california-highways
- techcrunch.com/2026/08/14/apple-proposes-to-take-a-15-cut-of-purchases-made-outside-the-app-store