All episodes

AI & Tech Daily

AI Agents Move Into the Infrastructure Layer

18:07

NVIDIA is pushing AI-agent containment beyond model behaviour and into runtime and hardware controls, with an open runtime boundary and a BlueField-based reference design for independent enforcement. We also cover two actively exploited NetScaler flaws, the Pentagon’s upheld exclusion of Anthropic from relevant procurement, Cloudflare’s agent-oriented CLI, DigitalOcean’s managed agent platform, Meta’s enterprise AI push and a new real-user web performance dataset. In What Changes for You, Claude Sonnet 5.5 arrives across major cloud APIs and paid GitHub Copilot plans, with rollout, billing and migration details developers need to check.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

Agent Safety Leaves the Model

An AI agent with permission to run code can do real damage before its own safety logic notices. NVIDIA wants the boundary around that agent to live somewhere the agent can’t rewrite or ignore.

Our main story today looks at NVIDIA’s Open Agent Safety Platform, and what changes when agent containment moves out of the model and orchestration code into the runtime and, eventually, dedicated hardware. The launch combines OpenShell, a broadly available secure runtime, with a Sentry reference design built around NVIDIA’s BlueField-4 data-processing units.

OpenShell is the part developers can begin examining now. It creates a runtime boundary around an agent, traces what the agent does and enforces defined policies. NVIDIA says it can run across NVIDIA Vera, Arm and Intel CPUs, so the basic software story isn’t confined to one processor family. That matters for organisations already testing agents on mixed infrastructure: the control layer can, in principle, sit around the workload rather than requiring the model to make every safety decision correctly.

The separation also changes how a failure can be handled. A malicious instruction or an incorrect tool result may still steer the agent towards an unsafe action. An external runtime policy has another chance to refuse that action and leave a trace of what the agent attempted. That doesn’t tell us whether the policy will be easy to write or precise enough for messy production work. It does give security engineers a control point outside the chain of prompts, model responses and tool calls they’re trying to supervise.

Sentry goes further. The reference design puts monitoring and enforcement on a BlueField-4 DPU, a separate processor commonly used for infrastructure work such as networking and security. The important architectural choice is independence. If the agent process is compromised, behaves unexpectedly or simply follows a bad instruction, the enforcement component isn’t sitting inside the same software boundary. NVIDIA says Sentry can detect and quarantine an agent that crosses a defined boundary within milliseconds. That is the company’s claim, though, and there’s no independent production validation in the material released so far.

There are also two different availability stories here. OpenShell and initial developer resources are available. Several associated features and products in the wider announcement come with when-and-if-available caveats, and NVIDIA hasn’t set out firm production timing for all of them. A design can be technically compelling and still need a lot of operational proof: false positives, policy management, visibility during an incident, performance overhead and the behaviour of the quarantine path all matter once an agent is doing useful work. The complete Sentry approach also means buying into specialised NVIDIA infrastructure.

My read is that the architectural direction is more important than the millisecond figure. As organisations give agents credentials, shells and permission to change systems, asking the agent to police itself becomes a weak control. A separately enforced runtime boundary gives security teams somewhere to set policy and observe behaviour even when the model or harness gets it wrong. Organisations can test that idea with OpenShell now, but Sentry should be treated as an emerging design that still needs availability details and independent evidence before it becomes a production control.

NetScaler Flaws Under Active Attack

That longer-term control problem can wait for a test environment. Two NetScaler flaws need a much faster response.

Citrix disclosed eight vulnerabilities in NetScaler ADC and Gateway products on 27 September, and confirmed that attackers had already exploited two critical unauthenticated remote-code-execution flaws on appliances that had not been mitigated. The two are CVE-2026-88771 and CVE-2026-88772, and both carry a CVSS 4.0 score of 9.5.

The first affects NetScaler ADC and Gateway deployments in their default configuration. The second can lead to remote code execution or denial of service when DTLS is enabled, including the default configuration for a VPN virtual server. These are internet-facing systems that often sit directly in the path of remote access and identity traffic, so unauthenticated code execution gives an attacker a valuable position before a user ever reaches an internal application.

Citrix has directed customers to fixed releases, including 14.1-73.37 and 13.1-64.23 or later. Its managed cloud services have already been patched, which narrows the immediate operational burden to customer-managed appliances. The public reporting doesn’t establish how many victims there are, who is behind the attacks or how broad the campaign has become. Confirmed exploitation is enough to change the response, even without those answers.

For an operator, installing the fixed release closes the known vulnerable path from that point onward. It doesn’t prove nobody used the flaw beforehand. The sensible incident boundary therefore includes both the update and a review for signs of earlier compromise, particularly on exposed gateways that were reachable while the vulnerable build was running. That’s my practical takeaway: for high-value remote-access infrastructure, patching and compromise assessment are two parts of the same job. Treating the successful upgrade as the end of the incident leaves the most important historical question unanswered.

When Model Rules Meet Defence Procurement

Security boundaries are also being drawn through contracts, and a US court has just reinforced one of them.

A divided United States Court of Appeals for the District of Columbia Circuit rejected Anthropic’s challenge to the Pentagon’s supply-chain exclusion on 25 September. The immediate result is that Anthropic remains barred from the relevant military procurement while separate litigation over a broader government response continues.

The dispute followed Anthropic’s restrictions on using its models for lethal autonomous weapons and mass domestic surveillance. In a two-to-one decision, the majority held that the department had adequate grounds to treat the company as a national-security supply-chain risk. The judgment is specifically about the Pentagon procurement decision. It doesn’t establish a general rule for every government agency or every use of AI.

That distinction matters because a separate federal ruling in California blocked a broader government-wide designation and restrictions affecting contractors. The two proceedings leave a split practical position: the defence procurement exclusion stands, while the reach of a wider response remains contested. Reuters reported that Anthropic was considering seeking further review, so this legal path may not be finished.

For AI suppliers, model-use policies are becoming part of product compatibility. A company can set restrictions for reasons it considers important, but a defence customer may decide those limits conflict with what it intends to procure. That choice then flows through prime contractors, integration partners and any project that depends on the excluded supplier. Agencies and contractors, in turn, have to work with uncertainty while the broader legal boundaries are still being tested.

My read is that procurement teams can no longer treat a model provider’s safety policy as a page for the legal file after the technical evaluation. It can determine whether the product is eligible for the work at all. The court has upheld one concrete exclusion, not settled the wider argument, and suppliers selling into defence will need to examine permitted-use terms alongside price, capability and deployment architecture.

Cloudflare Builds a CLI for Agents

At a developer’s terminal, the same shift towards capable agents is changing the shape of familiar tools.

Cloudflare has released an open-source command-line interface called cf in open beta. It’s designed for developers and coding agents, with structured access to more than 3,000 Cloudflare API operations and JSON output by default. Cloudflare contrasts that with roughly 280 command paths in Wrangler, its established Workers CLI.

The wider coverage changes what an agent can discover and operate without a developer first translating every task into a known command. The CLI includes natural-language command discovery, while the machine-readable output gives a coding agent something more dependable to inspect than prose formatted for a person. Workers configuration moves to a TypeScript file named cloudflare.config.ts, another sign that Cloudflare sees this as more than a thin command wrapper.

The beta is global, but it isn’t a complete replacement for Wrangler yet. Some JavaScript deployments that rely on esbuild, along with Rust and Python Workers, still delegate parts of the work to Wrangler. Cloudflare says Wrangler will be maintained for 18 months after the cf beta ends, but it hasn’t said when the beta will finish or when those delegated paths will become native. Anyone adopting it now is choosing a migration period, not a finished handover.

There’s a useful security consequence hidden inside the convenience. An interface covering thousands of API operations gives an agent a much larger action surface. That can remove tedious integration work, but a broadly privileged token paired with a mistaken generated command can also affect far more than a single deployment. My view is that cf is worth testing where agent-driven administration has a real payoff, with credentials scoped to the exact account, service and task. Beta compatibility is one limit; permission design is the more enduring one. A better agent interface doesn’t reduce the need to decide what that agent is allowed to touch.

Managed Agent Infrastructure Arrives

If managing the harness itself isn’t where you want to spend engineering time, DigitalOcean is offering to take on more of it.

DigitalOcean launched Managed Agents in public preview on 22 September. The service combines isolated agent sessions, model inference and managed tool access in one cloud platform. It supports familiar harnesses including Claude Code, Codex CLI, OpenCode, Hermes and LangGraph, as well as custom OCI containers.

DigitalOcean says each session runs in an isolated harness environment and can retain state while paused. Compute is charged for active CPU time, which is a useful model for agents that spend part of their life waiting for instructions or external systems. The advertised catalogue is large: more than 75 hosted models and over 16,000 tools from more than 500 providers through the Model Context Protocol, or MCP. MCP is the common interface many agent systems use to discover and call tools.

Putting execution, inference and tool brokering behind one managed service removes a fair amount of plumbing. A developer doesn’t have to assemble a persistent runtime, connect multiple model providers and host every tool gateway before an agent can complete a useful task. Paused state also opens the door to work that lasts longer than one interactive session without paying continuously for an active CPU.

Those benefits move the trust boundary rather than making it disappear. The provider becomes part of the path for code execution, credentials, model access and third-party tools. The very broad MCP catalogue also means developers need to understand which tools an agent can call and what authority each connection carries. And because this is a public preview, the claims about performance, security and total cost come from DigitalOcean; production reliability and real workload economics haven’t yet been independently established.

For developers, my assessment is that Managed Agents could be a useful way to validate persistent agent workflows without building the whole control plane first. The trade-off is deeper lock-in around session state, permissions and model access. Before using it for sensitive or durable work, the preview needs to prove that its isolation and cost model hold up under the workload you actually run.

Meta Sets Up an Enterprise AI Business

The platform contest is drawing in another very large supplier, although this announcement is still mostly a statement of intent.

Meta has created a new Enterprise Platform business to package its AI stack into products and services for organisations. It appointed CJ Desai as Chief Enterprise Platform Officer and named four initial areas: Muse, Meta Business Agent, Muse API and Muse Code. That product list points towards agents, developer access and business-facing services under one enterprise operation.

What Meta hasn’t supplied is just as useful for judging the news. There’s no pricing, no firm release timetable, no detailed deployment architecture and no general customer-availability information. The claims about platform capability and business value come from Meta, and there isn’t yet a generally available product against which buyers can test them.

The signal is still meaningful. Meta wants to compete for enterprise AI spending with a broader platform offer rather than relying only on models or consumer products. But an enterprise buyer can’t compare security boundaries, integration effort, support or total cost from a list of names. My takeaway is to treat this as a new prospective supplier entering the procurement map, not a migration trigger. The decision-grade information will be release dates, technical boundaries, contractual terms and evidence from running services. Until those arrive, the size of the ambition tells us more than the readiness of the products.

A Daily View of Real Web Performance

Here’s a more concrete release for developers who care about what a website feels like outside the lab.

Cloudflare has published BEACON, a public dataset built from billions of anonymised real-user performance measurements across 10,000 large websites on its network. The data updates daily through BigQuery and includes distributions for Core Web Vitals, along with component measurements broken down by geography, browser engine and device type.

Core Web Vitals capture loading performance, responsiveness and visual stability from the user’s point of view. Lab tools are useful because they produce repeatable tests, but field data shows what happened across real devices and networks. BEACON gives developers and researchers a large sample for comparing those two views, or for examining whether an apparent performance problem is concentrated in a region, browser engine or class of device.

Cloudflare says it removes domain names and URL paths, aggregates the records and discards cells with fewer than five observations. Those controls reduce the detail available about any one site or visit, while preserving broader distributions for analysis.

The limit is the sample. These are the 10,000 largest sites on Cloudflare’s network, not a census of the web. Smaller websites and sites served elsewhere may behave differently, so a result from BEACON shouldn’t be presented as the experience of every user online. My view is that the dataset is most useful as a strong external baseline alongside a site’s own real-user monitoring, not as a replacement for it. It can show whether an issue resembles a wider pattern, but the measurements from your own users still answer the question you ultimately care about.

What Changes for You

One model release is immediately usable in tools many developers already have open.

Anthropic released Claude Sonnet 5.5 on 28 September through its own platform, Amazon Web Services, Google Cloud and Microsoft Azure. The API identifier is claude-sonnet-5-5. GitHub has also made it generally available across paid Copilot Pro, Pro+, Max, Business and Enterprise plans, with access rolling out through supported editors, the CLI, coding agent, mobile app and website. Administrators can disable the model, so an eligible plan doesn’t guarantee that it’s enabled in a managed organisation.

Anthropic says Sonnet 5.5 completes work more than 30 per cent faster and can cost up to 30 per cent less per task than Sonnet 5. Those are Anthropic’s measurements, and the list API prices haven’t fallen: they remain US$2 per million input tokens and US$10 per million output tokens. Any saving therefore depends on the model finishing a real task with fewer tokens, less tool time or fewer retries. Copilot usage is billed at provider list prices, so checking the billing path matters before making it a default across a development group.

There are migration details as well. Anthropic says some high-risk cyber requests can visibly fall back to Sonnet 5, some biological requests may be blocked incorrectly, and applications that use thinking-off mode may need a new between_tools setting. Results will vary with the workload and with safeguard intervention.

For a developer, the useful change is broad access rather than another benchmark claim. You can compare the model in an existing API or Copilot workflow now. Check whether the rollout has reached your surface, whether an administrator permits it, and whether tool-calling behaviour still matches your application before swapping an established production model.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. investor.nvidia.com/news/press-release-details/2026/NVIDIA-Launches-Open-Agent-Safety-Platform-to-Secure-Agents-From-Testing-to-Deployment/default.aspx
  2. support.citrix.com/external/article/CTX697096
  3. bleepingcomputer.com/news/security/citrix-admins-warned-to-shut-down-netscalers-over-2-exploited-zero-days
  4. media.cadc.uscourts.gov/opinions/docs/2026/09/26-1049-2194984.pdf
  5. investing.com/news/economy-news/us-appeals-court-declines-to-block-pentagons-blacklisting-of-anthropic-4917802
  6. blog.cloudflare.com/cloudflare-cf-cli-launch
  7. investors.digitalocean.com/news/news-details/2026/DigitalOcean-Launches-Managed-Agents-Bringing-Agent-Execution-Tool-Access-and-Inference-Together-on-One-Cloud/default.aspx
  8. about.fb.com/news/2026/09/launching-meta-enterprise-platform
  9. blog.cloudflare.com/how-fast-is-the-web
  10. anthropic.com/claude-sonnet-5-5
  11. github.blog/changelog/2026-09-28-claude-sonnet-5-5-in-github-copilot