Back to the show

AI & Tech Daily

AI’s New Contest: Access, Agents and Custom Chips

19:50

Anthropic is opening Claude Mythos 5 to enterprise repository scanning while keeping the underlying cyber model behind a restricted interface—a useful test of how frontier access may be governed. Jesse also examines an exploited maximum-severity Entra ID flaw, Alibaba’s costly AI-cloud expansion, Waymo’s first purpose-built 5-nanometre driving chip, OpenAI’s national-security oversight initiative, urgent Cisco fixes, and Supermicro’s export-compliance investigation. In What Changes for You: Google brings Antigravity’s agentic development tools into eligible Gemini Enterprise plans, with central controls and substantial compliance limits.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

Claude’s Controlled Cybersecurity Opening

Anthropic is letting more defenders use its strongest cyber model—but only through a narrow security service where every proposed patch still needs a human yes.

That restriction is the development worth sitting with. On 21 August, Anthropic put Claude Mythos 5 into the public-beta Claude Security service for repository scans. Claude Enterprise customers can use it under standard token billing, with no extra Claude Security fee. Yet they don’t receive general direct access to Mythos 5. Anthropic controls the interface, the task and the form of the output.

For a security team, the workflow is fairly concrete. The service inspects a code repository and reports suspected vulnerabilities, their Common Weakness Enumeration classification, confidence and severity assessments, and suggested fixes. A person must review and approve any patch. That human gate is important because the output can affect production code, and because Anthropic hasn’t published independent evidence showing how effective Mythos 5 is at finding real flaws or how reliably its safeguards constrain the model. A confident-looking finding is still a lead to validate, not proof.

The structured output should make triage easier: a team gets a suspected weakness, an assessment of urgency and a proposed remediation in one place. It also creates several points where human judgment still matters. Engineers need to confirm that a reported flaw is reachable in the real application, that its severity fits the deployment, and that the suggested patch doesn’t break expected behaviour. Anthropic’s approval gate prevents an automatic code change; it can’t make those judgments on the customer’s behalf.

Anthropic says access will widen through its Cyber Verification Program and partner cyber-defence tools over the coming weeks. That means this isn’t universal availability today. The company also announced a US$35 million Defender Advantage Fund, initially offering model credits to open-source security projects, although it hasn’t named the first recipients. There’s useful potential there: open-source maintainers often protect widely used software without having enterprise security budgets. But until recipients and results are visible, the fund is a commitment rather than an outcome.

The bigger idea is controlled capability. Model releases are often framed as a choice between broad access and complete refusal. Anthropic is testing a third route: let approved users perform a bounded, defensible task, expose structured findings, and withhold the underlying model. Partner integrations could extend that pattern without giving every operator the same freedom they’d have through general model access.

My read for organisations is that access to frontier systems may increasingly arrive as permissioned workflows, not a model endpoint you can use however you like. That can make advanced defence tools available sooner, and it can make oversight easier, but it also concentrates judgment in the provider. Anthropic decides which tasks are allowed, which outputs are shown and who qualifies. Security teams gain capability while accepting dependency on those controls.

So the opportunity is real but bounded. Enterprises can put a stronger scanner into repository review without a separate service fee, then compare its findings with existing tools and human analysis. They shouldn’t treat the model’s name as independent validation, and they remain accountable for every fix they approve. The access gate may reduce some misuse risk; it doesn’t transfer responsibility for the code.

Microsoft’s Entra ID Blind Spot

That controlled-access question is one side of security. The other is what customers can see when the provider itself gets hit.

Microsoft disclosed CVE-2026-69836 on 20 August, an unsafe-deserialization flaw in Entra ID that allowed unauthenticated remote code execution. It carries a CVSS score of 10.0, the top of the scale. Microsoft also says attackers exploited it before disclosure, with no privileges or user interaction required.

Unsafe deserialization means a system turns supplied data back into an object without safely constraining what that object can cause the software to do. In the worst case, as here, a remote attacker can make the service execute code. That combination—no authentication, no click and remote execution—explains the maximum score.

Microsoft says its hosted Entra ID service is fully mitigated, so customers don’t have a patch to install. That removes an urgent maintenance job, but it doesn’t answer the operational questions. Microsoft hasn’t identified the attacker, the exploitation period, which tenants were affected or what the intrusions achieved. Without an exploitation window or affected-tenant count, customers can’t readily map the public notice to a defined period in their own records. They know the service flaw is closed; they don’t know whether their organisation was part of the attack.

For organisations using Entra ID, my practical judgment is that “no customer action required” should be read narrowly: there’s no software update to deploy. Security teams may still need to review identity and security telemetry for activity they can’t explain, preserve relevant records and watch for more detail from Microsoft. The exact scope of that review will depend on their environment because the public disclosure doesn’t yet define the campaign.

That distinction also matters for leadership. A patch status can be green while an incident question remains unresolved. Teams need a way to communicate that uncertainty without claiming compromise and without treating mitigation as proof that nothing happened.

Cloud identity shifts patching to the provider, which is valuable. It also makes customers dependent on that provider for evidence about a provider-side incident. Until Microsoft supplies more detail, remediation is complete at the service layer, but the impact assessment remains open.

Alibaba Pays for Vertical Integration

From a security gap with little disclosed impact, the scale changes completely: Alibaba has put numbers on the cost of building an AI stack.

For the June quarter, Alibaba reported 45 per cent year-on-year growth in AI cloud and compute services, taking revenue in that business to RMB48.437 billion, or about US$7.139 billion. The company says AI-related products achieved triple-digit annual growth for a twelfth straight quarter and reached 35 per cent of external cloud revenue. Those are company-reported figures, but they show AI demand becoming a meaningful part of Alibaba’s cloud business rather than a small experimental line.

Alibaba is also pushing further down into the hardware. It says its Zhenwu processors, including the M890, are commercially used by more than 650 external customers across over 20 industries. That matters in a market where access to leading Western accelerators can be scarce or restricted. A cloud operator with its own usable processors has more control over supply, system design and pricing choices.

There’s a large bill attached. Quarterly capital expenditure rose 75 per cent from a year earlier to RMB67.678 billion, while reported profit fell 75 per cent as AI investment expanded. Capital spending and profit move for many reasons, so those figures don’t prove that custom silicon will pay off. Nor do customer counts independently establish the performance or efficiency of the chips. They do show how much financial pressure accompanies the race to combine cloud services, models and processors.

For Alibaba Cloud customers, more domestic infrastructure could improve availability and give workloads an alternative hardware path. My caution is that the same integration can deepen lock-in: software, models and chips become parts of one supplier’s environment, making future migration harder.

For competing cloud providers, the threshold is rising. Model quality alone doesn’t settle the contest when a rival controls the data centres, accelerators and services around it. Alibaba is spending heavily to own more of that chain. The near-term result is faster cloud growth paired with a sharp hit to profit; the long-term efficiency benefit remains unverified.

Waymo Builds Its Own Driving Chip

A similar integration bet is showing up on the road, where latency and power use carry very physical consequences.

Waymo has disclosed its first purpose-built 5-nanometre ASIC for autonomous-driving workloads. An ASIC is a chip designed for a defined set of tasks rather than general-purpose computing. Waymo says this one processes data from cameras, lidar and radar, and runs the neural-network workloads inside its driving system. The company claims more than 1,000 TOPS—trillion operations per second—for front-end sensor processing and machine learning.

The disclosure also makes clear that “custom” doesn’t mean built alone. Waymo named AMD, Micron, NVIDIA, Samsung, Sandisk, Socionext and TSMC among its compute partners. Designing a specialised processor still depends on a broad network for manufacturing, memory, storage and other system components. Waymo may control more of the architecture while still relying on specialised suppliers to turn it into a working in-vehicle platform.

The potential payoff is tighter optimisation. Hardware shaped around known sensor and inference workloads can reduce latency and energy use compared with assembling everything from general-purpose parts. In a vehicle, faster processing and lower power demand can improve system efficiency. But Waymo didn’t disclose detailed power consumption, the number of vehicles using the ASIC or an independent performance comparison.

Those missing figures also make comparisons difficult. A throughput number is more informative when we know the power needed to reach it, the actual workloads used and the earlier system it replaces. Waymo hasn’t supplied that context, so the claim establishes intended scale rather than measured advantage over a competitor.

That keeps the 1,000-TOPS figure in its proper place. Throughput is a technical claim, not evidence of safer driving. Safety depends on the full system: sensors, software, training, validation, operational limits and how the vehicle handles failures.

My takeaway for autonomous-system builders is that custom silicon may become a competitive requirement at scale, but it raises the entry cost. Waymo can tune across hardware and software; smaller rivals may struggle to justify that design effort. The chip could make the system more efficient. The disclosure alone can’t tell us whether it makes the ride safer.

OpenAI Funds Oversight Tools

Control over an AI system is only half the issue. Someone also needs enough evidence to challenge how it was used.

OpenAI has launched a one-year, US$5 million initiative to help democratic oversight bodies examine government use of AI in national-security decisions. The programme, announced on 18 August, offers training, technical assistance and credits. Planned pilots would let authorised reviewers inspect records of AI-assisted decisions, including the inputs, outputs and tool use involved.

That kind of trace can be far more useful than a final answer in isolation. If an AI system contributed to a consequential government decision, reviewers need to know what information it received, what it produced and which tools it called. Otherwise, scrutiny can collapse into asking whether somebody liked the result. OpenAI says participating institutions would control the evidence and their findings, an important stated safeguard. It could also let a reviewer separate the model’s contribution from the actions of the human operator and any connected tool, rather than treating the whole process as one opaque AI decision.

There are large blanks. The company hasn’t named participating legislative, judicial or other oversight bodies, given a pilot timetable, or published independently defined measures for success. And this is a voluntary programme designed and funded by a major model supplier, not a binding accountability regime. That creates an obvious independence question even if the technical support proves useful.

My assessment for public institutions is that decision traces are newly worth demanding whenever AI affects national-security work. They can make review more specific and can expose where a model, operator or tool influenced an outcome. But access to logs is only a starting condition. Oversight also needs authority, relevant expertise and freedom to publish or act on findings. Those conditions can’t be supplied by model credits.

The initiative could help oversight bodies build that technical capacity. We can’t yet judge its reach because no participants or outcomes have been announced. The test will be whether reviewers gain durable, institution-controlled scrutiny, not simply training and credits supplied for one year.

Cisco’s Shared-Patching Lesson

Here’s a more immediate job for security teams: Cisco’s cloud upgrade doesn’t finish the patching at the customer edge.

Cisco released an August hardening update for Secure Workload covering five groups of internally discovered vulnerabilities. Two carry CVSS scores of 10.0. The flaws affect both software-as-a-service and on-premises deployments regardless of configuration, and Cisco says there are no workarounds. It also says it isn’t aware of malicious exploitation.

For on-premises operators, the fixed releases are 3.10.9.1 and 4.0.4.16. Cisco has upgraded its own SaaS clusters, but SaaS customers may still have work to do: relevant Agent and Connector components under customer control need updating too. That distinction is easy to miss when the central service is managed by the vendor.

The operational takeaway is simple and specific. Secure Workload teams should inventory the deployment, move on-premises systems to a fixed version, and verify every customer-managed Agent and Connector rather than treating the SaaS notice as proof that the whole environment is covered. With no workaround, delaying the update means retaining the exposure.

My broader read for organisations is that managed services divide maintenance; they don’t erase it. Cisco can secure its clusters, while software running at the edge of the service remains the customer’s responsibility. The lack of known exploitation is reassuring, but two maximum-severity scores and no alternative mitigation make verification more useful than reassurance.

Supermicro’s Export-Control Test

The next risk isn’t a software flaw. It’s whether controls around high-end hardware work before a shipment becomes a legal case.

Supermicro says it has completed a board-led investigation after two employees and a contractor were indicted over alleged export-control violations. The company reported further terminations across sales, technical support and business development for compliance and conduct failures, and says it has adopted recommendations to strengthen its export-compliance programme.

According to Supermicro, investigators found no evidence that current senior management knew of the alleged diversion, and nothing in the matters reviewed made its financial reporting unreliable. Those are significant findings, but they come from a company-led investigation. Government investigations and criminal proceedings continue, so the external case isn’t resolved. Supermicro also hasn’t identified the additional staff it dismissed.

For AI-server suppliers, distributors and customers, the consequence is more scrutiny of the route between an order and the final user. End-user screening, reseller controls and escalation processes have to work across sales and support, especially as restrictions on advanced accelerators tighten. A signed policy at headquarters won’t reveal much if staff can’t recognise or report a suspicious transaction.

My judgment is that a finding of no senior-management knowledge doesn’t demonstrate that the control system worked. Regulators and customers will want to understand how the alleged diversion was detected, where internal controls failed and whether the new measures prevent a repeat. Until the government proceedings develop, it would be wrong to treat the company investigation as the final word.

The commercial pressure cuts both ways: suppliers want to serve global demand, while buyers need confidence that orders won’t be delayed or questioned because an intermediary can’t establish the true end user. Stronger verification may make some sales slower, but weak verification now carries a much larger legal and reputational cost.

What Changes for You

For developers, one of these control-stack shifts is already landing inside familiar tools.

Google made Antigravity available on 20 August through eligible Gemini Enterprise subscriptions: Standard, Plus and Standard Emerging Market licences. Developers at qualifying organisations can use the agentic development tools from an IDE or the command line. A Visual Studio Code extension is available now; Visual Studio, JetBrains and Zed integrations remain previews.

What changes is less about a new chat window and more about governed agent work. Administrators can set project-level spending caps, control sandbox, browser and Model Context Protocol connections, and retain audit logs with prompts, responses and metadata. A developer can let an agent operate across more of a coding workflow, while the organisation sets boundaries around what it can reach, what it can spend and what gets recorded.

That makes Antigravity more practical for teams that couldn’t accept individually managed agents with inconsistent controls. My read is that central constraint is the useful product change: it can make agentic development easier to approve and harder to use invisibly. The retained logs also create a privacy and governance consideration because prompts, responses and metadata become organisational records.

There are serious limits. Access depends on an eligible licence and billing setup. Google says this Antigravity offering doesn’t support FedRAMP, ITAR, IL4 or IL5, SOC 1, 2 or 3, ISO 27001, or ISO 42001. That can block adoption outright for regulated or security-sensitive work. Several editor integrations and finer administrative controls are also unfinished, and Google hasn’t supplied independent productivity evidence.

So eligible developers can use the tools now, particularly in Visual Studio Code and the command line, but organisations should judge the agent by the work it completes under their policies—not by the promise of autonomy. For teams outside the listed compliance regimes, the controls may lower the barrier to a managed trial. For teams inside them, availability in the subscription doesn’t make the product ready for the workload.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. claude.com/blog/bringing-claude-mythos-5-to-more-defenders
  2. news.bloomberglaw.com/us-law-week/anthropic-makes-claude-mythos-5-available-for-security-scans
  3. msrc.microsoft.com/update-guide/vulnerability/CVE-2026-69836
  4. bleepingcomputer.com/news/microsoft/microsoft-warns-of-max-severity-entra-id-flaw-exploited-in-attacks/amp
  5. alibabagroup.com/en-US/document-2026456290057781248
  6. apnews.com/article/8a30302d23a96fc7b9aab664b9c1897d
  7. waymo.com/blog/2026/08/look-under-our-trunk
  8. openai.com/index/strengthening-democratic-oversight-in-national-security
  9. sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-hardening-csw1-shSvndWP
  10. ir.supermicro.com/news/news-details/2026/Supermicro-Announces-Completion-of-Independent-Investigation-and-Continued-Enhancement-of-Export-Compliance-Program/default.aspx
  11. assets.theregister.com/2026/08/21/20261
  12. cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers
  13. docs.cloud.google.com/gemini/enterprise/docs/ai-developer-tools-overview