Back to the show

AI & Tech Daily

Astra’s Cyber Warning and the New Limits of Frontier AI

18:42

OpenAI restricts work on its unreleased Astra model after preliminary testing raises the possibility of critical cyber capability. Anthropic adjusts Fable 5’s biology safeguards, AMD moves into specialised inference silicon, and Australia prepares national energy standards for AI data centres. Rippling confronts runaway token spending, an Amazon project receives an extraordinary emissions ceiling, and researchers uncover systemic weaknesses across Polish public websites. Plus, GitHub adds isolated Copilot worktrees, file rewind and enforceable enterprise MCP allowlists.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

Astra Hits a Cyber Threshold

OpenAI has paused some work on an unreleased model because it can’t yet rule out the possibility that the system could independently find zero-day vulnerabilities and carry out novel cyberattacks.

That’s the development worth sitting with, because the significant part isn’t a new benchmark or a dramatic demonstration. It’s that OpenAI has treated a preliminary capability warning as serious enough to restrict how its own researchers can use Astra.

On 7 August, the company said early internal evaluations could not rule out Astra reaching what it calls the Critical cybersecurity threshold. OpenAI defines that level in unusually concrete terms: independently developing zero-day exploits against many hardened systems, or executing a novel, end-to-end attack after receiving only a high-level goal.

A zero-day is a vulnerability for which defenders don’t yet have a fix. Finding one is difficult; turning it into a reliable attack against a hardened target adds several more layers of reasoning, experimentation and operational skill. OpenAI has not said Astra definitely possesses those abilities. The classification is preliminary, the model remains unreleased, and the underlying result hasn’t been independently reproduced. The company also says Astra had no involvement in the earlier Hugging Face incident.

Still, OpenAI has imposed stronger controls wherever Astra is being tested. Those include isolated test environments, restricted access to networks and tools, tighter protection for the model weights, sandboxing, and universal monitoring of Astra’s agentic applications. Activities that couldn’t meet the new standard were paused.

The distinction between an ordinary chatbot and an agent is important here. A chatbot can describe possible steps. An agent connected to tools can inspect systems, run code, adapt to results and continue towards a goal. If a model approaches critical cyber capability, the danger isn’t confined to the text it generates. The surrounding access, credentials, network position and ability to act become part of the safety case.

OpenAI now plans to involve government agencies and selected safety organisations before deciding how systems at this level can be tested or deployed. That introduces difficult questions about who gets access, what evidence outsiders can inspect, and whether the controls work against determined misuse rather than only during supervised evaluations.

My read is that the milestone here is operational, not competitive. A frontier developer is acknowledging that at least one model may need controls closer to those used around dangerous systems than the lighter safeguards attached to consumer software. Organisations preparing to deploy autonomous agents should notice that shift. Model capability alone can no longer determine readiness; the workspace, network boundary, monitoring and ability to stop the system may decide whether deployment is responsible at all.

Biology Filters Become More Selective

With that warning in view, Anthropic is testing the other side of the safety problem: reducing unnecessary restrictions without opening the riskiest capabilities.

The company updated Fable 5’s biology classifiers on 7 August. Anthropic says the change produced about 85 per cent fewer biology-related fallbacks across its products, allowing more benign health, clinical and educational questions to remain with the more capable model.

That should affect ordinary requests such as asking for an explanation of symptoms, help interpreting laboratory results, or support learning biology. Previously, a classifier could respond to biological language by transferring the request to a less capable model even when the user’s intent was routine. Anthropic says that should now happen substantially less often.

The restrictions haven’t disappeared. Work involving virology, toxicology, molecular design and other professional dual-use fields can still be diverted to Opus 5. Dual-use means the same expertise could support legitimate research or harmful biological activity, depending on the intent and execution. Anthropic therefore says Fable 5 still isn’t available for professional biology research or drug development unless a future trusted-access pathway is introduced.

Both the claimed reduction in fallbacks and the underlying risk judgement come from Anthropic’s own testing; there’s no independent audit in the briefing. Nor does a lower fallback rate establish the quality or medical reliability of the answers that now get through.

For the general public, the immediate result should be fewer frustrating blocks around normal biology and health questions, though model output still isn’t a substitute for clinical care. The more interesting editorial judgement is that useful safety controls have to separate context and capability rather than treating every specialist term as suspicious. Anthropic is trying to make that boundary less blunt, while keeping the most consequential professional work behind a lower-capability route.

AMD Buys Specialised Inference

The capability race also has a physical constraint: once models are deployed widely, somebody has to run every one of those requests economically.

AMD signed a definitive agreement on 6 August to acquire Taalas, a Toronto company specialising in inference chips. Inference is the stage where a trained model processes a live request and produces an answer. At large scale, that repeated work can consume enormous amounts of computing capacity and memory bandwidth.

AMD says Taalas designs around specific inference dataflows, aiming to reduce bottlenecks found in more general-purpose architectures. The plan is to integrate that technology into future accelerator and system-level products alongside AMD Instinct GPUs, EPYC CPUs, ROCm software and its Helios rack-scale systems.

That wording leaves plenty unresolved. AMD hasn’t disclosed the purchase price, provided a product timetable or identified a system customers can order. The transaction still faces customary closing conditions and regulatory approval. Its claims about future performance and efficiency are also forward-looking company statements, not independently tested product results.

Even with those limits, the acquisition says something about where hardware competition is heading. General accelerators remain valuable because they can support changing models and workloads. Specialised silicon can sacrifice some flexibility in exchange for better efficiency on recurring operations. As inference volume rises, power, cooling, memory movement and cost per request become at least as important as the ability to train a larger model once.

For organisations planning AI infrastructure, there’s nothing new to purchase from this deal today. My assessment is that the procurement question will increasingly move beyond which company sells the largest general accelerator. Buyers will need to consider the mix of models they actually run, how stable those workloads are, and whether specialised components can reduce ongoing inference costs without creating an awkward dependency on one architecture.

Australia Sets Terms for AI Power

Australia, for its part, is trying to set the energy terms before the largest AI projects are already locked in.

Energy minister Chris Bowen said on 5 August that the Commonwealth would legislate national minimum standards requiring new AI data centres to rely on additional renewable generation. He made the commitment despite opposition from Queensland and the Northern Territory.

Under the proposal Bowen described, a data centre powered only by gas wouldn’t meet the standard and wouldn’t be allowed to proceed. Gas could still serve as backup or firming power when renewable supply isn’t available. Operators would also be expected to pay their grid-connection costs, minimise water consumption, improve energy efficiency and put at least as much energy into the grid as they use.

The connection-cost provision tackles a less visible concern. The Australian Energy Market Commission has warned that without rule changes, some costs created by large data-centre connections could be spread across other electricity customers. A facility might bring investment and demand, while households and smaller businesses absorb part of the network upgrade bill.

None of this is settled law yet. The legislation hasn’t been enacted, and the detailed definitions, registration process, enforcement mechanism and treatment of projects already under development remain unresolved. Those details will decide whether the policy creates a firm national baseline or a collection of negotiable exceptions.

The political calculation is fairly clear. Australia has renewable resources, available land and growing appeal to data-centre investors. Bowen is trying to use that appeal as bargaining power before AI demand entrenches new fossil-fuel generation or shifts infrastructure costs onto the public.

Developers considering Australian sites should now treat renewable supply, flexible demand, water use and grid charges as core design constraints rather than matters to solve late in planning. My view is that this could make some projects harder or more expensive to establish, but also produce a more honest price for them. Cheap computing isn’t genuinely cheap if its power connection and environmental load are quietly carried elsewhere.

The AI Bill Meets Management

Computing costs don’t stay abstract for long. Rippling says its own AI enthusiasm produced a spending curve that management could no longer ignore.

The company has launched an AI Spend Console that attributes costs to employees, teams and roles, routes requests between models, and tries to compare consumption with work output. Rippling told TechCrunch its AI expenditure had been growing by 80 per cent month over month and was on track to equal 40 per cent of its research-and-development payroll budget.

The spending was highly concentrated. According to Rippling, between 10 and 15 per cent of employees generated about 60 per cent of its AI costs, while one engineer was spending $50,000 a month. Those figures are company claims and haven’t been audited, but they illustrate how flat access to powerful models can generate a very uneven bill.

Rippling says model routing and spending controls reduced July’s token cost to 37 per cent of April’s, while usage in both periods remained close to 600 billion tokens. Routing can send a routine job to a cheaper model and reserve expensive systems for work where their extra capability may justify the price.

The harder part is deciding whether the spending produced value. Rippling’s console uses proxies including pull requests, lines of code and the amount of review rework. Those are measurable, but none reliably establishes productivity or software quality on its own. A valuable engineer may delete code, prevent a bad feature or spend days resolving a subtle design problem. A prolific stream of generated changes may create more maintenance work later.

The product can be bought with Rippling or as a stand-alone offering, with additional usage-based AI costs. For enterprises, that creates a commercial way to govern expenditure across several models. It also creates detailed monitoring of individual workers.

My judgement is that workplace AI is entering its cost-accounting phase. Access will increasingly depend on whether organisations can defend the expense, not merely show adoption. Sensible model routing could make heavy use cheaper. But if management confuses visible activity with value, cost control can slide into surveillance and reward the easiest work to count rather than the work that genuinely helps.

Private Power, Public Emissions

A striking Texas project shows what can happen when a data centre avoids one infrastructure problem by creating another.

Reporting published on 8 August says Amazon’s planned data centre in Pecos County has an on-site natural-gas power plant permitted to emit as much as 33 million tonnes of carbon dioxide each year. TechCrunch, citing reporting by The New York Times, says that ceiling would exceed the emissions of any existing power plant in the United States.

That number needs careful handling. A permit ceiling is not a prediction. The facility is planned rather than operating, and its eventual construction, generation mix, utilisation and actual emissions could all differ substantially from the permitted maximum.

Amazon has confirmed the centre would use new on-site generation and says it wouldn’t increase electricity costs for Texas households. That addresses one criticism of hyperscale projects: they can compete for existing grid capacity and help drive costly upgrades. But private generation doesn’t erase the impact; it changes where the impact appears.

The context is uncomfortable for Amazon’s climate commitments. The company previously reported that its overall carbon emissions rose 16 per cent in its latest reporting year, despite a pledge to reach net zero by 2040.

For infrastructure planners and policymakers, the lesson is to examine the whole system rather than accepting grid independence as the end of the discussion. My reading is that private power may protect consumer electricity bills, yet transfer the burden to emissions and local permitting. The 33-million-tonne ceiling gives regulators a concrete measure of just how large that transferred risk could become.

One CMS, National Exposure

The next security story is less futuristic, and that’s precisely why it deserves attention: old shared software can create risk at national scale.

Two researchers presenting at DEF CON reported that a country-wide scan found security weaknesses affecting more than 10,000 Polish public entities and roughly 250,000 websites. The exposed sectors included courts, hospitals and airports.

One focus was PAD CMS, an end-of-life content-management system used for public websites. The researchers said its vulnerabilities allowed passwordless access to more than 300 sites. They also reported a separate flaw affecting websites belonging to about 245 courts—roughly two-thirds of Poland’s judiciary.

CERT Polska has documented several PAD CMS vulnerabilities, including unrestricted uploads of dangerous files, cross-site scripting and cross-site request forgery. In plain terms, those flaws can allow hostile content onto a server, inject scripts into pages, or trick an authenticated user’s browser into performing an unwanted action.

Exposure doesn’t prove that every site was exploited or breached. The overall counts come from the researchers’ conference presentation, and the briefing didn’t identify a complete public methodology or an independently validated dataset. The findings should therefore be read as evidence of widespread risk, not a claim that a quarter of a million compromises occurred.

The scale still reveals a structural weakness. Courts, hospitals and airports may look like separate targets, but a neglected common supplier can connect their security fate. Once software reaches end of life, vulnerabilities accumulate while responsibility becomes fragmented among agencies, contractors and local administrators.

Public bodies need coordinated asset discovery, replacement funding and vulnerability reporting rather than treating every affected website as an isolated repair. My assessment is that the hardest part isn’t patching a single flaw. It’s finding every inherited installation and assigning someone the authority and budget to retire it before a shared dependency becomes a national incident.

What Changes for You

There is one practical change developers can use now, with an important split between experimental workflow features and enforceable enterprise policy.

GitHub expanded Copilot’s agent workflow on 7 August with concurrent sessions, an experimental isolated-worktree command and file rewind. In the Copilot command-line interface, the new slash-worktree command can create a separate Git worktree for an agent task. That gives the task its own working directory and branch state, reducing collisions with changes in your main workspace.

The slash-rewind command can restore both the conversation and changed files to an earlier point while preserving later edits. Copilot’s app and command-line interface also have stronger support for concurrent or side sessions, and Visual Studio Code can send element-level browser feedback to an agent.

For working developers, the immediate difference is that parallel agent jobs can be separated more cleanly and unwanted edits can be recovered with less manual untangling. The limitation is maturity: the worktree command remains experimental, and GitHub hasn’t published evidence showing how much these features reduce real incidents. Isolation also doesn’t remove the need to review what an agent changed.

Enterprise owners have a firmer control. GitHub has made centrally managed allowlists for Model Context Protocol servers generally available. MCP servers connect agents to tools and data. Administrators can allow or deny a remote server by its URL, or a local one by its exact command. The policy fails closed when its configuration can’t be verified, and initially covers the Copilot app, Copilot CLI and Visual Studio Code. GitHub explicitly warns that a server’s display name isn’t a security control.

For developers inside managed organisations, that may mean an unapproved integration simply stops working. For security teams, it makes tool permissioning enforceable rather than a request in a setup guide. The useful shift is that agent safety is moving into the harness itself: isolated workspaces, recoverable changes and verified tool boundaries. Those controls make ambitious agent workflows more practical, but only within the clients covered by the policy and the limits of configurations administrators can inspect.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. openai.com/index/responding-next-frontier-critical-cyber-capabilities
  2. techcrunch.com/2026/08/07/openai-says-it-slowed-astra-model-development-over-security-concerns
  3. anthropic.com/news/improving-fable-5-s-biology-safeguards
  4. newsroom.amd.com/news/amd-acquires-taalas-ai-inference
  5. theguardian.com/australia-news/2026/aug/05/chris-bowen-vows-to-override-queensland-and-nt-to-force-renewables-powered-ai-datacentres
  6. techcrunch.com/2026/08/07/after-rippling-blew-millions-on-ai-in-months-it-built-an-employee-roi-tool
  7. techcrunch.com/2026/08/08/planned-amazon-data-center-could-become-the-biggest-climate-polluter-in-the-u-s
  8. techcrunch.com/2026/08/07/security-researchers-scanned-the-polish-web-and-found-courts-hospitals-and-airports-at-risk-of-hacks
  9. cert.pl/posts/2025/09/CVE-2025-7063
  10. github.blog/changelog/2026-08-07-github-copilot-weekly-releases-august-3
  11. github.blog/changelog/2026-08-06-mcp-allowlists-in-enterprise-managed-settings