All episodes

AI & Tech Daily

AI Agents Get Their Own Computers — and Their Own Security Problem

18:43

OpenAI is pushing agents beyond individual chats with Dots, GPT-6.1 Sol and cloud-hosted Codex environments. We examine what persistent, tool-using assistants change for users, and why permissions and memory now belong in ordinary setup. Also: six major AI developers sign a voluntary US audit accord, AMD agrees to buy World Labs for US$8.2 billion, Australia's cyber agency warns about compromised AI credentials, Anthropic's reported IPO figures expose the cost of frontier AI, H Company releases open-weight computer-use agents, and new previews open Android Studio to third-party coding agents and Google Cloud API Gateway to streamed model responses.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

The Agent That Stays on the Job

An AI assistant can now keep working after you leave the conversation. That convenience gives its memory, app access and cloud computer much more weight.

Our main story is OpenAI's move from session-bound chat to persistent agents, and what happens when delegating work also means delegating access. OpenAI began rolling out Dots on 29 September. Each Dot runs on GPT-6 Astra, has its own computer in the cloud and can use connected apps. It can form memories from content in those apps, follow a schedule and do proactive work between conversations.

That changes the basic unit of interaction. A chat starts when you type and usually stops when you leave. A Dot can continue as an ongoing worker: following a research topic, handling recurring administration or preparing something before you've asked again. The value is continuity. You don't have to rebuild the context and reissue the job every time.

The trade-off follows directly from that design. A persistent agent can remember more, reach more and act for longer. Connecting an app isn't merely a way to make one answer better; it can supply material for memory and give future work a route back into that service. People who receive access need to review the connected apps, the information being retained and the work the agent is allowed to perform. In my view, that review has become part of normal setup, alongside choosing a model or adjusting notifications.

Memory, app access and scheduled work also need to be considered separately. A user might be comfortable with a Dot remembering material from one connected service but not with it acting there on a schedule. Another task may justify proactive work while needing a narrower set of source material. The product's usefulness comes from combining those capabilities, yet the safest boundary for each one can differ. That makes a single approve-everything decision a poor fit for continuing work.

Access is limited for now. OpenAI says Dots are rolling out gradually to Pro users outside the European Economic Area, Switzerland and the United Kingdom, to Business Premium users, and through an administrator-enabled beta for Enterprise. Independent evidence about reliability in everyday work isn't available yet. OpenAI also hasn't announced the usage terms that follow the initial one-month allowance period. So the interesting capability is real, but its steady-state cost and dependability remain open questions.

Two related releases extend the same idea into model and coding infrastructure. GPT-6.1 Sol is available through the API with a 1,050,000-token context window. Standard short-context pricing is US$2 per million input tokens and US$10 per million output tokens. That large context gives developers room for unusually long working histories or source collections, although a large window doesn't remove the need to control what enters it. OpenAI's published model information also includes tool support and data-residency limits, which organisations will need to match against their own requirements.

Codex Cloud, released for authorised Enterprise members, provides isolated cloud workspaces that can keep coding while a developer's local computer is asleep. The practical shift is similar: the job belongs to a managed remote environment rather than the laptop and the live session. For engineering leaders, that can make longer tasks easier to delegate, but it also makes workspace authorisation and access policy operational concerns rather than optional extras.

My read is that Dots are the clearest signal here. The assistant is becoming a continuing worker with a computer, memory and credentials. If the rollout proves reliable, continuity could remove a lot of repetitive setup. The price of that convenience is a more serious permissions model, because the agent doesn't forget its reach when the chat window closes.

A Voluntary Audit Baseline

That persistent access raises a bigger governance question: who checks the companies building the most capable systems?

Executives from Anthropic, Google, Meta, OpenAI, Nvidia and xAI signed a voluntary accord with the US President on 29 September. It calls for robust internal controls, assessment by an independent external auditor, and board-level review of both internal and external audit reports. The text also says those measures could later be put into law or regulation.

For the signatories, this creates a public commitment to something more structured than a company publishing its own safety claims. An external auditor is meant to provide separation from the people developing the systems, while board review puts responsibility at the top of the organisation. Those are useful elements of a governance framework, especially as advanced models gain access to tools and longer-running tasks.

But the missing details determine how much confidence the accord deserves. It is voluntary, and the President described it as morally binding. No overall overseer had been named when the agreement was reported. The audit standard, the required scope, what gets disclosed and what follows a serious finding are all unspecified. Two audits can carry the same label while examining very different things, so the word independent only gets us part of the way.

For organisations buying AI services, the accord may become an early industry baseline, but it isn't yet a compliance standard they can safely build around. My test would be whether customers and the public eventually get enough visibility to understand what was examined, where controls failed and whether failures had consequences. Without that, a board can receive a report and the outside world still learns very little.

The accord is still significant because six large developers have accepted the principle of external review and board accountability in the same document. Its credibility now depends on turning those broad commitments into comparable audits rather than six private interpretations of good practice.

AMD Buys Its Way into Spatial AI

From rules around frontier models, let's shift to the hardware bets being made underneath them.

AMD entered a definitive all-stock agreement on 28 September to acquire World Labs for about US$8.2 billion. The transaction is expected to close by the end of 2026, subject to regulatory approvals and the usual closing conditions. World Labs founder Fei-Fei Li is due to become AMD executive vice-president and chief scientist once it closes.

World Labs develops spatial-intelligence models that generate, reconstruct and simulate interactive 3D environments from text, images and video. That work has potential uses in robotics and simulation, where an AI system needs some representation of spaces, objects and how they behave rather than only generating text or flat images.

For AMD, the unusual part is bringing model research of that kind directly inside the chipmaker. Hardware companies normally learn a great deal from the workloads their customers run. Owning the research group gives AMD another source of information: researchers working on emerging physical-AI problems can help expose where compute, memory, software and system design are getting in the way. AMD says the team will help shape future hardware, software and systems. It hasn't announced a resulting processor or a product timetable.

That makes the US$8.2 billion price a strategic wager rather than evidence of a finished product advantage. AMD is buying closer visibility into workloads that could matter for robotics and interactive simulation. It still has to convert that knowledge into chips and software that customers prefer, and do it after winning regulatory approval and integrating an expensive research organisation.

For companies planning infrastructure, there is nothing concrete to buy on the back of this announcement. The useful signal is that AMD sees spatial models and physical AI as important enough to influence its future systems from inside the company. What I'd watch is whether World Labs research begins to appear in AMD's software stack and accelerator design, because that is where the acquisition either becomes differentiated infrastructure or remains a very costly view of the future.

AI Credentials Become Privileged Access

There is a much more immediate task for anyone already running agents inside an organisation.

The Australian Signals Directorate issued guidance on 28 September that treats access to advanced AI services as a security-sensitive organisational asset. ASD says malicious actors are obtaining access through compromised API keys, authentication tokens and user sessions, as well as vulnerable applications and third-party arrangements.

A stolen key can obviously run up model usage, but the possible damage is wider. ASD says unauthorised access can consume credits, disrupt legitimate work, generate harmful material or be used in attempts to distil model capabilities. Once an agent is connected to enterprise tools, a compromised agent identity may also provide access to organisational data and systems downstream. The credential isn't only a meter for tokens at that point. It can be a key to whatever the agent has been allowed to reach.

The agency's recommendations are practical and familiar from service-account security. Keep an inventory of AI services and assign an accountable owner. Apply least privilege, use phishing-resistant multifactor authentication, store secrets in managed systems, and enforce spending and rate limits. Monitoring and incident response also need to recognise AI-specific consumption, because an unexpected surge can be a security signal as well as a large bill.

My practical takeaway for organisations is to put production agents in the same access reviews as other privileged identities. Record which model credential each agent uses, which tools sit behind it, who can rotate it and what limits contain misuse. An agent with permission to query internal data, update a workflow and spend against an API account has a meaningful bundle of privileges, even if its interface still looks like a chat box.

Spending limits deserve particular attention because they provide a second boundary when a credential escapes. They won't protect connected data, but they can restrict consumption and make abnormal use easier to spot. Least privilege does the complementary job: if an agent only has the tools needed for its task, one compromised identity doesn't automatically expose every connected system.

ASD didn't quantify how common these compromises are or identify affected Australian organisations, so the guidance isn't evidence of a measured local surge. It is still a timely reframing. As agents become persistent and tool-using, their credentials belong in the privileged-access inventory, not in an informal list of developer API keys.

The Bill Behind Frontier AI

The security controls may be familiar. The financial scale behind frontier models is anything but.

Reuters reviewed Anthropic's confidential IPO prospectus and reported rapid revenue growth alongside enormous costs and long-term commitments. According to that reporting, Anthropic's 2025 revenue rose twelvefold to nearly US$4.6 billion, while operating losses exceeded US$8 billion. Compute and infrastructure spending reached US$7.33 billion, more than half of reported operating expenses.

That compute figure is larger than the year's reported revenue, even before the other operating costs are considered. Revenue and infrastructure spending measure different sides of the business, so it isn't a profit calculation by itself. It does show how far current model economics depend on funding capacity and an expectation of much larger future demand.

The reported near-US$42 billion net loss needs careful handling. Roughly US$34 billion of it came from accounting charges tied mainly to financing liabilities, rather than operating cash expenditure. That distinction doesn't make the operating economics cheap; it does stop the headline net loss from being mistaken for a same-sized cash bill from running models.

The most striking forward-looking figure is US$518 billion in cloud, compute and infrastructure obligations over coming years. Reuters also reported that nearly one-quarter of 2025 revenue came from two customers. Anthropic declined to comment, and the prospectus wasn't public. IPO timing, valuation and the disclosed figures could change before a final filing.

For cloud providers and chip suppliers, those obligations represent potential demand on a remarkable scale. For enterprise customers, they also show how the frontier-model market is being built: providers are committing to infrastructure far ahead of present revenue in the expectation that demand and capability will catch up. A small number of large customers then carries more weight in that equation than a broad revenue figure might suggest.

My read is that procurement decisions need to account for this capital structure without pretending we can predict the winner. Long commitments can secure the compute needed to improve a service, but they also create pressure for high utilisation and durable revenue. Customer concentration adds another dependency. If one of two major customers changes course, the effect can be disproportionate.

None of this tells us that Anthropic's technology is weakening, or that the reported plan cannot work. It tells us frontier capability is being financed through obligations vastly larger than the business's current annual revenue. The prospectus figures are a useful reminder that model competition isn't only a research contest. It is also a bet that future customers will support infrastructure commitments being made now.

Open Weights for Computer Use

For developers who want more control over an agent, H Company has released a different kind of option.

Its Holo4 models are designed to operate through graphical interfaces, code, Model Context Protocol connections and APIs. The release includes a 27-billion-parameter dense model and a 35-billion-parameter mixture-of-experts model. Both main models are available through the H Models API, while downloadable weights come in BF16, FP8, NVFP4 and four-bit GGUF formats.

That range matters because computer-use agents have largely been experienced as hosted services. Downloadable weights give developers the option to inspect and run the model in an environment they control, subject to the hardware and engineering needed for a model of this size. H Company also published task trajectories and its benchmark methodology, which gives teams more material for understanding how the agent reached an outcome.

There are still two reasons to test carefully. First, the company acknowledges that model harnesses and task subsets differ, so performance figures aren't directly comparable in every case. Second, its demonstrations consumed roughly 1.3 million to 2.4 million tokens per task. Long computer-use workflows can therefore carry substantial context costs, whether those costs show up as an API bill or as demand on self-hosted infrastructure.

The useful development is inspectability, not a proven production result. Open weights and published trajectories make it easier to examine behaviour than a closed endpoint alone. They don't establish reliability on a company's own interfaces, permissions or edge cases.

For developers considering Holo4, my judgement is straightforward: reproduce the relevant workflow with your own harness and measure completion, token use and failure recovery together. A model that can complete a graphical task in a demonstration may still be uneconomic or brittle in a long-running application. Holo4 broadens the available choices, and the disclosure around trajectories helps, but deployment confidence has to come from the work you actually need it to do.

What Changes for You

Two developer previews could remove some awkward plumbing, provided you're comfortable working ahead of the stable release.

Google has added Bring Your Own Agent support to the Android Studio Rabbit 2 canary. Codex, Claude Agent and Google Antigravity are the initial featured options, but the connection is based on Agent Client Protocol, so any compliant agent can plug in. The agent can receive Android-specific project graph and build context, and use IDE-native build diagnostics, Compose previews, SDK tools and emulator controls under granular permissions.

For Android developers, that means changing agents no longer has to mean giving up the context and testing controls built into the IDE. You use your existing provider plan or API key, so pricing and data terms still follow the provider you choose. The main limitation is maturity: Rabbit 2 is a canary preview, Google hasn't given a stable-release date, and reliability across third-party agents hasn't been independently established. It suits evaluation work more than a production workflow you can't afford to disrupt.

Google Cloud has also added request-and-response streaming to API Gateway in Public Preview. It supports HTTP/2, HTTP/1.1 chunked transfer, Server-Sent Events, WebSockets and bidirectional gRPC, including token-by-token output from language models. Developers can now put a streaming model endpoint behind the managed gateway without buffering the full response or maintaining a separate streaming proxy.

There is one configuration catch: streaming has to be enabled when the gateway is created, and that mode can't be changed later. With no general-availability date announced, the sensible migration path is a new test gateway rather than trying to reshape an existing one. If the preview performs well, responsive AI applications gain a managed route for streaming and one less custom component to operate.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. help.openai.com/en/articles/20001530-getting-started-with-your-dot
  2. developers.openai.com/api/docs/models/gpt-6.1-sol
  3. help.openai.com/en/articles/10128477-chatgpt-enterprise-and-edu-release-notes
  4. apnews.com/article/trump-ai-anthropic-musk-595796511f110fc006cca0d01329733e
  5. ir.amd.com/news-events/press-releases/detail/1299/amd-to-acquire-world-labs-to-advance-the-future-of-ai-compute
  6. ir.amd.com/financial-information/sec-filings/content/0000002488-26-000182/amd-20260926.htm
  7. cyber.gov.au/about-us/view-all-content/alerts-and-advisories/protect-your-organisations-ai-services
  8. ca.marketscreener.com/news/anthropic-s-ipo-prospectus-shows-sweeping-ai-vision-surging-costs-ce785addd98afe2c
  9. huggingface.co/blog/Hcompany/holo4
  10. android-developers.googleblog.com/2026/09/build-your-way-use-any-ai-agent-in-android-studio.html
  11. docs.cloud.google.com/release-notes