Back to the show

AI & Tech Daily

Persistent Agents Meet Their Permission Problem

18:22

OpenAI's Dots turn ChatGPT agents into persistent workers with connected apps and their own cloud computers, raising immediate questions about permissions, monitoring and price. We also look at AMD bringing MCP into embedded engineering, GitHub's AI security taskflows finding 24 Android flaws, Anthropic's machine-readable provenance marks, a new bipartisan governors group on AI, safer Cloud Storage transfers by default, Washington's new 'Super Intelligence' terminology, and GitHub Copilot taking control of desktop apps.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

Dots Keep Working

An AI agent that keeps working after you close the chat is far more useful. It also has far more time to make a mess with whatever you've allowed it to touch.

Our main story today is OpenAI's rollout of Dots, and what persistent agents mean for the people handing them access to real apps, information and actions. A Dot is a ChatGPT agent with an ongoing goal, connections to other applications and a dedicated computer in the cloud. Instead of giving it a task, waiting for the answer and beginning again in another conversation, you can set work that continues between chats. You also decide which actions it may take on its own.

That changes the shape of delegation. A conventional chatbot mostly waits for you. A persistent agent can keep gathering information, operating software and advancing a job while you're doing something else. For recurring research, administration or project work, continuity removes a lot of the setup that makes current agents feel like clever demos rather than dependable colleagues. The dedicated cloud computer is important too: the work isn't tied to an open tab or a laptop that needs to stay awake.

Access is narrow at launch. OpenAI began a gradual rollout on 29 September for eligible Pro and Business Premium users aged 18 or over in supported markets. Pro users in the European Economic Area, Switzerland and the UK aren't included at launch. Enterprise access is in beta and an administrator has to enable it. OpenAI is excluding Dot usage from eligible plan allowances for the first month, but it hasn't published what the usage terms will be after that. That missing number could determine whether Dots become everyday workers or something people reserve for higher-value jobs.

Persistence also makes permissions part of the work itself. Giving an agent access to a calendar is different from letting it create appointments; connecting a document store is different from allowing edits; and a permission that seems reasonable for a one-off job may be too broad for an agent operating over days. People will need a clear view of what each Dot can see, what it can change and when it should stop for approval. Monitoring can't be an occasional clean-up exercise when the agent is designed to continue without a fresh prompt.

There are still big unknowns. The rollout is gradual, later pricing is unpublished and there isn't independent evidence yet about reliability over long periods of autonomous work. OpenAI has established the product shape, not proved that a Dot will stay on course through a messy week of changing information and unexpected application behaviour.

My read is that persistence is the step that makes agents genuinely useful for ongoing work, but it also ends the era when permissions could be treated as setup trivia. For an eligible user, the sensible unit of delegation is no longer merely the task. It's the task, the accessible data, the allowed actions and the conditions that bring the agent back to you.

MCP Reaches Embedded Engineering

That shift becomes more interesting when the agent can reach tools built for a very specific profession.

AMD has launched Ross, an agentic assistant for embedded-system engineering. It connects language models to AMD development tools, documentation and expert-authored workflows through Model Context Protocol servers and agent skills. MCP is the connection layer that lets an agent discover and use supported tools in a structured way. Here, that can include searching technical documentation, running supported commands and generating code, along with work across hardware design, debugging, power optimisation and embedded AI.

The useful detail is that AMD says Ross is client-agnostic. A development group can choose its language model, IDE and command-line environment rather than adopting one fixed assistant interface. The downloadable pieces currently include a Vivado AI extension for Visual Studio Code and Vivado MCP servers for Windows and Linux, as well as local knowledge bases. AMD says it plans to add more tools and workflows each month.

This pushes agent integration beyond the office apps and web services where MCP first gained attention. FPGA work spans code, hardware configuration, specialist tools and an enormous amount of documentation. A general coding assistant can explain a concept or draft a module, but it normally can't inspect the state of the engineering toolchain or execute the tool commands that move the design forward. Ross is an attempt to bridge that gap.

Coverage is the key limitation. AMD's development-speed claims are vendor claims, not an independent benchmark, and the list of supported components is still expanding. Engineers also remain responsible for checking recommendations and actions. In hardware development, a plausible answer isn't enough; the result still has to pass the normal engineering review.

For embedded developers, the promising part is less the name of the assistant than the arrival of a supported path into the toolchain. If AMD keeps exposing useful operations through MCP, teams may be able to automate more of the repetitive search-and-command work. Until then, the value will vary according to whether Ross reaches the particular tools and stages that consume a team's time.

Security Agents Scale the Search

A specialist workflow can also be valuable when the model itself is an unreliable judge.

GitHub Security Lab says it found and disclosed 24 vulnerabilities in Android applications using open-source, task-specific AI auditing workflows. The approach breaks an audit into targeted stages and looks directly at Android attack surfaces, including exported components, intents and WebViews. Reported findings included exposure of private location data and a chain that could lead to account takeover.

The number is striking, but the method is more useful than the scorecard. Security researchers encoded the way they search into repeatable taskflows. Instead of asking a general agent to find vulnerabilities across an entire application, each stage narrows the question and directs attention to a known class of weakness. That gives the model a job closer to an expert's search procedure, with evidence flowing into the next step.

GitHub is also unusually direct about the limits. The models often estimated severity badly and generated false positives. Researchers had to validate whether a suspected issue was actually exploitable, using proofs of concept or debugging. The 24 findings are reported by GitHub Security Lab, and the announcement doesn't include an independent audit of every result. Running the published workflow also isn't free: it requires a GitHub Copilot licence, premium model requests and potentially substantial token use.

For an application-security group, that leaves a practical division of labour. An agent can repeat a detailed search strategy across more code and surface candidates that deserve attention. A skilled researcher still decides whether the data flow is real, whether an attacker can reach it and how urgently the developer needs to respond. The false positives aren't a footnote because triage time is part of the cost.

I think this is the stronger near-term case for AI security agents. They can scale the search habits of good researchers without pretending that model confidence equals exploitability. Organisations evaluating these tools should measure useful, validated findings against review time and model cost. A long list generated quickly is only an advantage if the people doing the verification get a better use of their day.

Claude Adds Provenance Marks

Finding AI-generated material sounds simpler than finding a software flaw. In practice, the evidence is just as easy to overstate.

Anthropic has documented machine-readable marking for outputs from supported Claude models worldwide. Supported text can carry embedded watermarks, while supported files can use C2PA Content Credentials, an industry format for recording provenance information with digital media. Models launched in the European Union on or after 2 August 2026 are intended to support marking from launch. Where the relevant platform has implemented it, those marks apply across Anthropic's products and services from AWS, Google Cloud and Microsoft Foundry.

Coverage isn't complete. Several older models aren't supported, and part of the Amazon Bedrock rollout is scheduled to finish on 12 October. Anthropic's watermark detector is also a private preview available only to eligible organisations. Even when a detector is available, Anthropic warns against treating the output as a verdict. A detected mark isn't conclusive proof of provenance, and the absence of a mark doesn't prove that a person wrote the material.

There are technical reasons for that caution. Editing or translating text can damage an embedded signal. Very short passages may not contain enough information for reliable detection. File metadata can be stripped as content passes through publishing, messaging or document systems. A provenance chain is useful only while the systems handling the content preserve it.

That still gives publishers and regulated organisations something they haven't consistently had: a machine-readable signal supplied at the point of generation. It could support disclosure, content handling and audit processes, particularly for files carrying Content Credentials. But the policy built around the signal needs room for uncertainty. Automatically rejecting an application, claim or article because one detector fired would ask the technology to make a decision Anthropic says it cannot make conclusively.

The sensible organisational view is evidence, not detection theatre. Preserve the marks where systems can, combine them with access logs and declared workflows, and don't turn an imperfect signal into a binary label for authorship. The infrastructure can improve transparency. It can't reconstruct provenance after every trace has been removed.

Governors Seek a Shared AI Framework

The controls around AI are also being negotiated well beyond the companies building it.

Maryland Governor Wes Moore and Indiana Governor Mike Braun are forming a bipartisan group of US governors to develop an artificial-intelligence framework. Governors from both major parties are involved, and more are expected to join. For now, though, the announcement doesn't include the framework itself, a full list of participants or a timetable for producing it.

State governments have real levers. They buy technology, set rules for activities within their borders and deal directly with local effects such as data-centre development. Developers and operators can therefore face different obligations as they move from one state to another. A coordinated framework could reduce some of that patchwork if participating states turn common ideas into compatible procurement rules or legislation.

There is an obvious boundary. Many AI services, data flows and companies operate nationally or globally, while state authority is limited. A governors group can't settle every question left open by federal law, and an agreed framework wouldn't automatically become the same enforceable rule in every state. The membership, scope and capacity to translate a document into consistent action all remain unknown.

For AI developers and data-centre operators, this is a policy signal rather than a new compliance obligation. It says a group of state leaders wants a more coordinated position while federal legislation remains unsettled. The useful response for organisations is to watch for actual text: definitions, procurement conditions, reporting expectations and the parts states intend to implement. Until those arrive, claims that the group has solved fragmentation would be well ahead of the evidence.

My judgement is that coordination could make planning easier, even if it also produces new requirements. One coherent state agenda is easier to design for than a collection of conflicting ones. The hard part begins when governors have to agree on details and carry them through their own institutions.

Cloud Storage Checks the Journey

Here’s a smaller change that fixes the sort of gap most developers would rather never discover in production.

Google has enabled end-to-end checksumming by default in the latest versions of all Google Cloud Storage SDKs. A checksum is a compact value calculated from data. If the value calculated after transfer doesn't match the one from before, something changed on the journey.

Cloud Storage already calculated a checksum on the server, but an application that didn't supply one left an important gap before the data arrived there. With the new default, the SDK calculates an upload checksum and sends it with the object when the application hasn't provided one. The current SDKs can also verify checksums during downloads. Range reads over gRPC use built-in range checksums, so applications fetching only part of an object still get integrity protection suited to that operation.

The catch is pleasantly ordinary: deployed applications have to update their SDKs. Google's announcement doesn't list the minimum affected version for every supported language, and an existing service doesn't gain the new default merely because Google announced it. Developers need to check the dependency version in the application they actually ship.

This is the kind of platform default I like. It reduces the chance of silent in-transit corruption without requiring every application developer to design a separate verification path. For teams moving important objects, the practical job is to update, confirm the checksum behaviour in their chosen SDK and keep any deliberate custom handling intact. A good default does the most work when it reaches production.

Washington Renames AI

Language can change overnight even when the underlying technology doesn't.

A US executive order dated 29 September directs federal departments and agencies to replace the terms 'Artificial Intelligence' and 'AI' with 'Super Intelligence' and 'SI' in official non-statutory communications and policy documents, to the extent the law permits. The order initially defines the new labels by pointing back to the existing statutory definition of artificial intelligence. So the terminology changes immediately in the relevant federal material, while the legal concept begins from the definition already on the books.

The direction isn't retroactive. Agencies don't have to rewrite previously issued regulations, contracts, grants or historical documents. The White House science and technology adviser has 60 days to propose legislative language for a federal definition of Super Intelligence. That future wording could be significant, but it hasn't been published yet.

For organisations working with US agencies, the immediate consequence is document friction. A current policy paper may say SI while the statute it relies on says AI. Existing contracts and grants may retain the older term. Technical standards, international material and ordinary industry usage will still need to be mapped to the new federal language. Search, records management and legal review all have to cope with both sets of words for some time.

It would be a mistake to infer a capability jump from the label. The order's interim definition points to the same technology covered by the existing statutory term; it doesn't say that today's systems have suddenly reached a new technical threshold. The important unknown is whether the proposed legislation introduces a materially different definition and, if so, how it interacts with existing law.

For policy and compliance staff, this is now a terminology change to track carefully, not a reason to rewrite historical instruments that the order explicitly leaves alone. Clear cross-references will matter more than enthusiastic relabelling. Otherwise two people can discuss AI and SI as if they're different systems when, under the interim definition, they're pointing at the same field.

What Changes for You

One new agent capability is immediately useful to developers, but it deserves the same care as any other privileged desktop tool.

GitHub has released computer use in public preview for Copilot CLI and the Copilot app on macOS and Windows. Copilot can read content that is visible on screen or exposed through accessibility interfaces, then click, type, scroll, drag and navigate through application workflows. That gives an agent a route into legacy or desktop-only software with no API, command-line interface or MCP integration.

For a developer, that could connect an otherwise isolated configuration tool, test utility or internal application to a larger automated workflow. The integration barrier drops because the software doesn't need a purpose-built connector first. Copilot asks for approval before it controls an application, and users can review permissions that persist. Organisations can disable the capability.

The limitation is that this remains a public preview. GitHub hasn't supplied independent task-success or reliability measurements, and visual control can fail in ways an API call doesn't: a changed layout, an unexpected dialogue or the wrong focused window can redirect an action. The required screen-recording, accessibility and application-control permissions also give security teams a new privileged path to assess.

If you use Copilot on supported desktops, the change is access to software that agents previously couldn't operate. Treat each approved application as a meaningful permission boundary, especially where a click can alter data or trigger a real process. Computer use can make old tools automatable; preview status means the person running the workflow still needs a clear view of what Copilot is about to control.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. help.openai.com/en/articles/6825453-chatgpt-release-notes
  2. newsroom.amd.com/news/amd-ross-agentic-ai-embedded-design-development
  3. amd.com/en/support/downloads/ross-agentic-ai.html
  4. github.blog/security/how-we-found-24-android-vulnerabilities-using-our-open-source-ai-security-agent
  5. support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
  6. apnews.com/article/ai-technology-trump-data-centers-wes-moore-ed0337966ba4f598c7e7bb6ad93ba6f3
  7. cloud.google.com/blog/products/storage-data-transfer/enabling-end-to-end-checksums-in-cloud-storage
  8. whitehouse.gov/presidential-actions/2026/09/inaugurating-the-era-of-super-intelligence
  9. github.blog/changelog/2026-10-01-github-copilot-can-now-interact-with-desktop-apps