AI & Tech Daily
Agents Accelerate Faster Than Their Guardrails
Anthropic says Claude now leads 26% of its model research and development, putting a sharper number on AI helping to build its successors. We examine the limits of that claim, the monitoring burden created by 30,000 concurrent internal agents, a new embedded-evaluation deal with Accenture and California's push for independently tested safeguards. Also: sovereign Gemma 4 inference on AWS, MLPerf's first end-to-end RAG and edge-agent tests, active exploitation of exposed Vite development servers, Google's experimental household agent, and the open-source Mantis security harness.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
Claude Enters the R&D Loop
Anthropic says Claude now leads more than a quarter of its model research and development. The people building frontier AI are already relying on AI to help build what comes next.
That is the development worth sitting with, because the number is specific enough to change the governance conversation. On 17 September, Anthropic published prototype measures of AI-led research and development inside its own laboratory. It says Claude led 26% of model R&D work in August, and either collaborated on or led more than 90%.
The word "led" needs some care. Anthropic isn't saying Claude independently chose the research agenda, ran the experiments and approved the result. The company says Claude was not fully autonomous in any measured part of the work. Even the lead category still involves human supervision. But it does mean the model is doing enough of the substantive work for Anthropic to classify its role as more than assistance.
There is a meaningful difference between a model suggesting code in an editor and an agent carrying a research task through several steps. The latter can plan work, use tools, inspect results and continue until it reaches a stopping point. Anthropic's measure is an attempt to capture how much initiative the system takes across that process. The rating remains contestable, but the activity being measured is broader than autocomplete or a researcher asking an occasional question.
The scale around that work is just as revealing. Anthropic reports roughly 30,000 research and engineering agents operating concurrently on its main internal platform. Their actions pass through both online monitoring, while work is happening, and offline monitoring afterwards. That sounds less like one clever chatbot beside a researcher and more like a production system that has to schedule, observe and constrain a large digital workforce.
The implication is that supervision at this scale can't depend on a person reading every action as it happens. Monitoring systems have to decide what looks unusual, when to interrupt and which records a person should inspect later. Anthropic hasn't claimed that this solves the oversight problem. Its disclosure shows how quickly monitoring is becoming part of the basic infrastructure for research rather than an optional review layer added at the end.
Anthropic is also trying to measure how compute is allocated and how quickly AI involvement is changing. That gives governments and independent evaluators something concrete to ask of other frontier laboratories. If several labs published comparable measures, observers could begin tracking whether AI-assisted research is accelerating, where people still intervene and whether oversight capacity is growing with deployment.
We're not there yet. These are self-reported figures from one company. Anthropic constructed the automation index and used Claude for part of the judging. It openly says the method isn't standardised or directly comparable with other laboratories, and the results haven't been independently reproduced. A figure such as 26% can look more settled than the rating system beneath it.
My read is that the headline isn't self-improving AI operating alone. It's that a frontier laboratory now depends on supervised AI work at a scale large enough to shape the pace of its own model development. For organisations governing advanced agents, the hard question shifts from whether a person remains somewhere in the loop to whether monitoring and intervention can keep up with thousands of agents working at once. Common measures and external verification are the next useful step. Without them, the numbers are a valuable disclosure, but not yet a reliable league table.
Evaluators Move Inside the Lab
That brings us from measuring the work to testing the systems doing it. Anthropic and Accenture are proposing a much closer form of outside evaluation.
On 18 September, the companies announced a non-exclusive arrangement under which evaluators from Accenture's Faculty business will work inside Anthropic. They are expected to red-team models, assess alignment and test safeguards with access described as comparable to an employee's. Each company expects to invest at least one billion US dollars in evaluation capacity over five years. Anthropic will directly fund Accenture's work, while also pursuing pilots with nonprofit evaluators under other funding arrangements.
The access is the interesting part. An evaluator looking only at a public model endpoint sees the product after many internal choices have already been made. Someone embedded during development can examine failure modes earlier, observe how safeguards behave in realistic internal workflows and follow a problem through the systems around the model. That could reveal risks a periodic external audit misses.
It could also make evaluation more continuous. A test performed just before release captures one model and one configuration. Embedded evaluators may be able to see how behaviour changes while a model, its tools and its safeguards are still being developed. That makes it easier to find the mechanism behind a failure, rather than merely recording that the finished system produced one.
But the announcement arrives before the operating rules. Anthropic says the evaluator-access standards, reporting arrangements and a settled independent funding model don't yet exist. We don't know how findings will be published, what happens when the evaluator and laboratory disagree, or which conflict controls will apply when the laboratory is paying for the work. Employee-like access can make an evaluation more technically credible, while commercial dependence can make its independence harder to judge. Both can be true.
For frontier labs and regulators, this is a concrete model for continuous scrutiny rather than a one-off safety test. Its value will come from the details that haven't been settled: the evaluator's freedom to investigate, the right to report uncomfortable findings and independent checks on the process itself. A billion-dollar commitment creates capacity. It doesn't automatically create trust.
California Seeks Verifiable Controls
The same gap between a safeguard and proof that it works is now showing up in California's policy agenda.
On 18 September, Governor Gavin Newsom issued an executive order accelerating implementation of the state's new independent-verification and AI-auditor laws. The order also convenes experts to recommend stronger safeguards for frontier models within two months. Among the ideas they are being asked to consider are onsite independent evaluators, third-party verification of safety filings and an emergency shutoff mechanism that has itself been independently tested.
The group will also consider whether the definition of a critical safety incident should cover loss-of-control events. That wording matters because incident definitions determine what a developer has to detect, record and report. It brings failures involving agents and control systems into view, rather than treating safety as a property of a model response in isolation.
Independent verification of safety filings could change the compliance work as well. A laboratory would need evidence another party can inspect, not only a written account of its own controls. That could pull logging, test coverage and access to internal systems into the regulatory process. The exact scope will depend on recommendations that are still to come.
These aren't immediate new duties for AI companies. The possible requirements are recommendations under development, and any additional binding obligations would need further legal implementation. The expert group hasn't yet said how a shutoff would be designed, what systems it would cover or how an independent test would work.
That last point is technically difficult. A frontier model may be served across distributed infrastructure, connected to tools and used by agents carrying out longer jobs. A switch that stops one endpoint is only useful if operators can also contain the surrounding work, credentials and queued actions. That's an inference from the architecture, not a detail settled by the order.
Developers operating in California should prepare for audits that reach further into operating practice and for incident controls that need demonstrable evidence behind them. The encouraging shift is from accepting a laboratory's assurance to asking whether an evaluator can verify it. The unresolved part is whether the final rules define controls that work across the whole agent system, rather than a neat mechanism around one model.
Sovereign Gemma on AWS
There's also a more immediate infrastructure choice for regulated organisations in Europe. AWS has added three Gemma 4 variants to Amazon Bedrock in its European Sovereign Cloud.
The models became generally available on 17 September through Bedrock's next-generation inference engine. AWS is offering OpenAI-compatible Responses and Chat Completions interfaces, so an application already using those common API shapes may be able to change its endpoint and credentials instead of rewriting the integration. That lowers the practical cost of testing a different managed model service.
API compatibility doesn't guarantee identical behaviour. Models can interpret prompts differently, support different features and produce different output even when the request format is familiar. Still, keeping the integration shape removes one source of migration work and lets a developer focus testing on the model and service controls rather than rebuilding the client first.
AWS says inference, customer content and customer-created metadata stay in the eusc-de-east-1 region, with no global cross-region inference. For a European organisation dealing with residency rules, that turns sovereign AI from a broad policy promise into a configuration it can assess and deploy. Gemma is an open-weight model family, but here the customer is buying managed inference rather than operating the weights itself.
That distinction affects control. Running open weights on infrastructure you operate can offer more choice over the software stack, while a managed service removes much of the operational load. Bedrock customers get the managed endpoint and AWS's regional design, not unrestricted control over how the service itself is run.
The boundaries are important. This launch covers one model family in one sovereign region. Usage is billed per token, and AWS says limited retention can still apply to some models for abuse detection. The residency and operator-access assurances are AWS service claims; the briefing found no independent assessment of this new integration. An organisation still needs to map those terms against its own data classification and regulatory duties.
The trade-off is fairly clear. Managed sovereign inference can make deployment easier while keeping data handling within a defined EU region. It also ties that control to AWS's service design, available models and regional limits. For regulated buyers, this is newly worth evaluating, but it isn't the same as owning the complete stack.
Benchmarks Follow the Workflow
Hardware comparisons are changing too, and this one is overdue. MLCommons has started measuring complete retrieval and agent workflows, not only isolated model calls.
MLPerf Inference version 6.1, released on 16 September, adds end-to-end tests for retrieval-augmented generation, usually shortened to RAG, and edge agentic inference. RAG combines document retrieval with generation so a model can answer using a selected body of material. The new test covers document ingestion, retrieval, reranking and question answering. That is much closer to the system an organisation actually operates than a raw tokens-per-second figure.
Each stage can change the outcome. Faster generation doesn't help much if document ingestion is slow, retrieval selects the wrong material or reranking adds enough delay to make the application frustrating. Measuring the chain also makes accuracy gates more important, because the fastest answer has little value when the evidence entering the model is poor.
The edge test uses a multi-turn coding workload under fixed limits for memory, processing, power and context. That matters because an agent doing useful work can make several model calls, carry growing context and use local resources over time. Peak throughput on one prompt says little about whether the complete loop stays responsive and accurate on a constrained device.
Power and memory limits are particularly relevant at the edge. A system may need to sustain a conversation and coding task without the resources of a data centre accelerator. That exposes trade-offs hidden by a short, isolated inference run, including how the workload behaves as context accumulates.
Thirty organisations submitted results in this MLPerf round, including first reviewed entries for five recent or preview processors and accelerators. The new categories are still thin, though. MLCommons reported five systems for the edge-agent test and only one system for each of the end-to-end RAG tasks. Vendors also choose which systems and workloads they enter, so an empty column doesn't necessarily tell us a product performed badly. It may not have been submitted.
For infrastructure buyers, the useful change is the unit of measurement. Ask how the full application performs with its retrieval stages, tool loop, context growth and accuracy threshold intact. The first results are too sparse for sweeping vendor rankings, but the benchmark is pointing procurement towards evidence that resembles real work. A fast model call can still sit inside a slow or inaccurate system.
Exposed Vite Servers Under Attack
A quick but urgent security note for developers: check whether a Vite development server is reachable from a network it shouldn't be.
Singapore's Cyber Security Agency issued a patch-now alert on 17 September after F5 Labs reported mass scanning for CVE-2026-39364. The flaw can let an unauthenticated attacker retrieve files that Vite's server configuration was meant to deny, when those files sit inside otherwise allowed directories. Exploitation requires the development server to be exposed to a network.
Affected releases are Vite 7.1.0 through 7.3.1, Vite 8.0.0 through 8.0.4, and Vite-plus 0.1.15 and earlier. Patched releases include Vite 7.3.2, Vite 8.0.5 and Vite-plus 0.1.16.
F5's honeynet recorded about 32,000 raw events grouped into 807 attacks during August. The probes looked for environment files, cloud credentials and infrastructure state. That telemetry shows active scanning, although it doesn't tell us how many systems were successfully compromised. The files being sought can provide the credentials or configuration needed for a much larger cloud intrusion.
The practical response is to patch, remove external access to development servers and rotate relevant secrets if a vulnerable instance was reachable. Containers, reverse proxies and permissive cloud rules can quietly turn a local convenience into an internet-facing service. The label "development" offers no protection once the network path exists.
An Agent Joins the Household
Now shrink the scale from cloud infrastructure to the family calendar. Google Labs is testing what happens when an agent serves a household instead of one person.
Google expanded its experimental CC agent on 17 September from individual daily planning to household collaboration. Each agent receives a separate verified Google account. Up to six adults can share selected information and tasks with it, and CC can assemble a shared daily brief, update calendars and task lists, pre-fill forms and maintain household memory using Gmail, Chat, Drive, Docs and Calendar.
A separate account is more than a cosmetic choice. It gives the agent an identity that can be granted access, observed and, at least in principle, removed without handing over one family member's credentials. Google says each person chooses what to share, and the agent asks permission before acting or sharing information outside the group. That is a useful pattern for multi-user agents because the people affected by an action aren't always the person who requested it.
Shared memory makes those boundaries harder than they first appear. A useful daily brief may combine information from several calendars and messages, but the resulting summary can reveal details to someone who never had access to the original item. The permission model has to govern what the agent infers and repeats, not only which file or message it can open.
The limits are substantial. CC is an early experiment for people aged 18 or older in the United States, using personal Google accounts, and new users have to join a waitlist. It isn't currently available to Australian households. Google also hasn't published independent evidence about real-world reliability, permission failures or privacy performance.
A household agent could reduce the low-level coordination that fills family chats and calendars. It could also expose one person's messages, schedule or preferences for another person's convenience if consent becomes broad or hard to see. My test for this category is straightforward: each person needs to understand what the agent can access, what it has shared and how to withdraw that access. The account model is a promising start; the experiment hasn't yet shown that the permission experience holds up in daily family life.
What Changes for You
For builders who want something they can inspect now, Google has released Mantis, an Apache-2.0 toolkit and reference harness for agentic software security.
Mantis provides a structure for agents that find, reproduce, triage and patch vulnerabilities. Google says related internal agents scan changes across hundreds of millions of lines of code and prevent hundreds of vulnerabilities each month. Its described triage stage combines agent findings with structural checks and reportedly reaches more than 92% precision in under a minute. Those are Google's internal claims, not independently audited results, and they may not transfer to a different codebase, model or security programme.
The useful change is access to the harness. Security teams and experienced AI builders can inspect the repository, see how the stages fit together and adapt the approach for controlled code-review experiments. That can help continuous security review keep pace with a growing volume of AI-generated code. Apache-2.0 licensing also gives organisations room to modify and integrate the code, although operating it safely remains their responsibility.
The repository is equally direct about the limitation: Mantis is a demonstration, not a supported production product. It needs a strongly isolated environment, a threat model suited to the codebase, deterministic validation and expert review of every finding and proposed patch. Giving an agent exploit reproduction and code-changing tools without that isolation would create a new security problem while trying to solve the old one.
So this is useful for practitioners who already know how to build the containment and validation around it, not a switch that makes an application secure. The promising division of labour is for agents to search broadly and propose fixes, while deterministic controls and security experts decide what is real and what gets merged.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- anthropic.com/institute/measuring-pace-of-ai-development
- apnews.com/article/anthropic-claude-ai-model-self-improvement-4d3a7430f57cbc7c39e1c5f2b7d7e132
- anthropic.com/news/accenture-embedded-evaluation
- gov.ca.gov/2026/09/18/governor-newsom-issues-executive-order-to-accelerate-independent-oversight-and-advance-the-creation-of-an-ai-kill-switch
- aws.amazon.com/blogs/security/run-open-weight-models-on-aws-bedrock-in-aws-european-sovereign-cloud
- mlcommons.org/2026/09/mlperf-inference-v6-1-results
- csa.gov.sg/alerts-and-advisories/alerts/al-2026-124
- github.com/vitejs/vite/security/advisories/GHSA-v2wj-q39q-566r
- blog.google/innovation-and-ai/models-and-research/google-labs/cc-expanding-to-groups
- cloud.google.com/blog/topics/systems/using-ai-agents-to-secure-google-infrastructure
- github.com/google/mantis