All episodes

AI & Tech Daily

AI Agents Reach the Lab, but Human Control Still Sets the Pace

18:11

OpenAI says its internal agents have reached an automated research intern milestone, but frequent human intervention and rising compute use show where autonomy still stops. Jesse also covers a proposed US superintelligence ban, New York City's classroom AI moratorium, a vast identity-document exposure under FBI investigation, a Sydney AWS-Azure private link, fresh publisher litigation, Microsoft's new transcription model and AWS Lambda diagnostics for coding agents. Jesse's Medium article, Stop Letting AI Coding Agents Guess: https://medium.com/@jesse_19696/from-business-requirement-to-working-code-c912b588a291

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

The Automated Research Intern

OpenAI says its agents now produce 3.1 agent-workdays for every human workday. More than half of their successful longer tasks still need a person to step in.

The useful part of this claim is the gap between output and autonomy. On 6 September, OpenAI said its internal systems had met the company's own definition of an automated research intern: an agent that can complete a well-defined research task, under human direction, that would take a skilled researcher several days.

That is a meaningful operational threshold. A research group doesn't need an artificial scientist with its own agenda to change how much work it can attempt. Give an agent a bounded experiment, an engineering task or an analysis job, and it can run while the researcher works on something else. OpenAI says its research organisation reached 3.1 agent-workdays for every human workday by the middle of August. Its experiments per active experimenter also hit the highest level recorded since tracking began in January 2025.

The constraint shows up in what happened during those runs. Among successful tasks that ran for four to eight hours, more than half required at least one human intervention. The person might still have to redirect the work, correct a mistake or supply missing judgement. OpenAI also says greater compute contributed to the higher experiment rate. More agent work means more inference, and inference is a real operating cost.

The intervention rate also changes how to read the 3.1 figure. Agent-workdays describe activity, not three extra human colleagues arriving ready to make decisions. If a majority of successful longer runs still interrupt a researcher, parallel work can create a new queue of reviews and course corrections. That may still be a worthwhile trade when an agent handles hours of bounded work between interventions. The gain depends on how well tasks are specified and how quickly a qualified person can respond.

There are two limits on how far to take the result. These are OpenAI's preliminary internal measurements, not independently audited results. And agent runtime or experiment count isn't the same thing as scientific progress. A laboratory can run more experiments without every experiment being useful, and evidence from one highly technical organisation doesn't tell us how the same systems perform in a different research environment.

Even with those limits, the management question has changed. Research leaders can plausibly plan for more work in parallel, but they can't budget as though the agents are independent staff. They need enough compute, clearly bounded jobs and researchers with time to supervise. Safety controls also have to operate at the speed of the extra experiments, not the old human-only pace.

My read is that governance may become the bottleneck before model capability does. If one researcher can launch several days of agent work, the scarce resource becomes informed attention: deciding what to run, noticing when a path has gone wrong and checking the result before another system builds on it. The automated research intern is useful now. It is still an intern.

A Proposed Ceiling on AI

That leaves the lab with a supervision problem. In Washington, two lawmakers are arguing for a much harder limit.

US Senator Bernie Sanders and Representative Greg Casar announced on 3 September that they plan legislation to permanently prohibit artificial superintelligence and pause advanced AI development until a new federal regulator has established safety rules. Their proposed framework would create a cabinet-level AI regulator able to monitor frontier systems and supervise the removal of dangerous capabilities. It also calls for international coordination and export controls intended to prevent superintelligence development outside the United States.

There is no new law or compliance deadline. The full bill text hadn't been released or formally introduced when the proposal was reported, so key definitions and enforcement details remain unknown. It is also unclear whether the measure can attract enough support in Congress. A press-release outline tells us what its sponsors want, not the exact boundary a laboratory would face.

One unresolved term is artificial superintelligence itself. Until the text appears, we don't know how the sponsors would distinguish a permanently prohibited system from advanced AI covered by the temporary pause. That distinction affects which developers, models and research programs could fall inside the framework. The regulator's process for identifying and removing dangerous capabilities is also still only an outline.

Still, the framing is notable. Much of AI policy focuses on uses: discrimination, privacy, copyright, safety in a particular sector. This proposal starts with a capability ceiling. It says some systems could be too powerful to build, regardless of the application offered at launch.

For frontier laboratories and investors, my assessment is that the immediate burden is political uncertainty rather than compliance. They now have an explicit US proposal aimed at stopping a class of development, and its unresolved terms make long-range planning harder. Whether the bill advances or not, capability limits are now part of the live policy argument around increasingly autonomous systems.

New York Draws an Age Line

The control question looks very different when the users are children. New York City has drawn a firm age line for one school year.

New York City Public Schools adopted a moratorium for the 2026–27 year that blocks student-facing generative AI from pre-kindergarten through eighth grade. High-school use is narrower rather than unrestricted: students first complete two 45-minute AI-literacy modules, then they can use one of five centrally approved pilot programs under teacher supervision. Companion chatbots are prohibited at every grade level.

The policy leaves room for assistive technology and specified accessibility uses. Teachers can use approved AI for planning and operations, but not for grading, behaviour monitoring or decisions about whether a student progresses. That separation matters because it distinguishes a teacher using a tool to prepare work from a system making consequential judgements about a child.

This is the largest US school system choosing a year of controlled experimentation over routine access for younger students. Families and teachers get clearer boundaries during the new school year, while education software vendors face a practical test: a product being available doesn't make it approved for a classroom. High-school pilots also make AI literacy a condition of access rather than an optional lesson after students begin using the tools. Vendors now have to fit centrally approved pilots and age rules rather than relying on a single product experience across the school system.

The policy isn't a permanent answer. It runs for one school year, and the city hasn't decided what follows the pilots and evaluation. My view is that other education systems now have a concrete policy model to examine: keep younger students away from general-purpose generative tools, teach older students how the systems work, and limit their use to supervised programs while evidence develops. That may make classroom adoption slower, but it gives schools a clearer way to learn without treating every age group as the same.

The Cost of Storing Identity

Some data risks don't disappear when a service goes offline. A licence scan can remain useful to a criminal for years.

Security journalists have verified authentic identity documents inside a newly launched criminal service, and the FBI has opened an investigation into an apparent breach involving identity-verification provider IDScan. The service claimed more than 153 million US and Canadian driver-licence records, along with other identity documents.

That total hasn't been officially confirmed. KrebsOnSecurity did, however, verify multiple records with the people named in them and observed roughly 11.5 million result pages, each showing about 15 entries. IDScan said it was investigating, but it hadn't confirmed that it was the source or provided a substantive public explanation. The criminal site later went offline. That is welcome, but it doesn't show whether copies were recovered or destroyed.

A password breach is serious, but a person can change a password. A driver licence contains durable identity details and often a photograph, address and date of birth. Depending on what was taken, the risk can include identity fraud and physical safety concerns. At this point there is no authoritative list of affected people, so the scale, the suspected source and individual exposure all remain unresolved. People cannot yet determine from an official notice whether their own documents are among the records. That uncertainty is part of the harm.

The organisational lesson is uncomfortable. Expanding identity and age checks can produce central stores of documents that are exceptionally valuable to an attacker. My judgement is that any organisation collecting those scans needs to treat retention as part of the security design, not an administrative afterthought. Keeping less data, for less time, and isolating verification records can reduce what a single breach exposes. The investigation still has to establish where these records came from, but verified documents appearing in a criminal search service are already evidence of the harm created when identity checks turn into a long-lived data store.

A Private Sydney Cloud Link

For Australian cloud architects, a quieter infrastructure change could remove a fair amount of network plumbing.

AWS and Microsoft have opened public-preview interoperability between AWS Interconnect and Azure Multicloud Interconnect. One of the four available region pairs connects AWS in Sydney with Azure Australia East. The service provides managed private networking between selected regions across the two cloud platforms.

Building that link has traditionally involved customer-managed circuits, routers and routing sessions. In the preview, the providers coordinate more of that provisioning. An organisation running applications or data across both clouds can test a private connection without assembling the entire cross-cloud network path itself.

There are important boundaries. Azure caps the preview connection at one gigabit per second and provides no service-level agreement. Customers need accounts with both providers, an Azure ExpressRoute virtual network gateway and network address spaces that don't overlap. General-availability timing, final pricing, higher bandwidth options and real-world operational performance haven't been announced.

That makes this useful for evaluation, not an automatic home for a production workload that depends on a contractual uptime target. It may simplify the work for network engineers, but the architecture still spans two providers, two sets of controls and two commercial relationships. Without an SLA, architects have no preview commitment to use as the basis for workload availability. The one-gigabit ceiling also puts a known limit around the designs that can be tested. Both constraints need to sit beside the convenience of provider-managed provisioning.

For Australian organisations already committed to both hyperscalers, my take is that a managed link makes a multicloud design easier to test and potentially easier to operate. The trade-off is deeper dependence on both control planes at once. The preview removes some physical and routing complexity; it doesn't remove the need to understand failure modes, address planning or which provider owns each part of an incident.

Publishers Add Another Copyright Case

The legal boundary around AI training is also being contested one publisher at a time. Two more have gone to court.

The Seattle Times and Newsday filed a federal complaint on 4 September against OpenAI and Microsoft. They allege that their journalism was copied without permission or compensation for model training, retrieval and generated outputs. The complaint includes copyright claims, allegations involving copyright-management information, trademark claims and related claims under state law.

The publishers also allege that the companies bypassed paywalls and produced outputs that substitute for their work. Those are allegations. They haven't been tested in court, and the defendants hadn't filed substantive responses when the briefing was prepared. The case joins existing publisher litigation; filing it doesn't create a new legal precedent.

For AI companies and news organisations, the practical effect is another active dispute hanging over how journalism is acquired and used. Each additional action raises the commercial pressure for clearer licensing arrangements, even while courts are still working through the underlying legal questions. My assessment is that legal uncertainty itself is becoming a cost: publishers have to decide whether to license, litigate or try both, while model providers have to price in the possibility that current data practices face tighter limits.

Microsoft's New Transcription Option

There is a more immediately testable release for developers working with recorded speech. Microsoft's latest model is now in public preview.

MAI-Transcribe-2 is available through Azure Speech in Microsoft Foundry. It handles 60 languages and includes speaker diarisation, which separates speakers in a recording, along with word-level timestamps, keyword biasing and a choice between clean and verbatim output. Microsoft set an introductory price of ten US cents per audio hour through 31 December 2026. Pricing after that date hasn't been disclosed.

Microsoft describes the model as the world's fastest, most accurate and cheapest speech-recognition option. The independent picture is more restrained. Artificial Analysis currently measures a 2.0 per cent word-error rate and ranks it second for accuracy on its non-streaming test set. That is a strong result, but it doesn't support every part of Microsoft's unqualified claim.

For developers building captions, meeting tools, call analysis or voice-agent systems, the combination of low introductory pricing, timestamps and speaker metadata makes the model inexpensive to evaluate. Public preview also means behaviour and production terms may change.

My recommendation is to compare it with real audio from the intended workload. Accents, background noise and specialist vocabulary can matter more than a leaderboard average. The new model could make high-volume multilingual transcription cheaper to test now; choosing it for production depends on results from the recordings your users actually produce.

What Changes for You

One AWS release gives coding agents a useful new job today, with a sensible boundary around what they can touch.

AWS has added serverless diagnostics to its managed MCP Server. MCP is a standard way for an AI tool to call structured external capabilities. A compatible coding or operations agent can now inspect a Lambda function and its connected AWS resources, gathering an incident picture through read-only tools.

The capability can compare current metrics with a seven-day baseline and retrieve logs, deployed configuration, recent changes and summaries of X-Ray traces. It covers connections to API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions. That can replace the manual job of opening several services and assembling the same evidence before diagnosis begins.

It works within the caller's AWS account and remains read-only, so the agent can investigate without gaining permission to change production. The managed server runs in Northern Virginia and Frankfurt, although it can inspect resources across commercial AWS regions. X-Ray evidence is available only when tracing was already enabled. AWS also hasn't published independent evidence showing how accurately the rule-based diagnosis identifies real root causes.

For working AWS developers, this makes agent-assisted incident triage more practical now. The best part of the design is the permission boundary: diagnosis becomes easier without bundling it with remediation. You still need the relevant AWS access, an MCP-compatible agent and enough operational judgement to check the assembled evidence.

I've written a Medium article called Stop Letting AI Coding Agents Guess. A detailed business conversation can somehow arrive at a developer's desk as three lines in Jira. That leaves the coding agent to fill the gaps, and plausible isn't the same as correct. My recommendation is to keep the full requirements where the BA records them in Confluence, link that page from the assigned Jira work, and give the agent that business context alongside the existing repository and a targeted prompt. Spec Kit then gives the work a useful shape: specification, planning and tasks.

Take notification controls. If the requirement doesn't say who may pause notifications, the model can invent a tidy rule that looks perfectly sensible in code. The developer needs to stop there and clarify that business rule with the owner, then update the source context before implementation continues. The point isn't to make the prompt enormous or expect perfect output. It's to make sure the agent sees the same approved requirement the developer sees, and to use structure where structure helps. Three Jira lines may be enough to assign the job. They aren't enough to explain it. Silence is inference.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. openai.com/index/research-acceleration-view-inside-openai
  2. sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development
  3. axios.com/2026/09/03/bernie-sanders-superintelligence-ban-ai-pause
  4. schools.nyc.gov/about-us/policies/guidance-on-artificial-intelligence
  5. apnews.com/article/647f6a968eea0399521b7934418b1aff
  6. krebsonsecurity.com/2026/09/fbi-probes-service-selling-153m-drivers-licenses
  7. techcrunch.com/2026/09/02/it-sure-looks-like-hackers-breached-a-major-id-card-verification-service
  8. aws.amazon.com/about-aws/whats-new/2026/08/aws-announces-AWS-interconnect-multicloud-microsoft-azure-preview
  9. learn.microsoft.com/en-us/azure/multicloud-interconnect/availability-limits
  10. itpro.com/cloud/cloud-computing/multi-cloud-with-aws-and-azure-just-got-a-whole-lot-easier-thanks-to-a-new-interconnect-service
  11. storage.courtlistener.com/recap/gov.uscourts.nysd.672142/gov.uscourts.nysd.672142.1.0.pdf
  12. techcrunch.com/2026/09/05/seattle-times-and-newsday-are-the-latest-publications-to-sue-openai-and-microsoft
  13. microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world
  14. artificialanalysis.ai/speech-to-text/non-streaming
  15. aws.amazon.com/about-aws/whats-new/2026/09/aws-mcp-server-serverless
  16. docs.aws.amazon.com/agent-toolkit/latest/userguide/capabilities.html