All episodes

AI & Tech Daily

AI Labs Put Outside Evaluators Inside the Frontier Race

19:35

Anthropic has promised independent AI safety evaluators employee-like access, and OpenAI has reportedly agreed to the same principle. Jesse examines what would make that oversight credible, then covers California’s child-safety rules for companion chatbots, critical Check Point VPN fixes, Intel’s first million High-NA EUV wafer passes, the Google Cloud–Accenture deployment push, smarter Copilot reviews and AI-branded phishing. The final section explains what Google’s generally available Agent Development Kit for Kotlin changes for JVM and Android developers.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

Outside Eyes Inside AI Labs

Two leading AI labs say independent safety evaluators should work almost like employees. The hard part now is turning that unusual access into oversight people can trust.

That’s the development worth spending some time on, because the promise is far more concrete than another public statement about responsible AI. On 12 September, Anthropic chief executive Dario Amodei published a three-part proposal for slowing unchecked gains at the frontier of AI capability. Its first step is unilateral: Anthropic says it will give outside evaluators ongoing access to inspect safety practices, training pipelines, models and incidents.

The proposed arrangement goes well beyond inviting researchers in for a carefully managed model test. Amodei says evaluators should receive office access, company equipment and permissions broadly comparable to those held by an internal risk team. They could build a continuous view of how the lab operates rather than judging one release from a snapshot supplied by the company. Over time, that continuity could reveal how exceptions are handled and which concerns survive contact with real release deadlines. Anthropic also says the access would need safeguards for customer information, legal privilege, security and commercially sensitive material.

The Associated Press reported that OpenAI chief executive Sam Altman then committed OpenAI to the same employee-like access principle. emission gives the idea more weight, but it doesn’t yet amount to an industry system. Anthropic hasn’t named the evaluators, set a start date or published the full terms. OpenAI hasn’t supplied those operating details either. Common pacing standards, government coordination and any rules for the wider industry remain proposals.

There’s another important distinction in Amodei’s argument. He forecasts that much more dangerous swarms of AI agents could arrive within six to twelve months. That is his assessment of the risk, not an independently established timeline. You can accept the case for better scrutiny without treating one lab chief’s forecast as a settled fact.

For organisations buying or deploying frontier models, embedded evaluators could make safety claims more useful. An evaluator with persistent access may be able to test whether internal controls survive the pressure of a release schedule, whether incidents change practice and whether the published account matches the actual process. A polished safety report written after the decisions were made can’t offer the same view.

My test is fairly simple. Who are the evaluators? What can they inspect? Can they choose their own tests? What are they allowed to publish, including when the lab disagrees? And what happens if they find a serious problem before release? Until the labs answer those questions, this is a promising voluntary commitment rather than enforceable oversight.

That’s also the weakness in the wider pacing proposal. If one company slows while its rivals keep racing, commercial pressure can overwhelm good intentions. Embedded evaluation could make individual claims more verifiable, which is worthwhile. Without shared standards or enforcement, it may leave the competitive incentives largely intact. The next proof point isn’t another pledge. It’s whether outsiders receive enough independence to affect a real release decision.

Child-Safety Rules for Companion Bots

That leaves the labs with a credibility test. California has taken a more conventional route: putting child-safety duties into law.

The governor signed a package on 10 September that includes stronger protections for children using companion chatbots. SB 1119 is the measure focused on companion-chatbot safety, while SB 867 covers toys that include companion-chatbot functions. The official announcement says providers will need crisis protocols for suicidal ideation, parental controls and notifications when a child disables safety settings. It also calls for independent child-safety audits and annual risk assessments.

Those are significant obligations because companion products are designed for continuing interaction, not an occasional search or one-off question. A service can become part confidant, part entertainment and part advice channel. When the user is a child, failures around emotional dependence, crisis language or hidden safety settings carry a different weight. An annual risk assessment forces a provider to examine those risks as an operating issue, while an independent audit gives regulators and families something more substantial than the company grading itself.

The wider package also addresses child privacy, school data and addictive features in social media. Those provisions are part of the same political response, but they don’t prove that a chatbot crisis protocol or parental dashboard will work. The announcement also doesn’t provide a complete timetable for every requirement, so families shouldn’t assume all of these controls appear immediately.

For providers serving children in California, the practical work is now quite specific: build controls that can be audited, record the risks,ly, notify parents when relevant safeguards are switched off and prepare for crisis handling that stands up outside the company. Families gain clearer rights once the provisions take effect.

I think a legal safety baseline is better than leaving every service to choose its own threshold. The trade-off sits in the implementation. Age checks and parental controls can create privacy problems or lock legitimate users out when they’re clumsy. California hasly has defined the dutiesd the duties; providers still have to show they can meet them without collecting more sensitive information than the protection requires.

Critical Check Point VPN Fixes

From policy, let’s get very practical for anyone responsible for a network edge. Two VPN flaws need urgent attention.

Check Point released emergency fixes on 9 September for CVE-2026-85102 and CVE-2026-85103. Both involve the way affected products process certificates during VPN activity, and both can allow an unauthenticated attacker to run code remotely. CERT-EU gives each flaw a CVSS score of 9.8 and lists affected Security Gateway, Security Management Server and Spark Firewall versions in the relevant VPN configurations.

The first vulnerability concerns validation of certificate data during VPN negotiation. The second is a heap overflow in ASN.1 certificate decoding. ASN.1 is a structured data format commonly used in certificates and security protocols. A decoding bug at that point is especially uncomfortable because the appliance is parsing data supplied during a connection attempt, before an attacker has authenticated.

CheckCheck Point says Live Patch protection began rolling out on 9 September. For systems that aren’t covered automatically, the company recommends installing the latest applicable Jumbo Hotfix. It also says its own researchers found the vulnerabilities and that it had no indication of active exploitation when it disclosed them. That’s useful context, but it isn’t a reason to relax. A vendor’s view at disclosure can’t establish that exploitation has never happened, and Check Point hasn’t publicly described every precondition.

Operators need to inventory the exposed appliances, confirm whether Live Patch is actually active, and apply the hotfix where it isn’t. Internet-facing VPN gateways belong at the top of the queue. This isn’t routine housekeeping: unauthenticated remote code execution on a device that protects the network perimeter can give an attacker exactly the foothold the gateway is meant to prevent.

The sensible editorial call here is decisive patching backed by verification. Don’t mark the job complete because an automatic protection mechanism exists. Check the affected version, the configuration and the installed protection on each appliance, then retain the evidence your incident team would need if exploitation appears later.

High-NA EUV Enters Production

Now for a manufacturing milestone where the headline number needs a little unpacking.

Intel Foundry and ASML say Intel has processed more than one million wafer passes through High-NA EUV systems. High numerical aperture extreme ultraviolet lithography uses improved optics to print smaller features on advanced chips. The important change is that the technology has moved beyond tool certification and research: Intel says it is now using High-NA EUV in high-volume manufacturing on selected layers for a subset of Core Ultra Series 3 processors.

Intel reports that overlay, throughput and availability are meeting its expectations. Overlay is the accuracy with which one patterned layer lines up with another, so it’s critical when a chip is built from many tightly aligned layers. The companies also say comparable Intel 18A layers patterned with High-NA meet or exceed the performance of layers produced with conventional 0.33-NA EUV. Those are company results, but they indicate that the machines are doing production work rather than sitting in a development lab.

The one-million figure isn’t one million commercial processor wafers. It combines qualification, tool testing, research, development and production passes. Intel and ASML haven’t disclosed how that total divides between those categories, how many production wafers were saleable, or what the yield and cost look like. Those missing numbers are essential for judging whether High-NA is economical across more layers and products.

There’s also a practical mask constraint. Intel is using today’s six-inch masks with floor-planning or stitching techniques while the industry works towards a larger six-by-twelve-inch mask standard. That lets engineers learn now, but it shows the surrounding manufacturing system is still catching up with the scanner.

For Intel and possible foundry customers, the advantage is earlier process experience. Engineers get production data on alignment, availability and integration before broader High-NA adoption. My read is that this could become a manufacturing edge, especially if that learning shortens later process ramps. It’s too early to treat the cumulative wafer count as proof of commercial leadership, strong customer demand or attractive economics. The real milestone is narrower and still meaningful: High-NA EUV is now touching selected production layers.

A Thousand Engineers for Gemini

Better models don’t remove the messy work of fitting them into a business. Google Cloud and Accenture are staffing directly for that gap.

The companies launched an Accenture Gemini Enterprise business group on 8 September and plan to establish a one-thousand-person forward-deployed engineering workforce. These engineers are meant to work closely with customers on custom agentic AI systems, alongside Gemini-certified Accenture staff and Google Cloud specialists. The group sits inside Accenture’s existing Google practice.

Forward-deployed engineers are technical staff placed close to a customer’s operation. Their job isn’t limited to configuring a model endpoint. They work through data access, business processes, application integration, governance and the awkward gap between a convincing demonstration and something people can rely on every day. The partners say the group will develop implementation frameworks, reusable industry solutions and dedicated capability centres, with an emphasis on moving deployments into operations.

That tells us something about the present enterprise AI market. Access to a capable model is rarely the whole constraint. A company still has to decide which process is worth changing, connect the right systems, control what an agent can do, measure the result and support it after launch. TechCrunch framed the deal as part of a wider contest to solve those кухня implementation and return-on-investment problems. A target of one thousand engineers is a sizeable bet that services and integration will decide a lot of enterprise spending.

It is still a target. Google Cloud and Accenture haven’t published a staffing timetable, service pricing, independent customer outcomes or a measure of how many projects will reach production.ly. The announcement demonstrates commitment, not delivered value.

For large customers, the new group could make Gemini Enterprise easier to deploy because there’s a named path from product to hands-on implementation. The cost is potential dependence on a tightly joined cloud, model and consulting stack. My view is that buyers should judge the offer by operational outcomes and exit options, not the number of people attached to the partnership. The engineer count is most revealing as an admission: integration, governance and process redesign remain the expensive part.

Copilot Reviews Close the Loop

At a developer’s desk, one small source of review clutter may now clear itself.

GitHub updated Copilot code review on 11 September so a review thread can close automatically when a later commit addresses the underlying feedback. Findings that haven’t been fixed remain open. Copilot can also propose commit messages that reflect the specific change, and the review agent has access to a broader set of shell tools for checking its work. That can include builds includes builds, tests, targeted scripts and available API or tool calls, operating behind GitHub’s Copilot agent firewall.

GitHub also says its Lite review level now uses an ensemble of agents rather than relying on a single pass. The company reports improvements from its own experiments, but the changelog doesn’t provide enough methodology to assess the claimed quality and cost gains independently. The useful part is the observable workflow change, not a vendor performance number.

Automatic resolution could save developers from revisiting threads that are already obsolete. Tool-backed analysis can also produce a stronger finding than a text-only review because the agent may be able to reproduce a failure or run a focused check. That said, a review agent with a wider execution surface needs carefully bounded repository permissions. A successful command inside the review doesn’t replace the project’s own CI, and an automatically closed comment doesn’t make the merge decision.

I’d treat this as a reduction in review administration, with a chance of better evidence attached to findings. The developer and reviewer still need the normal protected checks and their own judgement. If the feature works as described, the nicest result may be mundane: fewer stale comments, less manual tidying and more attention left for the changes that remain unresolved.

AI Brands Become Attack Bait

The popularity of AI services has created a very ordinary security problem: attackers know the names people will click.

Microsoft published threat findings on 10 September describing campaigns that impersonated well-known AI products to steal payment details and credentials or deliver malware. One ChatGPT-themed phishing campaign sent as many as one hundred thousand messages in a single day, according to Microsoft’s telemetry, with the aim of capturing payment-card information. Researchers also saw a Claude-themed adversary-in-the-middle campaign, a fake AI plugin for Windows that delivered Vidar malware, and fraudulent DeepSeek installers distributed through GitHub.

An adversary-in-the-middle attack places the attacker between a user and the real service, often through a convincing sign-in page or proxy. That can let the attacker capture credentials and session information as the victim logs in. Fake installers and plugins take a more direct route: they borrow the excitement around a new tool, then persuade the user to run malicious software.

Microsoft is explicit that these are impersonation campaigns. The examples do not show that the genuine AI services were compromised. That distinction matters when a familiar logo appears in a security alert. The brand may be the lure, not the breached system. The campaign examples and prevalence data also come from Microsoft’s own threat telemetry, so they don’t tell us how representative the activity is across the entire internet or how much the campaigns overlap.

For individuals, the useful habit is to reach billing pages, downloads and plugins through the provider’s known official channel rather than a link in an urgent message. Security teams can reinforce that with domain checks and detection across email, web and endpoint activity. My take is that AI-themed phishing deserves attention, but not a special category of panic. Attackers are applying proven social engineering to products people are curious about. Verify the source, and don’t mistake a convincing brand imitation for evidence that the provider itself has been breached.

What Changes for You

There is one release this week that removes a genuine bit of friction for Kotlin and Android developers.

Google made Agent Development Kit for Kotlin 1.0 generally available on 9 September. ADK is an orchestration framework for building agents that use models, tools, sessions and memory. Kotlin teams can now use a first-party implementation with parity to the ADK 1.0 core, rather than bringing a separate Python agent stack into a JVM application. The core includes hierarchical agents, context compaction, human confirmation flows, resumable sessions and support for long-running tools.

The Kotlin Multiplatform core is designed to remain agnostic about the model, session store and memory provider. Java interoperability matters for existing JVM codebases, while the Android extensions connect with LiteRT-LM, Firebase AI Logic, Room and AppSearch. That opens a path to local or hybrid agents on Android, including workflows that keep some processing or state on the device.

The immediate audience is Kotlin, Java and Android developers who already have the skills and application context to build an agent workflow. Human confirmation and resumable execution are available in the same language and runtime as the rest of the product. That can simplify architecture, deployment and ownership for a team that didn’t want to operate a separate Python service.

There are clear limits. The on-device LiteRT-LM and ML Kit route is still labelled beta, and general availability of the core isn’t independent proof of production reliability. Google’s quickstart currently calls for Kotlin 2.1.20, KSP, a JVM 17 toolchain and a model API key. ADK Web is for development and debugging, not production.

For JVM and Android teams, I think the release makes a production trial newly reasonable, particularly where approval steps and session persistence matter. The distinctive on-device path needs more caution while it remains beta, and the quality of the finished system still depends on the selected model, permissions and deployment controls. The barrier is lower; the engineering work hasn’t disappeared.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. darioamodei.com/post/we-must-pace-the-frontier
  2. apnews.com/article/anthropic-ai-dario-amodei-d59552edcb27892d8ee4d98a48397706
  3. gov.ca.gov/2026/09/10/governor-newsom-signs-the-strongest-child-safety-chatbot-and-social-media-laws-in-the-nation
  4. apnews.com/article/california-social-media-safety-kids-online-harms-6063026d1b54a8537d639605c23aab80
  5. cert.europa.eu/publications/security-advisories/2026-012
  6. community.checkpoint.com/t5/General-Topics/Action-Required-Critical-Security-Advisory-VPN-Vulnerabilities/m-p/282004
  7. securityweek.com/check-point-patches-critical-vpn-vulnerabilities
  8. intel.com/content/www/us/en/newsroom/news/intel-foundry/intel-foundry-asml-accelerate-industry-readiness-for-high-na-euv.html
  9. newsroom.accenture.com/news/2026/accenture-and-google-cloud-deepen-partnership-with-formation-of-new-accenture-gemini-enterprise-business-group
  10. techcrunch.com/2026/09/08/google-cloud-races-to-catch-up-in-the-ai-deployment-wars-with-accenture-deal
  11. github.blog/changelog/2026-09-11-auto-resolution-and-analysis-updates-in-copilot-code-review
  12. microsoft.com/en-us/security/blog/2026/09/10/detect-and-disrupt-ai-themed-attacks-with-microsoft-defender
  13. developers.googleblog.com/announcing-adk-for-kotlin-10-building-production-ready-ai-agents-in-kotlin-android-and-beyond
  14. github.com/google/adk-docs/blob/main/docs/get-started/kotlin.md