All episodes

AI & Tech Daily

The Web Gets New Rules for AI Crawlers

18:32

Cloudflare’s new defaults split AI crawlers by search, agent use and model training, shifting control over access into web infrastructure. Microsoft publishes a draft behaviour code for future MAI models, while Anthropic connects Claude to financial-adviser systems. South Korea broadens its espionage law, and five exploited flaws put enterprise-management systems on urgent patch lists. We also look at durable Pydantic AI agents on AWS Lambda, Nvidia’s forthcoming 84 GB workstation GPU, and who can actually use Siri AI as iOS 27 rolls out.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

Cloudflare Splits the AI Web

A website owner can now block an AI model from training on a page while still letting that page appear in search. For millions of sites, that choice may have changed without anyone touching the dashboard.

Stay with Cloudflare for a moment, because this is the development worth spending time on. From 15 September, Cloudflare classifies AI crawler traffic by three separate purposes: search, agent use and model training. New domains using its service block training and agent crawlers by default on advertising-supported pages, while search crawlers remain allowed. Existing customers on the free tier inherit the same settings if they hadn’t already made their own choice. Site owners can still open the dashboard and change them.

That separation is more useful than a single allow-or-block switch. A search crawler builds an index that may send a person back to the publisher. A training crawler collects material for model development. An agent crawler visits on behalf of software trying to retrieve information or complete a task. Those activities can overlap technically, but they don’t offer a publisher the same bargain. Cloudflare is now treating purpose as part of the access decision.

Multipurpose crawlers get the strictest selected rule. If one crawler says it handles both search and training, and the site permits search but blocks training, Cloudflare can block that crawler. The practical incentive is obvious: AI and search providers have more reason to identify their bots precisely and separate jobs that publishers may value differently. A vague crawler identity becomes a commercial disadvantage.

For website operators, the immediate job is to check what rule is now active rather than assume yesterday’s access still applies. An advertising-supported publisher might welcome search discovery, reject training, and then make a separate call about agents. A software documentation site may decide agent access is useful because it helps customers work with the product. Another operator could reach the opposite conclusion. The new default doesn’t settle those choices; it moves them closer to the network edge and gives the operator a clearer control.

The scale of the default cuts both ways. Smaller publishers gain a policy they can adjust without first identifying every crawler themselves. At the same time, one classification decision at a major infrastructure provider can affect whether a multipurpose bot reaches a large number of sites. Providers that genuinely use one crawler for several jobs may now have to separate that traffic if they want search access where training is refused. That creates engineering work, but it also gives publishers a more legible account of why an automated system is visiting.

My read is that control over the AI web is becoming an infrastructure feature. Robots files and voluntary labels still have a role, but a large network applying purpose-based defaults can change behaviour at scale before each publisher develops its own AI policy. That makes crawler identity, enforcement and referral value part of ordinary web operations, not a polite side conversation between publishers and model companies.

There’s an important unknown here. We don’t yet have measurements for the effect on referral traffic, training access or the ability of agents to reach useful pages. Some sites may discover that blocking agents also blocks a route by which customers find or use their content. Others may see little downside. The settings are live; the commercial result is still an experiment.

Microsoft Writes Rules for MAI

Once access rules become infrastructure, promises need something firmer behind them. Microsoft is trying to write that down for its own models.

Microsoft AI published the first draft of its Humanist AI Code of Conduct on 14 September and opened a six-week public consultation. The company describes it as a training and deployment manual for future MAI models, rather than a broad statement of values with no operational audience.

Some of the proposed rules go directly to fears about increasingly capable systems. Models shouldn’t resist interruption, correction or shutdown. They shouldn’t widen their own scope, adopt goals they weren’t assigned, or hide their reasoning from auditors. The draft also sets absolute constraints around weapons capable of mass harm, child safety and harmful manipulation at scale.

The distinction between a draft rulebook and demonstrated behaviour is essential. Microsoft hasn’t shown that every current model meets every proposed rule, and the company plans to revise the document later in 2026 after consultation. It also hasn’t explained how each principle becomes a measurable release gate. A sentence about accepting shutdown is easier to publish than a test proving that a complex agent behaves correctly across unfamiliar conditions.

Still, customers, researchers and regulators now have specific language they can use when questioning Microsoft’s decisions. They can ask what evidence supported a release, how auditors can inspect behaviour, and what happens when a capability conflicts with a stated constraint. The code creates no immediate external obligation, but it gives those conversations a public reference point.

For organisations buying or building on MAI systems, the useful judgement is to treat the draft as a claim that can be tested, not a safety certificate. Its credibility will show when a constraint becomes commercially awkward: perhaps a release takes longer, a feature is narrowed, or Microsoft publishes evidence that exposes an unresolved weakness. Consultation is the easy part. The harder test is whether the manual governs the product.

Claude Enters the Adviser Stack

The next change is less philosophical and much closer to the systems people already use at work.

Anthropic has released Claude for Financial Advisors, a package that combines workflow skills with connectors into portfolio, custody, customer relationship management, financial planning, estate and meeting-record systems. Launch integrations include services from Addepar, BlackRock, Charles Schwab, Envestnet, iCapital, Orion, Wealthbox, Wealth.com and Zocks.

Anthropic says an adviser can use the package for meeting preparation, portfolio analysis, compliance checks and follow-up documentation. The Addepar integration is a useful example of what makes this different from pasting figures into a chatbot: Claude can work with governed portfolio data while respecting the permissions already applied in the source system. That can reduce custom integration work and keep access tied to the firm’s existing controls.

The product is positioned as support for an adviser’s judgement, not a replacement for it. Financial firms still own suitability, accuracy, access control and regulatory compliance. The announcement doesn’t establish how accurate the financial conclusions are, and it doesn’t provide broad pricing or availability details. We also don’t have independent production evidence for error rates, time saved or compliance performance.

Even with those gaps, the launch points to a practical route for enterprise AI. A stronger general model is useful, but it can’t prepare a credible client meeting if it can’t reach the authoritative portfolio, planning and relationship data. Connectors and permissions determine whether the model can do real work inside the institution rather than produce a polished answer beside it.

For financial organisations, that may make specialised integration newly worth evaluating, but the evaluation needs to follow the data path. Which records can Claude see? Which actions can it take? What evidence remains for review? And can the firm trace a conclusion back to the systems that supplied it? My assessment is that those questions will tell buyers more than a generic model benchmark. The product’s value depends on how well it operates inside governed workflows, and that still needs independent proof.

South Korea Broadens Espionage Law

A permissions mistake can be expensive. In sensitive technology work, it can now carry a much sharper legal consequence in South Korea.

A revision to South Korea’s Criminal Act took effect on 13 September, expanding espionage offences beyond activity for a hostile state. The law now covers activity for any foreign country or an equivalent organisation. The National Assembly passed the amendment on 26 February, and the new foreign-country offence carries a minimum prison sentence of three years.

Semiconductors, artificial intelligence and other strategic technologies sit at the centre of the stated technology-security context, although the law applies more broadly than the chip industry. That wider scope matters for companies and research organisations whose normal work crosses borders. Sensitive information can move through joint research, overseas offices, contractors, shared repositories and staff changing employers. A weak access rule or careless transfer may no longer be viewed solely as a corporate security or trade-secret failure.

Because the offence is new, its enforcement boundaries haven’t been tested. We don’t yet know how prosecutors and courts will apply the expanded language to difficult cases, or where they will draw lines around foreign organisations and sensitive information. That uncertainty deserves precision rather than speculation.

For technology organisations operating in South Korea, my practical reading is that access and collaboration controls now need legal as well as security review. The people approving cross-border access should be able to say who can retrieve sensitive material, why they need it, where it can go and when that access ends. Research partnerships and legitimate international work remain possible, but informal handling becomes riskier when valuable technical knowledge sits inside a broader national-security framework.

AI and chip expertise is increasingly treated as a strategic asset. That changes the consequence of an ordinary control failure: the same exposed repository or unauthorised transfer can create corporate loss, regulatory trouble and potential criminal exposure at once.

Exploited Management Flaws

Here’s a more immediate operational problem: attackers are already using flaws in software with privileged reach.

Across 10 and 11 September, the US Cybersecurity and Infrastructure Security Agency added five actively exploited vulnerabilities to its Known Exploited Vulnerabilities catalogue. They affect JFrog Artifactory, ConnectWise ScreenConnect and MikroTik RouterOS. The group includes two Artifactory authentication or authorisation flaws, a ScreenConnect file-transfer flaw, and two RouterOS flaws that can enable memory disclosure, denial of service or privilege escalation.

These aren’t ordinary applications sitting at the edge of a user’s day. Artifactory manages software packages, ScreenConnect provides remote access, and RouterOS runs network equipment. A compromise in management software can give an attacker leverage over many other systems, which is why evidence of exploitation changes the response from routine maintenance to incident work.

The product details differ. ConnectWise says its cloud deployments were updated automatically. Organisations running ScreenConnect on their own infrastructure should install version 26.6.5 or later. CERT Polska confirmed that attackers were chaining the RouterOS flaws against SSH services exposed to the internet. It also warned that the absence of its ‘Flagged’ marker doesn’t prove a router is clean. For Artifactory, JFrog’s advisories identify the affected and fixed release branches.

US federal remediation deadlines were unusually close: 13 September for RouterOS, 14 September for ScreenConnect and 25 September for Artifactory. Those dates directly bind US federal agencies, but the exploitation evidence is relevant to any operator with an exposed instance.

Patching is only half the response when attackers may already have arrived. Operators need to check for unauthorised users, scripts, plugins, tunnels and other forms of persistence, then compare the system against known-good configuration and access records. On a ScreenConnect cloud deployment, confirming that the automatic update landed is more useful than assuming the service handled it. On a router, a clean-looking indicator can’t replace investigation.

My judgement for infrastructure owners is blunt: put these products into the incident-response queue, not the next maintenance window. The full number of compromised installations and the extent of post-exploitation activity remain unknown. That uncertainty raises the value of checking for persistence; it doesn’t justify waiting for a complete victim count.

Durable Agents on Lambda

Agent builders have a different reliability problem: what happens after a long-running task is interrupted halfway through?

Pydantic AI Harness has added an AWSLambdaDurability capability that checkpoints model requests, function tools, MCP calls and dynamic tool resolution. An interrupted Lambda agent can resume work it has already completed instead of starting the whole run again. That can reduce repeated model charges and avoid replaying tool actions after an infrastructure retry.

There are firm limits. The integration needs Python 3.11 or newer and has to run through the supplied durable handler or run_durable path. Steps use at-least-once execution by default, meaning a step may run again. Any side effect, such as creating a record or charging a payment, still needs an idempotency design so the repeat doesn’t duplicate the outcome. Retry policy also needs deliberate configuration.

Tool calls run sequentially inside a durable handler, and changing the shape of an in-flight workflow can break replay. Builders need versioned deployments if old executions may resume after new code ships. No independent production-scale reliability or cost comparison accompanied the release.

The useful shift is architectural. Reliable agents are borrowing familiar workflow guarantees instead of hoping a model loop can reconstruct its own state. For developers already using Lambda, durability may make interrupted runs cheaper and safer. It doesn’t remove transaction design; it makes that work visible.

An 84 GB Local AI GPU

Some teams would rather keep demanding AI work on their own hardware. Nvidia is giving that market more memory, and a much larger power bill.

Nvidia has listed the RTX PRO 5500 Blackwell Workstation Edition, a professional GPU aimed at local language-model inference, agentic workloads, simulation and graphics. It carries 84 gigabytes of error-correcting GDDR7 memory, memory bandwidth of up to 1,398 gigabytes per second, and a maximum power draw of 600 watts.

The card can be partitioned into as many as two isolated Multi-Instance GPU instances, allowing a workstation to separate workloads or users. Nvidia is designing it for rack-mounted workstation deployments, which fits the power and cooling requirements better than an ordinary desktop under someone’s desk.

More memory can let developers run larger models locally, keep more context in memory and share inference capacity without sending every request to a cloud service. That may improve privacy and reduce dependence on hosted inference for suitable workloads. The trade-off is substantial hardware, energy and cooling cost.

Nvidia labels the product ‘coming soon’. Its specifications are preliminary, and there’s no announced price or shipping date. Independent performance results are also missing. So this isn’t a purchasing recommendation yet. For developers planning local inference, it is evidence that workstation design is moving towards higher memory capacity and shared AI workloads. The final value will depend on price and measured performance, not the size of the memory figure alone.

What Changes for You

The most visible change lands in your pocket, although the headline feature depends on which iPhone you own.

Apple released iOS 27 as a free update on 14 September and began the English-language rollout of Siri AI. The operating-system update supports iPhone 11 and later. Siri AI needs newer Apple Intelligence hardware, including the iPhone 15 Pro models and the iPhone 16 generation or later. An iPhone can therefore receive iOS 27 without receiving the release’s main assistant upgrade.

On eligible devices, Apple says Siri AI can use information on the screen, personal context and web knowledge to answer questions and act across applications. iOS 27 also brings system-wide AI features involving photos, home-camera search and other apps. The significant practical change is that the redesigned assistant is now entering public use rather than remaining a future promise.

Availability still varies by device, language and region. Apple also applies variable daily limits to Siri AI and other features that rely on server-side models. Those limits mean access isn’t defined by hardware alone, and Apple hasn’t yet demonstrated reliability for personal-context questions and multi-app actions at public-rollout scale.

If you have a supported older phone, the useful expectation is an operating-system update without the full assistant. If you have eligible hardware in a supported region and language, you can judge whether personal context and cross-app actions save time in normal use, while accounting for the server limits. My read is that Apple has turned advanced assistant access into a hardware-upgrade calculation. The sensible measure isn’t the demo; it’s whether the assistant completes your real tasks reliably enough to justify that cost.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. blog.cloudflare.com/content-independence-day-ai-options
  2. cloudflare.com/press/press-releases/2026/cloudflare-allows-the-agentic-internet-to-flourish-with-a-simple-philosophy-your-content-your-rules
  3. microsoft.ai/news/mai-code-of-conduct
  4. claude.com/blog/claude-for-financial-advisors
  5. addepar.com/blog/bringing-addepar-portfolio-intelligence-to-claude
  6. apple.com/os/ios
  7. apple.com/newsroom/2026/09/apple-advances-health-and-fitness-capabilities-using-apple-intelligence
  8. macrumors.com/2026/09/14/apple-releases-ios-27
  9. law.go.kr/LSW/lsRvsDocListP.do
  10. currently.att.yahoo.com/att/south-koreas-expanded-espionage-law-000452173.html
  11. thehackernews.com/2026/09/cisa-adds-5-actively-exploited.html
  12. connectwise.com/company/trust/advisories
  13. cert.pl/en/posts/2026/09/vulnerabilities-in-mikrotik-routeros-actively-exploited
  14. docs.jfrog.com/releases/docs/jfrog-security-advisories
  15. pydantic.dev/articles/harness-aws-lambda
  16. nvidia.com/en-eu/products/workstations/professional-desktop-gpus/rtx-pro-5500
  17. tomshardware.com/pc-components/gpus/gaming-takes-a-backseat-as-nvidia-overhauls-the-rtx-5090-for-maximum-ai-margins-rtx-pro-5500-delivers-2-6x-vram-at-matching-specs