AI & Tech Daily
Claude Code Turns Autonomy On by Default
Anthropic is preparing to make Claude Code’s auto mode the default for Pro, Max and Team users, shifting routine permission decisions from people to a classifier. We examine the promised reduction in interruptions, the limits of Anthropic’s safety study and what developers need to check before 14 August. Also: adversarial clothing challenges surveillance detectors, OpenAI acquires NextSlide, Airbnb tests generative travel search, Spotify expands its licensed AI remix effort, Waymo opens Dallas robotaxis to the public, Google Wallet adds supervised balances for US children, and Hark previews an agent for operating websites without APIs.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
Claude Code Defaults to Auto
On 14 August, many Claude Code users will stop approving each action themselves. Anthropic’s classifier will increasingly make that call for them.
It’s worth staying with this development because it changes the normal relationship between a coding agent and the person supervising it. New Claude Code sessions on Pro, Max and Team plans will use auto mode by default unless the user or an administrator chooses another setting.
Auto mode replaces many individual permission prompts with a classifier that approves or blocks proposed actions. The immediate appeal is obvious. Coding agents can lose much of their usefulness when every routine command interrupts the flow and waits for another click. Anthropic is also removing the classifier’s processing overhead from customer billing, so paid users won’t be charged for the extra automated checks.
But this isn’t simply a faster permission screen. The safety boundary itself moves. In manual operation, a developer can see a consequential action, examine the surrounding context and decide whether it belongs. Under the new default, Anthropic’s classifier handles more of that judgment before the developer sees anything. That may make ordinary sessions smoother, while also making mistakes or poorly chosen policies less visible.
There are limits on the autonomy. After three consecutive blocks, or 20 blocks during one session, Claude Code falls back to manual permission requests. Enterprise, API and cloud-platform users remain opted out for now. Anthropic says it plans to broaden the default change next month, but those customers still have time before the setting changes for them.
Anthropic’s supporting safety result sounds striking: in an artificial study involving 1,053 paid professional testers, the company says auto mode caught 89 per cent of dangerous prompts, compared with 13.6 per cent caught by humans. That comparison deserves careful handling. It’s a company-sponsored result from constructed scenarios, not a measurement of incidents inside production repositories. Anthropic has also reported misses in synthetic adversarial testing, so the number doesn’t establish how the classifier handles a genuinely novel attack, a misleading project instruction or an unusual deployment environment.
There’s a human-factors wrinkle here too. Repeated approval prompts can become security theatre when people learn to accept them automatically. A classifier could plausibly outperform a tired developer on familiar patterns. The risk is that teams interpret fewer prompts as proof that fewer dangerous situations exist. They don’t. The intervention has simply moved earlier and become less observable.
My read is that the productivity benefit is credible, but it depends heavily on context. A personal prototype and a repository connected to production credentials shouldn’t necessarily share the same autonomy policy. Teams with sensitive secrets, deployment access or destructive tooling have a short window to review their defaults, deny rules and administrative controls before Anthropic makes the convenient option the normal one.
Clothing Versus Surveillance
That shift from visible approval to automated judgment has a physical-world cousin in surveillance.
Research presented at Black Hat and DEF CON tested adversarial clothing patterns against 11 computer-vision detectors. These patterns are designed to disrupt the models that identify people in images. The work included detector weights extracted from a deployed person-detection system, bringing at least part of the experiment closer to technology operating outside a laboratory.
The researcher reports that every detector was defeated digitally beyond a same-coverage control under at least one tested condition. That’s meaningful, but it doesn’t mean one magic shirt defeated all 11 systems. No single universal pattern has been demonstrated against the entire group at once.
The strongest reported result against the extracted production detector was 61.7 per cent non-detection across the project’s all-garment digital test set. Even that needs a large qualification: the tests used simulated printing and cameras with white-box access to the models. They weren’t trials of physical garments walking past live surveillance cameras. The results are self-published and haven’t been independently replicated.
So, no, this isn’t reliable invisibility clothing. Fabric, movement, lighting, camera angle, compression and model updates could all make the physical problem substantially different from a digital simulation. A person shouldn’t assume a printed pattern provides dependable protection from identification.
For surveillance operators, though, the security lesson is already useful. A detector’s confidence score is an attackable signal, not an objective record of what a camera observed. Systems that treat one model result as decisive can fail even when the underlying scene looks ordinary to a person. My assessment is that computer-vision operators should plan for deliberate evasion in the same way other security teams plan for hostile inputs: use layered controls, preserve review paths and avoid giving one automated detector more authority than its evidence deserves.
OpenAI Buys NextSlide
The next development is less immediate, but worth reading as a signal about where office software is heading.
NextSlide says it has joined OpenAI and that its team is now working on ChatGPT. Reporting describes the transaction as an acquisition completed earlier in 2026, although the exact date and financial terms haven’t been disclosed.
Before the deal, NextSlide built a product that converted prompts, notes, documents and research into editable presentations. That combination matters more than basic slide generation. A useful presentation tool has to organise source material, choose a narrative, create a visual structure and leave the result editable enough for a person to correct. Those are tightly connected tasks, and a specialist team can bring practical product knowledge that a general-purpose model alone doesn’t provide.
Still, the acquisition doesn’t give ChatGPT users a new presentation feature today. Neither company has announced an integration plan, a release date or even confirmed that NextSlide’s former product will reappear in recognisable form. It would be premature to treat the deal as evidence that ChatGPT can replace established presentation workflows.
The more defensible conclusion is about consolidation. OpenAI is assembling specialist capability around end-to-end content creation, while productivity suites are trying to make generation, editing and collaboration feel like one continuous job. For organisations, that could eventually reduce the friction between research and a finished deck. Right now, it mainly increases the pressure on Microsoft, Google, Canva and other productivity platforms to show that their integrated workflows are better than a collection of AI features. Buyers shouldn’t change tools on the strength of the acquisition, but competitors can’t dismiss it as a minor talent deal either.
Airbnb Tests Generative Search
Travel gives us a clearer look at what happens when generative systems meet an existing marketplace.
Airbnb is testing an optional natural-language search experience that produces personalised result titles and highlights alongside visual accommodation listings. Instead of relying only on traditional filters and standard listing text, a guest can describe what they want in ordinary language and see generated framing around the available properties.
For guests included in the test, that could make an awkward or highly specific trip easier to describe. For hosts, it creates another intermediary between what they wrote and what a prospective customer sees. A model may decide which property details best answer a request and present those details in generated language. That can be helpful when it works, but hosts have limited visibility into why one feature was highlighted while another disappeared from the framing.
Airbnb is pairing the search test with broad claims about its internal use of AI. The company says AI has cut its concept-to-launch cycle by as much as 60 per cent and helped it deliver nearly 80 per cent more features and improvements than during the comparable preceding six-month period. It also says 45 per cent of customer issues that begin with its AI agent finish without human intervention, while support cost per booking fell 16 per cent year over year.
Those figures all come from Airbnb. The company hasn’t provided independent evidence that AI caused the reported productivity and cost changes, and other business or staffing decisions could contribute. It also hasn’t specified the search test’s audience, duration or timetable for a wider release.
My concern isn’t that generated travel search is inherently worse than filters. It’s that generated descriptions add an opaque layer to a marketplace where presentation strongly influences bookings. Airbnb may make discovery and support faster, but it also becomes more responsible for how the model interprets hosts and steers guests. If the test expands, useful accountability would include clear labelling, ways to correct inaccurate highlights and enough visibility for hosts to understand how their properties are being represented.
Spotify Expands Licensed Remixes
Consent gets even more concrete when the raw material is somebody’s music.
Spotify has added Merlin to the licensing framework for a planned generative-AI tool that would allow users to create covers and remixes from participating artists’ work. Merlin represents a network of more than 30,000 independent labels and distributors, so its involvement potentially brings a substantial independent catalogue into the discussion.
Spotify says the model is based on artist or rights-holder consent, with credit and compensation built into participation. That approach is notable because many generative-music disputes begin after models or products have already used creative work. Negotiating a framework in advance gives rights holders a clearer place at the table.
There isn’t a finished public product to judge. The tool remains a limited research preview and is intended eventually to become a paid Premium add-on. Spotify hasn’t announced a price, a public launch date, eligible catalogues or the formula that would determine payments. Artists and listeners therefore can’t yet tell whether consent is genuinely granular, how easy it is to withdraw, or whether compensation would be meaningful.
There’s also a difference between licensing a recording and maintaining an artist’s control over identity and context. A technically authorised remix may still create difficult questions about attribution, style and whether an artist wants to appear inside a particular user-created work. Credit helps, but it doesn’t resolve every form of reputational risk.
For independent labels, Merlin’s inclusion creates a possible route into a market that might otherwise favour major rights holders. My view is that advance licensing is a more defensible starting point than releasing a remix tool first and negotiating after the damage. Its fairness, though, can only be assessed once Spotify discloses the payment structure, catalogue rules and practical controls. Consent is the beginning of the design, not proof that the finished system treats creators well.
Waymo Opens Dallas Access
Now, from software deciding what to do to vehicles doing the driving.
Waymo has removed the waitlist for its Dallas robotaxi service. Anyone inside the operating area can download the Waymo app and request a driverless ride without securing an invitation first. That shifts the service from a controlled-access trial into ordinary on-demand transport within its geographic boundary.
Public access is a meaningful autonomy milestone because it changes the customer relationship. A rider no longer needs to be selected as a tester or early participant. The service has to handle the less predictable behaviour, expectations and support needs that come with members of the public treating it as transport rather than a technology demonstration.
The opening remains bounded. Waymo isn’t yet carrying passengers to Dallas Love Field, where it is still testing. Airport trips are a significant practical use for ride services, and airports combine unusual road layouts, kerbside rules, congestion and changing pickup arrangements. Excluding Love Field leaves a noticeable gap even for people who live inside the wider service area.
Weather is another limit. Waymo previously paused service in several cities, including Dallas, during heavy rain and flooding. Suspending operations in dangerous conditions can be the responsible choice, but it means availability still depends on weather and the system’s operating limits. A publicly accessible robotaxi isn’t automatically a universally available one.
Waymo didn’t publish Dallas-specific comparative safety or intervention data alongside the waitlist removal, so broader access shouldn’t be treated as fresh proof of local safety performance. For Dallas residents and visitors, the immediate benefit is straightforward: driverless rides are now something they can request, not merely join a queue to try. The useful way to judge the service is as a real but bounded transport option, with airport coverage and severe-weather resilience still unfinished.
A Wallet Balance for Children
A smaller product change could still matter at the family level—though only in the United States for now.
Google has introduced a Wallet balance for supervised users under 18. Parents can add money, set a daily spending limit, review transactions and remotely lock the balance, giving a child a prepaid-style way to make contactless purchases without a conventional bank account.
The balance works only for in-person tap-to-pay transactions on supported Android or Wear OS devices. It can’t be used for online purchases or cash withdrawals. Parents also can’t require approval for every transaction, and transferred money can’t be withdrawn without closing the balance.
That combination puts the controls somewhere between cash and a full bank product. Parents gain central visibility and the ability to impose a daily ceiling, while children can make eligible purchases without asking at the checkout. The lack of per-purchase approval means families still need to choose a spending limit they’re comfortable delegating.
For eligible US families, the feature could offer useful independence with straightforward guardrails. It also ties family spending more closely to Google’s supervised-account and device ecosystem, concentrating additional payment activity and behavioural data inside one platform. Australian families can’t use it, and Google hasn’t provided a timetable for expansion beyond the United States. The trade-off is convenience and oversight in exchange for platform dependence, not a general replacement for a child’s bank account.
Hark’s Browser Agent Preview
For developers, one more autonomy pitch is arriving through the browser.
Hark has previewed Handoff, a browser-use agent that the company says predicts the next action needed to complete tasks on websites. The proposed attraction is coverage: because the agent operates the site itself, it could automate workflows even when a service doesn’t offer a dedicated API.
That could be valuable for organisations still relying on old administrative portals, supplier sites or internal tools that were never designed for automation. Conventional integrations can require custom connectors or brittle scripts. An agent able to interpret the visible interface might cover more systems with less bespoke work.
At the moment, that promise is substantially ahead of the evidence. Handoff is a preview with a waitlist, not a generally available developer service. Hark demonstrated only part of a task, which prevents an independent assessment of whether the agent completed the workflow reliably from beginning to end. The company is targeting release by the end of the northern summer, but hasn’t disclosed production pricing or detailed security controls.
Browser agents also inherit some awkward risks. Websites change without notice, authentication state can expire, and untrusted page content can try to manipulate an agent. Hark’s preview doesn’t yet provide enough information to assess prompt-injection defences, credential handling or recovery when an interface changes halfway through a job.
The prospect is useful because many important workflows still have no API. But for working developers, Handoff is something to monitor rather than adopt: there’s no broad access, no production cost model and no complete public reliability demonstration. The threshold for production use should be evidence that it fails safely and visibly, not simply that it can click through part of a polished example.
What Changes for You
Before we leave Claude Code, there’s one practical decision worth making ahead of 14 August.
If you use Claude Code on a Pro, Max or Team plan, new sessions are due to begin in auto mode unless you or your administrator select another default. The practical difference is fewer approval interruptions: Anthropic’s classifier can approve or block actions that would previously have waited for your decision, and the processing overhead for those checks won’t be billed.
The relevant limitation is control, not cost. Auto mode can fall back to manual prompts after three consecutive blocks or 20 blocks in one session, but that safeguard doesn’t make every repository equally suitable for the same setting. A local experiment, a work project with deployment access and a codebase containing sensitive credentials present very different consequences if an action is approved incorrectly.
So check which default your account or team policy currently selects, and whether stronger deny rules are appropriate for sensitive repositories. Choosing manual permissions isn’t a rejection of agentic coding; it’s a way to keep consequential actions observable where the cost of a mistake is high. Enterprise, API and cloud-platform users remain opt-in for now, although Anthropic plans a wider default change next month.
The real benefit of auto mode is reduced friction. The unresolved question is whether Anthropic’s classifier can recognise novel attacks and unusual production risks outside artificial tests. Until there’s stronger real-world evidence, the sensible autonomy level is the one matched to the repository’s actual access—not whichever setting happens to arrive as the default.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- claude.com/blog/auto-mode-default-in-claude-code
- techcrunch.com/2026/08/09/anthropic-is-turning-claude-codes-auto-mode-on-by-default
- sandbox.norecognition.org/research
- nextslide.ai
- techcrunch.com/2026/08/08/openai-acquires-presentation-startup-nextslide
- techcrunch.com/2026/08/07/airbnb-says-ai-is-helping-it-ship-features-faster-as-it-tests-a-new-search-function
- techcrunch.com/2026/08/04/spotify-adds-merlin-to-its-ai-music-remix-and-covers-effort
- techcrunch.com/2026/08/04/waymo-opens-up-robotaxi-service-in-dallas-to-everyone
- support.google.com/wallet/answer/17211296
- techcrunch.com/2026/08/06/google-wallet-now-lets-parents-set-up-secure-balances-for-their-kids
- techcrunch.com/2026/08/05/hark-previews-its-browser-use-agent-for-completing-tasks