AI & Tech Daily
When AI Proofs Become Too Big for Humans to Referee
Claude agents have produced an open, machine-checked formalisation of Fermat’s Last Theorem, pointing to verification as the trust layer for AI-generated mathematics. We also examine Microsoft’s blueprint for local-agent PCs, AWS’s Sydney agent registry, Canada’s principles for AI data centres, the US automated-vehicle roadmap and Qualcomm’s lower-cost edge processors. In What Changes for You: an actively exploited Chrome flaw needs a verified update, and NVIDIA’s beta PAIR router can spread local inference requests across compatible computers.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
A Proof Too Large to Referee
AI agents have turned a famous theorem into roughly 13 million lines of machine-checked mathematics in 11 days. The startling part isn’t a new proof; it’s the arrival of a different way to trust work at that scale.
This is worth sitting with, because the achievement and the limitation are really the same thing. Anthropic says collaborating Claude agents produced the first complete computer-checked formalisation of Fermat’s Last Theorem. The result is published as an open repository written in Lean, a proof assistant that checks whether every logical step follows from explicitly defined rules.
Fermat’s Last Theorem says there are no positive whole-number solutions to a to the power of n plus b to the power of n equals c to the power of n when n is greater than two. Anthropic’s agents haven’t found a different route or solved an open mathematical problem. They’ve formalised the established Wiles and Taylor–Wiles route so a computer can verify it. That distinction matters. It keeps an impressive engineering result from becoming a misleading claim about mathematical discovery.
The scale is hard to ignore. Anthropic reports about 29,500 intermediate theorems in the final proof. It says the work took 11 days and was carried out largely autonomously by multiple agents. The repository reports successful checks in Lean, compares its final statement with the one in the established Mathlib library, and includes a check using a second Lean kernel implementation. Those are important safeguards because a formal proof is useful only if the claimed theorem and the machinery accepting it are both what researchers think they are.
Replayability changes the review task. Another group doesn’t have to trust a fluent explanation or accept a huge chain on reputation. It can obtain the published artefact, run the required checks and ask a small kernel to reject an invalid inference. That doesn’t eliminate judgement: experts still need to inspect the definitions, assumptions and relationship to the intended mathematics. It does separate logical validity from persuasiveness.
Still, publication isn’t the same as broad validation. Independent mathematicians and formal-methods specialists had only just begun reproducing and reviewing the work by the briefing cutoff. Thirteen million lines also make this an unusually large artefact to inspect, maintain and understand. A kernel can confirm that steps obey the formal system; it doesn’t automatically tell a human which parts are elegant, reusable or conceptually illuminating.
My read is that verification, rather than generation alone, is the consequential advance. If AI systems can produce formal mathematics faster than people can referee it line by line, then a small trusted checker becomes the practical trust boundary. That pattern reaches beyond pure mathematics. Software proofs, hardware verification and safety-critical systems all depend on converting plausible output into claims that can be replayed and checked.
There’s also a shift in where scarce expertise goes. Researchers may spend less time constructing every formal step and more time specifying the right claim, choosing the right abstractions, auditing generated structures and improving the libraries later projects depend on. The 11-day figure suggests that possibility; it doesn’t yet establish a repeatable timetable for other difficult proofs.
For universities, research labs and formal-verification teams, the open repository is now a concrete test case: reproduce the build, examine the assumptions, and see which parts of the workflow generalise. If that scrutiny holds up, multi-agent formalisation could make projects that once required years of specialist labour far more achievable. But the immediate work is verification of the verifier’s result, not celebration of an AI mathematician replacing mathematical judgement.
Windows Defines the Local-AI PC
That giant proof leaves one practical question hanging: what sort of computer will developers use for the smaller agents running beside them every day?
Microsoft has put a specification around one answer with Project Zenith, a preconfigured Windows 11 experience intended to ship on developer-class computers. The first systems are planned around AMD’s Ryzen AI Halo platform. Microsoft’s stated floor is at least 64 gigabytes of unified memory and 250 gigabytes per second of memory bandwidth.
Unified memory gives the processor and graphics hardware access to a common pool, which is useful when a local model is too large for the dedicated graphics memory found in a conventional laptop. Microsoft says Zenith machines will be able to run models above 30 billion parameters locally. That’s a capability claim, not yet a full performance picture: there are no exact device launch dates, prices or independently tested speeds in the announcement.
The software configuration may be just as important as the hardware. Microsoft says the machines will arrive with development tools and settings, deeper container integration through the Windows Subsystem for Linux, and protections for agents that include operating-system-enforced identity and containment. The company is effectively treating an agent as a local process that needs a defined identity and boundaries, not simply a chat window with permission to roam through a developer’s files.
For developers and procurement teams, Zenith creates a more legible baseline. Buying a machine for local coding agents can shift from vague claims about an ‘AI PC’ to questions about memory capacity, bandwidth, tool compatibility and isolation. My practical take is that local inference may reduce recurring cloud use for suitable jobs, but the cost doesn’t disappear. It moves towards high-memory hardware and deeper commitment to a platform’s security model.
Zenith isn’t a downloadable setup for an existing PC, and Microsoft hasn’t said how widely it will extend beyond the initial AMD systems. For now, it’s a category definition rather than a product you can order. Even so, the specification gives the industry a clearer argument about what a serious local-AI development machine needs to be.
An Agent Registry Arrives in Sydney
Once agents move from experiments into teams, keeping track of what they can access becomes its own engineering problem.
AWS has made Agent Registry generally available, including in the Asia Pacific Sydney region. It’s a private catalogue for agents and the components around them: tools, skills, MCP servers and custom resources. Teams can work with a registry through the AWS console, command-line interface and software development kits, while developers can also query it as an MCP server.
The release adds infrastructure-as-code support, cross-account sharing and organisation-wide automatic detection of AgentCore runtimes and gateways. That turns the registry into more than a list someone maintains in a spreadsheet. A large organisation can discover supported AWS resources, attach governance to them and give builders a common place to find approved capabilities. In Australia, the Sydney availability also means that inventory can be established in-region rather than managed through a registry in another geography.
Cross-account sharing matters in a large AWS estate, where a useful internal tool may be built in one account and needed in another. Infrastructure-as-code support also makes the registry’s configuration reviewable and repeatable alongside the systems it describes.
There are boundaries. The service initially operates in five AWS regions, and the new automatic discovery is focused on supported AWS resources. A company with agents spread across several clouds and local systems will still have work to do if it wants one complete view. The registry is also part of AWS’s control plane, which can simplify operations while making that provider more central to how an organisation describes and approves its agent estate.
My assessment for technology leaders is fairly direct: agent governance is becoming normal cloud infrastructure. A shared inventory can reduce duplicated tools and expose shadow deployments before they turn into security or compliance surprises. The trade-off is that the catalogue only earns trust if teams actually register what they build and reconcile what the platform discovers.
The useful shift here is from asking whether an organisation permits ‘AI agents’ in the abstract to recording which agent, tool and server exists, who can find it, and how it is shared. That’s less glamorous than a new model, but it’s the kind of plumbing that determines whether agent deployments remain manageable after the pilot stage.
Who Pays for AI Data Centres
The infrastructure conversation gets much more physical when the agent registry gives way to the buildings powering all that computation.
Canada has published Responsible Data Centre Development Principles, setting a national baseline for how new projects should handle local benefits, electricity costs, water, environmental effects, transparency and strategic value. The framework doesn’t replace approvals run by provinces, territories, municipalities or Indigenous authorities. It is intended to complement those processes with a common set of expectations.
One principle deals directly with a question communities are already asking: who pays when a large facility needs a new grid connection or network upgrades? The document says costs attributable to the data centre should be paid by the project proponent, rather than shifted to existing electricity customers. It also calls for water use and environmental effects to be measured and reported transparently, and for host communities to receive information they can independently verify.
That combination is significant because AI infrastructure proposals are often framed through investment, jobs and national computing capacity. Canada’s principles put system costs and community evidence into the same assessment. Companies and cloud providers that signed the framework now have a public benchmark against which a proposed facility can be tested, even where the precise approval rules remain regional.
The limitation is enforcement. The announcement doesn’t create one national mechanism or a timetable for turning every principle into binding conditions. Its effect will depend on how regional authorities use the framework, what project agreements require, and whether the reported data is detailed enough for genuine scrutiny.
For data-centre operators, my reading is that electricity, water and local consent are moving from side issues to core project inputs. A proposal that can’t explain its grid costs or produce verifiable environmental information may face a harder approval path, even if its computing capacity is strategically attractive. For communities, the framework offers better questions to ask, but not an automatic answer.
That’s a useful model for other countries considering rapid AI infrastructure growth: account for the computing benefit and the physical burden in the same decision. The principles create that expectation. Regional approvals will determine whether it has teeth.
A Roadmap for Automated Vehicles
Policy is also trying to catch up with machines that act in public, although a roadmap is still a long way from a road rule.
The United States Department of Transportation has released a national automated-vehicle strategy covering current and planned federal activity from the 2026 through 2030 financial years. It sets four objectives: improve safety, provide regulatory certainty, support commercial deployment, and develop interoperable transport corridors across state lines.
The department says it will work towards its first safety standards for automated-vehicle competency. That phrase points to a basic gap in the market: developers and the public need comparable ways to judge what an automated driving system can actually do. A shared federal competency standard could provide a common reference across systems that are now described with different capabilities, operating conditions and levels of human responsibility.
But this release is a strategy, not the finished standard. It doesn’t settle the final safety tests, implementation dates or the relationship between new federal measures and differing state rules. Companies therefore have a clearer view of Washington’s direction, without yet having the concrete compliance obligations that would let them design to one final target.
For vehicle developers, the useful signal is that safety evidence, deployment rules and cross-state operation are being considered together. My judgement is that regulatory certainty could make commercial planning easier, but faster deployment will only carry public credibility if the eventual competency tests are measurable and demanding. A national label without a robust test behind it would add consistency in language, not confidence in the vehicle.
The next consequential developments will come from agencies converting these objectives into standards and regulations. Until then, the strategy narrows the policy questions but doesn’t resolve the engineering or safety arguments.
Edge AI Moves Downmarket
At the other end of the computing spectrum, AI is being designed into devices that may never look like computers at all.
Qualcomm has announced two processors, the Dragonwing Q-2390 and industrial IQ-2390, aimed at lower-cost connected devices. Both combine application processing, local AI, vision, connectivity and real-time control. The core package includes a quad-core Kryo processor, an Adreno 704 graphics processor and a real-time RISC-V microcontroller, with support for Linux, Android and the Zephyr operating system.
The industrial IQ-2390 adds features intended for harsher and more time-sensitive deployments. These include two Gigabit Ethernet connections with time-sensitive networking, plus operation across Qualcomm’s stated temperature range from minus 30 to 115 degrees Celsius. That puts potential uses closer to industrial controllers, cameras and automation equipment than to the premium phones where on-device AI is most visible.
Integration is the attraction. A manufacturer evaluating machine vision or local automation can consider one platform for general software, AI processing, connectivity and real-time tasks, rather than assembling those jobs from more separate components. Running inference at the edge may also reduce latency and keep some data on the device instead of sending every input to a remote service.
There’s no basis yet for declaring the performance or economics a win. Qualcomm hasn’t published comparative AI results or pricing. Customer engagement is at early-access stage, and evaluation kits aren’t expected until early 2027. Commercial product dates, unit costs and the size of models that will work well in practice are still unspecified.
For device makers, the sensible interpretation is that capable edge AI is spreading into more cost-sensitive hardware tiers, but design commitments should wait for measured performance and full pricing. A highly integrated chip can simplify a product, while also creating a silicon and software dependency that may last for years. The announcement makes evaluation worth preparing for; it doesn’t complete the production case.
What Changes for You
Two releases deserve a closer look because they change something you can act on now, rather than setting direction for later.
First, desktop Chrome has an urgent security update. Google released version 152.0.7977.82 or .83 for Windows and macOS, and 152.0.7977.82 for Linux, fixing 12 security issues. The company says an exploit for CVE-2026-85046 already exists in the wild. It’s a high-severity type-confusion flaw in V8, Chrome’s JavaScript engine, and crafted HTML can trigger code execution inside the browser sandbox.
Google hasn’t disclosed the targets, scale or full exploit chain while the update rolls out. Update desktop Chrome and then check the installed version, especially across managed fleets. A restart doesn’t prove the gradual rollout has delivered the fixed build. The notice covers desktop Chrome, not every Chromium-based product, so verify those separately. Active exploitation makes this a task for ordinary users and IT teams now, not routine maintenance for later.
For developers with more than one capable computer, NVIDIA has released a beta of Personal AI Router, or PAIR. It presents one local endpoint and sends separate Ollama or LM Studio inference requests to eligible paired machines on the same network. Existing agent harnesses using those interfaces can spread parallel work across otherwise idle computers without being rebuilt around a cloud service.
PAIR supports Windows, macOS and Linux, including GeForce RTX 20-series or newer systems, RTX PRO workstations, DGX Spark, and Apple M4 or newer hardware. Approved pairing and mutual TLS protect the connections, and NVIDIA says prompts and inference traffic remain local. That may reduce cloud-token spending and keep sensitive work within your network, provided you already own suitable machines.
The limitation is in the word ‘router’. PAIR assigns each request to one machine; it doesn’t pool GPU memory or split one model run across several computers, and every serving node needs the requested model. It’s also a beta, with no independent general benchmark behind NVIDIA’s configuration-specific demonstration. For technically capable builders running parallel local agents, it can make existing hardware more useful. It won’t make several small GPUs behave like one large one.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- anthropic.com/research/formalizing-fermats-last-theorem
- github.com/anthropics/fermats-last-theorem
- chromereleases.googleblog.com/2026/09/stable-channel-update-for-desktop_01882797386.html
- digital.nhs.uk/cyber-alerts/2026/cc-4844
- developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network
- docs.nvidia.com/local-ai/nvpair
- blogs.windows.com/windowsdeveloper/2026/09/04/announcing-project-zenith-the-ready-to-code-windows-experience
- aws.amazon.com/about-aws/whats-new/2026/08/aws-agent-registry-generally-available
- aws.amazon.com/blogs/machine-learning/manage-agents-tools-and-skills-at-scale-with-aws-agent-registry
- ised-isde.canada.ca/site/ised/en/canadas-responsible-data-centre-development-principles
- canada.ca/en/innovation-science-economic-development/news/2026/09/government-of-canada-launches-canadas-responsible-data-centre-development-principles.html
- transportation.gov/briefing-room/trumps-transportation-secretary-sean-p-duffy-unveils-new-national-av-strategy
- transportation.gov/policy-initiatives/automated-vehicles/america-leads-dots-national-strategy-automated-vehicles
- qualcomm.com/news/releases/2026/09/-qualcomm-introduces-dragonwing-q-2390-and-iq-2390-processors--e