All episodes

AI & Tech Daily

When AI Research Starts Producing Proofs

18:57

OpenAI says a massive agent swarm has produced and formally checked a proposed solution to the Navier–Stokes Millennium Prize Problem, but independent mathematical acceptance is still to come. Jesse examines what that claim could mean for research verification, then covers Meta’s action-taking Muse agent, Google’s evidence of faster AI-assisted cyber campaigns, US allegations of industrial-scale model distillation, Qualcomm and Amazon’s inference-chip partnership, and an urgent N-central security fix. The practical final section looks at more precise editing in ChatGPT Images 2.5 and NVIDIA’s early Rust toolchains for CUDA kernels.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

A Proof Produced at Machine Scale

A mathematical problem unsolved for generations may now have an AI-generated answer. The harder question is whether the research world can verify work produced at this scale.

That’s why this is the development worth spending time on. On 8 September, OpenAI published what it describes as a complete analytical solution to the Navier–Stokes Millennium Prize Problem, backed by a formal proof in Lean. The claim is specific: a smooth, forced, three-dimensional fluid flow can develop a singularity in finite time. Put more plainly, the equations can begin with a well-behaved flow and still reach a point where that smooth mathematical description breaks down.

Navier–Stokes equations describe how fluids move. Mathematicians have long wanted to know whether their three-dimensional solutions always remain smooth, or whether they can blow up. The Clay Mathematics Institute made that question one of its seven Millennium Prize Problems. Its website still lists Navier–Stokes as unsolved, which is the status that counts while independent review is pending. OpenAI has made a major claim and released material for scrutiny. It hasn’t ended the process.

The way OpenAI says it reached the result is almost as striking as the proposed proof. An internal model, described by the company as more capable than GPT-6 Astra, coordinated roughly ten thousand agents working concurrently. OpenAI reports that the search took 88 hours, followed by another 17 hours for formalisation and verification in Lean, with about 130 billion output tokens generated during the Navier–Stokes effort. Lean is a proof assistant: it checks whether each formal step follows from stated definitions and rules. That provides a stronger artefact than an ordinary model-generated explanation, but formalisation doesn’t remove every question. Reviewers still need to examine whether the formal statement matches the original mathematical problem and whether all assumptions and constructions hold.

The sheer volume changes the shape of review. No mathematician is going to retrace 130 billion generated tokens as a conventional working notebook. The published analytical proof and its Lean counterpart have to carry the argument, while independent experts probe the choices that led there. Formal verification can narrow the space for a hidden logical slip. It can’t decide by itself that the chosen definitions faithfully capture every condition in the prize problem. That alignment between the human statement and the machine-checked statement is where careful mathematical reading still earns its keep.

There’s also a provenance issue. OpenAI says it began after hearing rumours of related work and didn’t access specific user data for the project. It also says it can’t rule out de-identified product data having indirectly improved its models. That disclosure doesn’t establish that outside work was used in the proof. It does show why research institutions will want a clearer record of how systems discover, combine and validate ideas when thousands of agents are involved. A result can be logically correct while questions about intellectual contribution remain open. At this scale, the record of the search may be too large to serve as a useful explanation of origin.

For mathematics organisations and AI labs, my view is straightforward: treat this as a substantial proof package to test, not a solved prize to celebrate. If specialists find an error, the release will still be a revealing stress test of agentic research. If the proof survives, the larger change will be a research process capable of generating results faster than existing review systems can absorb. Verification, attribution and independent reproduction then become part of the scientific infrastructure, not paperwork added after the interesting work is done.

Meta’s Agent Starts Taking Action

With that proof now in reviewers’ hands, the shorter news starts with the same shift at a practical scale: AI systems beginning to act.

Meta launched Muse on 8 September for adults in the United States, through dedicated apps, the web and WhatsApp. Muse can browse, fill in forms and carry out approved tasks inside a dedicated cloud virtual machine. It can keep working after the app closes, so the experience is closer to delegating a job than keeping a chatbot conversation open. That persistence is useful, but it also means the user may not be watching each intermediate choice as it happens.

The examples matter less than the authority behind them. An agent that can use services on your behalf may encounter email, account details, shopping decisions and personal preferences while it works. Meta says Muse asks for confirmation before sensitive steps such as sending an email or making a purchase. Users can restrict permissions for connected apps, inspect an audit trail, opt out of model training and keep Muse data separate from Meta’s advertising systems.

For now, the work runs in Meta’s Secure VM. The company says a stronger Confidential VM will arrive later in 2026, with an encryption key held only by the user. That future design could reduce how much trust has to be placed in the operator of the infrastructure, but it isn’t the architecture people are receiving at launch. Meta’s privacy and safety claims also haven’t been independently tested at large scale. Availability starts with people aged 18 and over in the US, and Meta plans subscriptions for heavier use, without detailed pricing yet.

The useful way to judge Muse is by the controls around an inevitable mistake. Can a person see what the agent did, stop it, recover cleanly and narrow its permissions without losing the convenience that made delegation attractive? For eligible users, multi-step chores may become easier. My concern is that permission design and audit trails are now everyday consumer product features, not specialist enterprise settings. Muse makes that trade-off tangible: one service gets enough context and execution authority to save time, and the quality of its safeguards determines how costly a mistaken action becomes.

Cyber Campaigns Accelerate

That action layer looks quite different when the operator is an attacker. The practical issue is speed, not a science-fiction system running an entire campaign alone.

Google’s Threat Intelligence Group says some threat actors have progressed from isolated AI prompts to agentic workflows that coordinate several stages of an operation. In one incident during the second quarter of 2026, an actor compromised cloud infrastructure, then planned, built and executed an agent-enabled mass credential-harvesting campaign in under six hours. The framework handled scanning, credential collection, operational errors and traffic routing with limited human involvement.

Six hours changes the defender’s working window. A cloud alert that once entered a queue for review later in the day can now sit beside automated activity that is already scaling across accounts and infrastructure. The attacker doesn’t need perfect autonomy to gain that advantage. A framework that keeps scanning, handles routine errors and reroutes traffic can remove the pauses that once slowed an operation whenever a script broke or a person had to coordinate the next step.

Google also observed attackers targeting model code, proprietary research, API credentials and cloud capacity that could be used for unauthorised AI workloads. Those assets belong in the same incident picture as conventional identities and servers. An exposed API credential may provide access to a model, while compromised cloud capacity can support further automated work. Defenders need to see those links while the incident is developing, not reconstruct them the following morning.

There is an important boundary around the finding. Google says it hasn’t observed threat actors operating fully autonomous offensive pipelines against targets in the wild. Its evidence shows AI being inserted into parts of existing operations. And the report reflects Google’s own investigations and visibility; it doesn’t tell us how common these methods are across the whole threat landscape.

Even with those limits, defenders have a concrete operational problem. My read is that organisations need detection and containment built for the shorter gap between initial access and scaled exploitation. Cloud accounts, developer tooling, API keys and AI infrastructure can’t sit in separate monitoring silos while an attacker’s framework crosses them automatically. Faster automation on the offensive side makes a human-paced handoff between several security queues riskier, even when people still direct the campaign.

Distillation at Industrial Scale

There’s another kind of extraction happening at the model boundary, and here the evidence comes with a clear qualification: these are allegations from US agencies.

The NSA, FBI and CISA released a joint advisory on 8 September alleging that six China-based AI companies ran coordinated, industrial-scale distillation campaigns against US frontier-model providers. The agencies say the activity began in late 2024 and generated billions of tokens through millions of requests. They describe traffic distributed across providers and accounts, routed through aggregators and proxies, with metadata sanitised and systems configured to fail over automatically between services.

Knowledge distillation itself is a legitimate technical method. A smaller model learns from the outputs or behaviour of a larger one, often to reduce cost or make deployment easier. The advisory distinguishes that ordinary practice from the alleged campaigns, which the agencies say were designed to extract restricted capabilities in breach of provider terms. The public material doesn’t include enough of the underlying evidence for outsiders to independently verify the scale or attribution, and it doesn’t include responses from the named firms. Those limits should stay attached to the claim.

The defensive guidance is still useful. The agencies recommend correlating behaviour across accounts and infrastructure, looking for unusual ratios in how services are used, and sharing indicators among model providers, cloud platforms and API intermediaries. A single account might resemble a heavy but ordinary customer. Millions of coordinated requests spread across many identities reveal a different pattern only when providers can connect them.

For model and API operators, I think the difficult part is getting that wider view without making legitimate research access miserable. Crude per-account limits are easier to evade when activity is deliberately distributed, and aggressive restrictions can catch developers doing valid evaluation or distillation work. The better controls will recognise coordinated behaviour across infrastructure while giving ordinary customers clear, predictable boundaries. This advisory pushes model protection closer to fraud detection: relationships and timing may tell you more than any one request.

A New Bet on Inference Silicon

Now for the physical layer beneath all of that computation. Qualcomm and Amazon have announced a broad partnership, although there’s no product for AWS customers to buy yet.

The companies plan to develop multiple generations of customised silicon for AI inference in Amazon Web Services data centres. Inference is the stage where a trained model answers a prompt or processes new input, and its cost becomes especially important when a service operates at high volume. The agreement also covers high-bandwidth optical connectivity, with the companies naming 1.6 terabits per second and later generations as a target. Fast links matter because a large AI workload often has to move data among many chips rather than run neatly on one device. Compute and networking therefore have to advance together if bigger deployments are to avoid spending too much time moving information around.

There’s a second part to the relationship: Qualcomm plans to increase its use of AWS infrastructure and Amazon Bedrock for electronic design automation. That puts AWS services into more of the chip-design workload while Qualcomm contributes prospective compute and connectivity products for Amazon’s data centres. The arrangement is broader than a one-off component order. Calling it multi-generational suggests both sides expect the work to continue as models, packaging and data-centre networks change.

What we don’t have is just as important. The announcement gives no architecture, benchmark, deployment date, order volume or customer pricing for the inference products. It tells us the two companies intend to build across several generations, not how the first generation will compare with hardware already operating at scale. Without those details, buyers can’t estimate performance per dollar, power demands or when the collaboration could affect available AWS capacity.

For AWS and organisations buying inference capacity, this could eventually create another source of custom silicon and interconnect technology beyond the established accelerator suppliers. My assessment is that supply diversity is the interesting possibility, but there’s no immediate change to procurement or cloud architecture. The partnership earns attention when delivered chips have measured performance, availability and prices. Until then, the 1.6-terabit optical target and multi-generation scope show ambition, not an operating advantage.

An Urgent N-central Patch

One item does require an immediate operational response. If your organisation runs N-able N-central on premises, this is a patch-and-investigate job.

N-able released N-central 2026.3 Hotfix 4, build 2026.3.1.14, on 5 September to fix CVE-2026-86218. The vulnerability is rated critical and allows remote code execution before authentication. On-premises installations older than that fixed build are affected and should be upgraded immediately. N-able says its hosted instances have already been patched, and applying the server hotfix doesn’t require endpoint agents to be upgraded. That last detail keeps the remediation focused on the N-central server rather than turning it into a fleet-wide agent rollout.

The reporting on exploitation conflicts. N-able’s public release notes say the company has no confirmation of exploitation in production. Canada’s Cyber Centre says N-able indicated that the flaw was being exploited in the wild. Neither public source provides detail about observed attacks, so there’s no sound basis for smoothing those statements into one definitive account. Operators should know that the disagreement exists. It may affect how urgently they search historical activity, but it shouldn’t alter whether they install the fix.

It doesn’t change the sensible priority. N-central is a remote-monitoring and management platform, so a compromised server may provide broad administrative reach into the environments it controls. Pre-authentication remote code execution removes the need for an attacker to begin with a valid login. The possible blast radius therefore outweighs the uncertainty over which exploitation description is current. A server at the centre of remote administration deserves a different response from an isolated application with little reach.

For an on-premises operator, I’d separate the work into two connected actions: get the server to build 2026.3.1.14, then examine recent access and system activity for anything suspicious rather than assuming the patch closes the historical question. Patching prevents exposure through the known flaw from continuing; it doesn’t tell you whether someone used it before the upgrade. Hosted customers don’t have that upgrade task because N-able says those systems are already fixed. The audience most affected is the organisation maintaining its own N-central server, and waiting for the two public accounts of exploitation to converge only gives a potential attacker more time.

What Changes for You

Two releases are worth putting into your working toolkit, with very different levels of maturity.

OpenAI has started rolling out ChatGPT Images 2.5 across all ChatGPT, ChatGPT Work and Codex tiers on desktop, mobile and the web. The new controls let you sketch a reference, begin with a template, attach comments to specific parts of an image and share prompts. For anyone iterating on a visual, the practical gain is more targeted revision: you can point at the area that needs work instead of repeatedly rebuilding the whole composition.

Developers also get two API models. GPT-Image-2.5 Flare is positioned as the faster option, while Sunburst takes longer in exchange for more precise editing. OpenAI claims generation latency up to 50 per cent lower than Images 2.0 and says reference subjects and earlier edits are preserved better. Those are company measurements, so real project testing still has to establish how reliably the improvements hold. API access is separately priced, staged account rollout may take time, and generated images retain C2PA provenance metadata plus invisible watermarking.

For a more experimental developer option, NVIDIA has opened cuda-oxide and cutile-rs, two early toolchains for writing GPU kernels in Rust. cuda-oxide exposes the lower-level SIMT model, where developers work closer to how threads execute. It requires Linux, CUDA 12 or newer, a GPU with compute capability 8.0 and a pinned nightly Rust toolchain. cutile-rs uses a higher-level tile model and leaves thread mapping and memory layout to the compiler; it needs stable Rust 1.89 or newer and CUDA 13.3.

Both compile towards NVIDIA’s GPU stack, through PTX or the CUDA Tile intermediate representation. NVIDIA says Rust’s ownership model can catch some memory and aliasing errors at compile time. That could make kernel work safer for Rust developers, but neither route is ready for production. Coverage is incomplete, APIs will change, cuda-oxide is early alpha, and performance and long-term interoperability remain unproven. Images 2.5 can enter real editing workflows as access reaches an account. CUDA Rust belongs in an evaluation branch where immature tooling can’t surprise a production build.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. openai.com/index/navier-stokes-solution
  2. claymath.org/millennium-problems
  3. about.fb.com/news/2026/09/introducing-muse-personal-ai-agent
  4. apnews.com/article/3a4572eb4cf4e95d8a0dfdad6e6ca065
  5. cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai
  6. nsa.gov/Press-Room/Press-Releases-Statements/Press-Release-View/Article/4592113/nsa-and-others-warn-china-based-ai-companies-are-distilling-us-frontier-ai-mode
  7. qualcomm.com/news/releases/2026/09/qualcomm-announces-multi-generational-product-collaboration-with
  8. documentation.n-able.com/N-central/Release_Notes/GA/Content/N-central_2026.3_HF4_Release_Notes.htm
  9. cyber.gc.ca/en/alerts-advisories/n-able-security-advisory-av26-885
  10. openai.com/index/introducing-chatgpt-images-2-5
  11. developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels