All episodes

AI & Tech Daily

AI’s Proof Problem: Who Checks the Breakthroughs?

17:58

OpenAI says an internal model has resolved more than 100 open maths problems, while a new independent advisory group considers how such results should be reviewed and communicated. We examine the unverified claim, the group’s limited authority and the strain AI-generated research could place on scholarly validation. Also: proposed international standards for automated AI research, a coalition tackling the language gap, NVIDIA’s power and cooling qualification program, Amazon’s block on Meta’s Muse shopping agent, a proposed US-China AI incident alert channel, GitHub’s enterprise credential inventory, and what Googlebook changes for laptop buyers and developers.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

The Maths Verification Bottleneck

A machine may have solved more than a hundred open maths problems. The extraordinary part is that nobody outside OpenAI has yet been shown the results.

Our main story today looks at OpenAI’s claimed mathematical breakthrough and the harder question behind it: who can verify a flood of machine-generated research without lowering the standard of proof?

OpenAI said on 21 September that an internal model had resolved more than 100 long-standing open problems across many areas of mathematics. The company says training began on 28 August, so the claim covers a striking amount of work in less than a month. But the announcement didn’t include the results. There are no proofs for outside mathematicians to inspect, no independent verification and no indication yet that the mathematical community accepts any of the claimed solutions. For now, the number is OpenAI’s claim, not an established scientific result.

Alongside that claim, OpenAI and nine mathematicians announced an independent Advisory Group on Mathematics and Artificial Intelligence. Its immediate job is to advise on how AI-generated mathematical results should be reviewed and communicated. The members are unpaid. They can publish their recommendations, offer advice without waiting to be asked and speak publicly about OpenAI’s effect on mathematics. That gives them a formal channel and some freedom to criticise the company’s approach.

There are firm limits. The group has no decision-making authority inside OpenAI, and it won’t advise the company on how quickly to advance its internal mathematical work. OpenAI still decides what it releases, when it releases it and whether it follows the group’s recommendations. The group’s own account describes its current task as working out principles and processes, rather than endorsing the company’s claimed solutions. Its practical influence is still untested.

That distinction is important because mathematical proof has always depended on scrutiny, not the reputation of whoever presents it. A correct-looking proof can hide a subtle gap. Specialists need the underlying argument, enough time to check it and a way to record objections and corrections. If AI systems can produce plausible results much faster than researchers can validate them, the scarce resource becomes expert attention. Publishing a bigger pile of candidate proofs doesn’t create a bigger supply of mathematicians able to check them.

There’s also a communication problem. A company announcement can travel around the world before the proof behind it reaches the people qualified to assess it. An independent advisory group could help OpenAI separate a model result, a checked proof and a result accepted by the field. Those are three different stages, and collapsing them into one headline would weaken confidence in all of them.

A workable release process may therefore need two speeds. Researchers could first identify which candidate results are complete enough and consequential enough to justify scarce specialist time. Full validation would still proceed at the pace the mathematics demands. Clear status labels would let other researchers see whether they are looking at an unchecked model output, a proof under review or a result that has survived expert challenge. That kind of triage wouldn’t prove anything by itself. It would stop the queue of claims from becoming the public record by default.

My read for research organisations is that verification capacity now deserves the same planning as model capability. If OpenAI eventually releases a large body of valid work, journals, universities and professional societies may need new review workflows simply to keep up. The advisory group is a useful opening, but its credibility will come from visible recommendations, disclosed results and normal mathematical challenge—not from the word independent in its name.

Rules for Automated Research

That maths claim now faces the slow work of proof. OpenAI is also trying to shape how advanced research systems are measured.

On 21 September, the company published a proposal for internationally compatible measurements, safeguards and incident-reporting protocols for frontier AI and increasingly automated AI research. This is a policy paper, not a new legal regime. It creates no licence, mandatory pre-release review or model approval, and no organisation acquires a fresh compliance obligation from it. Governments would decide whether any part of the proposal belongs in law.

OpenAI suggests using the existing network of AI safety institutes, including Australia’s, to develop common evaluations and reporting protocols. The attraction is fairly practical. If labs and regulators use the same terms for capability, safeguards and incidents, evidence collected in one country becomes easier to compare with evidence from another. A shared vocabulary can also reduce ambiguity when agencies compare incident reports.

The company also says fully autonomous recursive self-improvement isn’t happening today and shouldn’t be pursued unless human control and safety can be preserved. Recursive self-improvement means a system repeatedly improving the machinery used to build its successors. OpenAI is placing that possibility inside the standards discussion before it exists as a demonstrated current capability.

The gap is implementation. No government has adopted the proposal, no negotiating process has been announced and there’s no timetable. Voluntary definitions can improve internal practice, but they won’t produce comparable evidence if labs choose what to disclose or if independent institutions have no meaningful role.

For organisations working across borders, I’d treat this as a useful draft vocabulary rather than a compliance map. Its value will depend on safety institutes, researchers and governments testing the definitions and deciding what has to be reported. A standard that only its author uses is still a house rule.

AI’s Language Gap

Scale looks very different when the missing resource isn’t computing power, but language data people can trust.

Sixty organisations have announced a five-year commitment aimed at making AI services usable in the languages and voices of an estimated 3.4 billion people whose languages are underrepresented in current systems. The coalition’s starting point is blunt: of roughly 7,000 languages, only a small proportion have enough resources to support strong AI capabilities.

The planned work includes openly licensed language infrastructure, shared assessments, models and applications that builders can access, plus protections for privacy, consent and data sovereignty. In practice, better language infrastructure can mean carefully collected text and speech, agreed ways to test whether a model actually understands a language, and tools that developers can use without rebuilding the foundation for every application.

Nothing usable arrives simply because the coalition exists. Its detailed governance, workstreams and the commitments of individual members are due to be developed over the coming year. Funding allocations, delivery milestones and measurable targets for particular languages haven’t been settled. People can’t download a promised dataset or switch to a newly supported service today.

The unresolved issue is control. Language and voice data can carry identity, culture and information about a community. Open licensing may help developers build useful systems, but openness alone doesn’t answer who consented to collection, who decides acceptable uses, or whether the people represented can challenge a harmful deployment. The coalition acknowledges privacy, consent and data sovereignty; the real test is how those ideas appear in its governance and in each data project.

For the public, the potential benefit is access to services that work in the language people use at home, rather than forcing them into one of a few well-resourced languages. My judgement is that language-by-language goals and community authority will be more revealing than the coalition’s headline population figure. If local speakers help set collection and use rules, builders could gain better data without treating those communities as raw material.

Power and Cooling Get Qualified

The next constraint is much more physical: keeping racks powered and cool once the accelerators arrive.

NVIDIA launched DSX Ready on 21 September, a vendor qualification program for infrastructure products intended to meet applicable requirements in its DSX reference designs for AI factories. The first categories cover battery energy-storage systems and cooling-distribution units. More infrastructure and software categories are promised later.

The phrase AI factory refers here to the large data-centre systems used to train and run AI, not a single box full of chips. At that scale, power quality, stored energy and liquid cooling affect whether expensive computing hardware can operate reliably. A cooling-distribution unit moves and controls coolant between facility systems and the computing equipment. If that part isn’t matched properly to the rack design, accelerator specifications won’t rescue the installation.

NVIDIA named three initial suppliers in each of the two categories. The qualification paths differ: battery suppliers submit test data for NVIDIA to review, while cooling suppliers use a self-qualification suite. That means the DSX Ready label isn’t one uniform independent testing process. It is an NVIDIA-run program tied to NVIDIA’s own reference designs.

The company is careful about the boundary. Qualification doesn’t replace engineering for the actual site, and it doesn’t imply that the completed facility will be stable. Operators still have to account for their electrical system, building, climate, rack layout and operating conditions. The announcement also provides no independent evidence yet that the program cuts construction time, cost or failures.

For data-centre builders, the shortlist may reduce some early integration work by identifying products designed around the same reference requirements. I’d use it as a compatibility filter, not a finished facility design. The bigger signal is that power and cooling are being packaged as qualified components of the AI stack. The deployment bottleneck has moved well beyond finding enough accelerators.

When an Agent Meets a Closed Door

A capable agent still needs somewhere it’s allowed to act, and one large storefront has now drawn that boundary.

Amazon has blocked Meta’s Muse personal agent from shopping on its website. According to GeekWire’s report from 20 September, Amazon cut off access after unsuccessfully asking Meta to remove the store from Muse’s supported experience. Affected users began seeing a notice that access by an unauthorised AI agent violated Amazon’s conditions of use.

Amazon alleges that Muse doesn’t identify itself while browsing and may create risks around credentials, privacy and security. Meta had previously said Muse can’t see users’ passwords or payment methods and that credentials shared with the service are kept in secure storage. Meta didn’t immediately respond to the report about the block, so the credential handling and identification concerns remain competing company positions, not settled findings.

The immediate effect is straightforward: Muse users can’t currently depend on the agent to shop through Amazon. There’s no announced resolution process and no indication of how long the block will last. An agent can understand a request, compare products and navigate a shopping flow, yet the destination can still refuse the automated visitor.

That exposes a weak point in the idea of a general personal agent. Websites have their own conditions, fraud controls, account protections and commercial interests. If an agent acts without clearly identifying itself, the site may not know which permissions the user granted or who is responsible when a transaction goes wrong. If every destination sets a different rule, the agent’s advertised reach can change without the model changing at all.

For people considering consumer agents, supported services need to mean more than technically reachable services. My reading is that dependable shopping agents will require negotiated access, transparent identity and clear liability arrangements between the agent provider and the retailer. Until then, a familiar website can disappear from an agent’s working world overnight.

A Hotline for Serious AI Incidents

At national scale, the access question becomes an alert question: who calls whom when an AI incident crosses borders?

US Treasury Secretary Scott Bessent said the United States proposed a notification mechanism for AI incidents that could affect national security during discussions with Chinese Vice Premier He Lifeng. The proposal was raised during weekend talks ahead of planned leader-level discussions. Bessent described it as a way to improve transparency around shared AI threats.

There is no agreement yet. Neither side has announced an incident threshold, an operating body or a protocol for what information would be exchanged. The proposal imposes no reporting requirement, and acceptance and implementation remain unknown.

A narrow, reliable alert channel could still be useful. During a serious incident, delay and ambiguity can encourage each side to infer intent from incomplete information. A direct notification could reduce that risk if both parties agree on scope, verification and timely participation. Those conditions do most of the work. A channel that isn’t used promptly, or that lacks a shared definition of a nationally significant event, offers little reassurance.

For organisations watching international AI governance, this is a diplomatic opening rather than an operating safeguard. The practical milestone will be an agreed mechanism with named contacts and clear triggers. Until then, it remains a proposal to talk when something serious happens.

One View of Enterprise Credentials

There’s a smaller change for security teams, but it solves a very familiar incident-response problem.

GitHub now provides authorised Enterprise Cloud users with CSV and REST API access to an enterprise-wide inventory of credentials that can reach their environment. The inventory covers SSH keys, personal access tokens, OAuth access tokens, and GitHub App user and installation tokens.

Security teams can inspect who owns a credential, its scope and permissions, when it was created and expires, when it was last used, and which organisations or repositories it can target. They can also correlate the inventory with audit-log activity. During an incident, that replaces the awkward job of assembling separate lists across organisations and applications before responders can work out which credentials may be exposed.

The feature is available now for GitHub Enterprise Cloud. GitHub says Enterprise Server support is coming in a future release, but it hasn’t supplied a date. The distinction matters for organisations running their own server: the new central view isn’t available to them yet.

My practical take for Enterprise Cloud security teams is to incorporate the export into credential reviews and incident-response preparation now. Access to the inventory itself needs tight control because it maps powerful credentials and their reach. Used carefully, it should make the first hours of a credential incident less about finding the records and more about revoking the right access.

What Changes for You

One product announcement does change what you can buy and use now, though the biggest questions will need hands-on testing.

Google opened pre-orders on 21 September for Googlebook, a laptop platform built around Gemini features and a developer-ready Linux environment. Devices are due on Australian shelves on 5 October, although Google hasn’t published Australian local pricing in the announcement. Google lists prices from 899 dollars, and each device includes 12 months of Google AI Pro with five terabytes of cloud storage.

For everyday work, Magic Pointer can act on content you select on screen. Rambler restructures dictated notes and supports changing languages mid-sentence. Create My Widget generates small tools from a description. The important shift is that these features sit in the laptop workflow rather than waiting in a separate chatbot tab.

Developers get Google Antigravity, a full Linux terminal and support for tools including Claude Code on every device. That puts ordinary command-line work and agentic coding tools beside the consumer AI features, so the same machine is being pitched to people who want assistance with documents and to developers who want a proper local working environment.

The limitation is how much remains untested outside Google’s announcement. Performance and privacy claims come from the vendor, and Google hasn’t clearly separated local processing from cloud processing for every feature. Buyers are also taking on a substantial hardware cost and a deeper dependence on Google services, while the included AI Pro period lasts 12 months.

My view is that operating-system integration could make AI assistance more useful because it can work with the thing already on screen. But the buying decision should turn on how reliably those features work, what leaves the device and what the ongoing service costs look like—not on the novelty of generating a widget during a demo.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. openai.com/index/advisory-group-on-mathematics-and-ai
  2. terrytao.wordpress.com/2026/09/21/advisory-group-on-mathematics-and-artificial-intelligence
  3. openai.com/index/building-standards-next-phase-ai
  4. gatesfoundation.org/ideas/media-center/press-releases/2026/09/ai-language-partnership
  5. apnews.com/article/artificial-intelligence-anthropic-openai-gates-aefb021bede3b02c83890f65cd540fd0
  6. blogs.nvidia.com/blog/dsx-ready-ai-factories-power-cooling
  7. geekwire.com/2026/amazon-blocks-metas-muse-ai-assistant-in-new-standoff-over-agentic-shopping
  8. apnews.com/article/bessent-ai-xi-trump-china-trade-2c7f54f07e755f506d9db9b91df282bd
  9. github.blog/changelog/2026-09-21-github-enterprise-adds-credential-inventory-exports
  10. blog.google/products-and-platforms/devices/googlebook/googlebook-built-in-intelligence