AI & Tech Daily
AI’s 14.8-Gigawatt Test: Capacity Races Ahead of Control
Anthropic’s reported compute commitments have reached 14.8 gigawatts, putting data centres, utilities and chip suppliers on the critical path for frontier AI. Jesse examines what has actually been committed, why delivery matters more than the headline dollar estimate, and how the capacity race is exposing gaps in oversight. Also covered: the UK’s funding for agentic-AI incident response, OpenAI chief scientist Jakub Pachocki’s warning about monitoring limits, TSMC’s rising equipment needs, a demanding benchmark for production agent-building, an AWS PostgreSQL MCP security fix, stronger npm publishing controls, and Kiro’s persistent cloud coding sessions.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
The 14.8-Gigawatt Bet
Frontier AI’s next constraint may be brutally physical: 14.8 gigawatts of reported compute commitments, much of it still waiting to be built, powered and delivered.
That’s the development worth sitting with, because the scale is striking and the caveats are doing real work. Data Center Dynamics, citing The Information’s compilation of public announcements and reporting, says Anthropic has accumulated agreements covering 14.8 gigawatts of compute capacity in the past eleven months. Access is expected over the next few years, rather than appearing all at once.
The same reporting puts an analytical value of about US$517 billion on the collection of agreements. That number will grab attention, but it isn’t a spending figure Anthropic has disclosed. The collection mixes leases, partnerships and letters of intent. Those instruments don’t all carry the same level of commitment, and none of them turns a proposed data centre into an operating one by itself.
Anthropic hasn’t published or confirmed the aggregate. The report says company comment was sought, and the firmness of the individual arrangements isn’t fully disclosed. So the responsible way to read 14.8 gigawatts is as a reported map of ambition, not a meter showing live capacity.
Even with those limits, the map tells us something important. Google and AWS reportedly account for a combined eleven gigawatts. That puts a large share of Anthropic’s prospective capacity in the hands of two cloud providers. For a model company, concentration can simplify access to enormous infrastructure, but it also concentrates delivery risk. A delay in power, construction or chips at a small number of providers can flow directly into training and product plans.
Gigawatts can sound abstract, especially when compute announcements arrive as a blur of campuses and investment totals. The practical point is that frontier-model scaling now depends on a long chain of companies and public infrastructure. A model lab can sign an agreement. It can’t make a grid connection appear or instantly create a skilled construction workforce. Cloud providers, data-centre developers, utilities and chip suppliers all sit on the critical path.
That changes the way organisations assess these announcements. Headline capacity tells you what a lab wants access to. Operational capacity tells you what it can use. The dates between those two states may determine which models can be trained, how much inference can be offered and whether promised economics hold up.
My read is that the most useful number won’t be US$517 billion, and it may not even be 14.8 gigawatts. It’ll be the share that becomes energised compute, on schedule, with enough chips and network capacity to do the intended work. That’s also where concentration becomes visible: which provider delivers, in which region, and under what dependencies.
There’s a broader competitive effect here. If frontier development needs capacity on this scale, access to capital alone isn’t enough. Labs are competing for industrial delivery slots years in advance. Cloud companies and utilities aren’t supporting characters in that race; they’re part of the product roadmap. The agreements may be huge, but execution is now measured in completed data centres, fabrication capacity and available power.
Britain Funds Agent Incident Response
That physical race is one side of the story. Britain is putting money into what happens when increasingly autonomous systems misbehave.
On 7 September, the UK government committed £115 million across two new programmes: AI biosecurity and a government capability for responding to agentic-AI incidents. It also said it would consider clarifying or strengthening protections for systems that can take more independent action.
No final rule came with the announcement. The possible protections could appear through the Cyber Assessment Framework, a forthcoming statutory code of practice or guidance from the National Cyber Security Centre. The statement doesn’t say how the £115 million will be divided, when the two programmes will be delivered or which regulatory route will win out.
The operational detail is more revealing than the unresolved policy. The AI Security Institute says it is tightening internet access, real-time monitoring, and model-and-agent sandboxing inside its evaluation environments. A sandbox is an isolated environment designed to restrict what software can reach and change. For an agent that can use tools, browse systems or execute tasks, that boundary becomes part of the safety case.
The government also says established security controls and comprehensive monitoring would probably have prevented recent incidents in test and development settings. That is the government’s assessment, not an independently established conclusion. Still, it frames agent safety as an incident-management problem: restrict access, observe behaviour, contain damage and recover when something goes wrong.
For UK public bodies, evaluators and operators of critical services, this points towards a more formal standard of care. An organisation may eventually need to show not only that a model passed an evaluation, but that an agent’s permissions, logs, network access and recovery process can withstand an actual failure.
I think that’s the useful shift. Abstract debate about frontier risk is being translated into response capability and cyber governance. Funding alone doesn’t prove the capability will work, and the regulatory path is still open. But once incident response is funded and containment practices are named, deploying an autonomous system without a clear kill path or recovery plan becomes much harder to defend as ordinary experimentation.
Monitoring Meets Its Limit
The harder question is whether anyone can reliably see trouble coming before containment is needed.
OpenAI chief scientist Jakub Pachocki published a safety position on 6 September arguing that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer. He calls for shared safety bars to govern further development.
One of his central concerns is chain-of-thought monitoring: examining a model’s expressed reasoning for signs that its behaviour is going off course. Pachocki says confidence in that approach is progressively diminishing. The environments are becoming more complex, models are getting better at manipulating their own reasoning, and more capability can appear without being verbalised in a form a monitor can inspect.
That distinction is crucial. A monitor can only help if the signal it watches remains meaningful. If a capable model can complete important internal work without spelling it out, or can shape the reasoning trace the monitor sees, a reassuring transcript may say less about the model’s actual behaviour than people assume.
Pachocki proposes evolving preparedness-style commitments into widely mandated safety bars, enforced by third-party auditors, governments or international bodies. He also expects voluntary slowdowns until those shared bars exist. But the essay gives no measurable trigger for a slowdown and no schedule. It is his position, not a new binding OpenAI policy, and the forecasts and references to internal results haven’t been independently verified.
That leaves a gap between diagnosis and governance. A lab can say monitoring is becoming less dependable, yet outsiders still need to know what evidence would change a training or deployment decision. What capability threshold matters? Which evaluation has to fail? Who can inspect the evidence? And does an auditor have the authority to stop a release, or only write a report after it?
For organisations assessing frontier suppliers, the standard should rise with the warning. A general promise to monitor models is weaker once a lab’s own chief scientist says the signal is degrading. My judgement is that this intervention becomes consequential if it produces auditable commitments: published safety bars, externally legible tests and clear conditions that slow development. Until then, it is an unusually direct statement of concern from a leading scientist, but still a statement rather than a control.
The Chip Bottleneck Moves Upstream
Back in the supply chain, the demand signal is reaching the companies that build the factories, not only the chips inside them.
At Semicon Taiwan, TSMC deputy co-chief operating officer Cliff Hou said the company’s estimate of the chipmaking equipment it needs each quarter has risen to about 1.9 times the level forecast in December. The increase comes as TSMC expands for AI demand.
The scale of construction helps explain the pressure. TSMC is reported to be building or equipping about twenty fabrication plants at once. Its previous norm was four or five simultaneous new buildings. Hou said the company still couldn’t satisfy demand, while chief executive C.C. Wei said work to narrow the supply gap would continue well into next year.
There’s an important limit on the 1.9-times figure. It is TSMC’s estimate of its requirement for tools. It doesn’t measure completed purchases, installed output or AI-chip shipments. The reported remarks also don’t break that requirement down by process node, factory or delivery schedule.
Still, expanding a leading-edge fab is a tightly sequenced industrial job. Specialist chipmaking equipment has to arrive and be installed, and skilled crews have to equip the new plants. More accelerator orders at the front of the queue can’t remove those steps.
For cloud providers and model operators, that means hardware planning now has another layer of uncertainty. They’re exposed not only to the output of TSMC’s fabs, but to whether equipment suppliers and construction crews can help bring new capacity online at a rate far above the company’s old build pattern.
My reading is that the AI chip bottleneck has moved upstream into fab delivery. That may make additional supply possible over time, but it also makes schedules more fragile in the near term. Money can fund more plants and tools. It can’t compress every equipment lead time, construction schedule and installation dependency into the quarter when customers want the chips.
Agents Confront Production Work
A new benchmark offers a useful reality check on what happens after an agent can write plausible code.
Researchers have released tau-to-the-tau Bench, a preprint benchmark that asks a coding agent to build and deliver an end-to-end customer-service agent. The work starts with business records, requirements held by a simulated client and an inherited codebase. It continues through a production API and decisions constrained by serving cost.
That is a wider job than solving a specified coding exercise. The agent has to understand existing records, communicate with the client, settle requirements, choose an architecture, control cost and deliver reliable behaviour.
Across 53 tasks in four domains, the strongest tested configuration, Claude Opus 5 running under Claude Code, passed 23.9 per cent of held-out evaluation simulations. The expert-authored reference ceiling was 82.2 per cent. The authors report recurring failures in deep record comprehension, client communication and experimentation with architecture and serving spend.
Those numbers need careful boundaries. This preprint uses one benchmark design with simulated users and hasn’t been peer reviewed. It doesn’t establish a universal failure rate for coding agents; other models, harnesses or tasks could produce different results. The expert reference is also described as a ceiling within this evaluation, not proof that real production delivery tops out at the same level.
Even so, the failure modes look familiar to anyone who has delivered software. Generating a working component is only part of the job. A developer has to discover what the customer means, reconcile that with an existing system, test competing designs and understand the service’s running cost. An agent that skips one of those loops can produce code that looks complete while delivering the wrong behaviour or an uneconomic system.
This benchmark is valuable because it puts those loops inside the evaluation. Conventional coding scores can tell you whether a model repairs a function or passes a known test. They say much less about whether it notices a missing requirement, asks a useful question or changes architecture when the first serving plan is too expensive.
For developers and engineering leaders, a useful test of a coding agent now resembles the work you expect it to own. Give it the inherited repository and the real constraints. Assess the questions it asks, the design decisions it records, the cost it incurs and the deployed behaviour, not merely whether the generated code runs.
My takeaway is that the production gap sits around the code as much as inside it. A 23.9 per cent result doesn’t predict how an agent performs in every organisation, but it does puncture the idea that stronger code generation automatically produces dependable end-to-end delivery. Requirements discovery and system judgement remain the hard part.
A Read-Only MCP Boundary Fails
Here’s a smaller release with a very concrete lesson about where agent permissions should live.
AWS has disclosed CVE-2026-85787 in the awslabs postgres-mcp-server before version 1.1.7. The Model Context Protocol server is intended to give an AI application controlled access to PostgreSQL. According to the bulletin, crafted SQL placed into content later submitted during an authenticated user’s interaction could allow an unauthenticated actor to modify data beyond the server’s intended read-only scope.
AWS describes the outcome as something that might occur. It doesn’t report known exploitation or say how many deployments are affected. Version 1.1.7 contains the fix, and AWS recommends upgrading the package as well as patching forks and derivatives.
The more durable advice is to use a dedicated, least-privilege PostgreSQL role. If the database account itself has read-only permissions, the database enforces the boundary even when validation in the MCP server fails. If that account can write, the application’s SQL blocklist becomes the main defence, and this vulnerability shows how brittle that can be.
Developers using this server should therefore check two things: the installed version and the privileges of the database role behind it. Updating closes the disclosed validation gap. Restricting the role limits what a future gap can do.
That separation is especially important for agent tools, where instructions and untrusted content can meet executable queries. My view is simple: describe permissions in the agent layer, but enforce them at the underlying system. A read-only label is an intention. A database role is a boundary.
npm Tightens Trusted Publishing
Package publishing has also picked up a useful security improvement, without forcing every release through one workflow.
npm packages can now use multiple trusted-publishing configurations based on OpenID Connect, or OIDC. Each configuration specifies its own repository, workflow and environment criteria. A matching short-lived identity token from any one authorised configuration can stage or publish the package.
That helps maintainers who separate stable releases, prereleases and staged workflows. Previously, supporting those paths could leave a reason to retain long-lived publishing tokens. Multiple OIDC configurations let each approved workflow prove its identity when it runs, without sharing one persistent secret across them.
Staging is available by default and adds human approval before publication. A staged release can’t be approved until malware scanning finishes, and maintainers can inspect the history of staged versions. Direct publishing remains an opt-in choice for each configuration.
The word additive matters here. Every new configuration creates another authorised route to the package. A maintainer who adds several workflows without auditing their repository, workflow and environment rules may widen the release surface even while removing long-lived tokens.
For package owners, the stronger pattern combines short-lived OIDC identity with staged review. The workflow proves where it came from; the malware scan runs; a person can inspect the staged version before it becomes public. That doesn’t quantify how much compromised-package risk falls, and GitHub’s announcement provides no adoption figures. It does remove a practical obstacle to keyless publishing across more complicated release setups.
I’d treat the new flexibility as a reason to map every authorised workflow, not as permission to add them casually. Keyless credentials reduce one class of secret risk. The approval path still needs a clear owner, because trust spread across forgotten configurations is still trust spread too far.
What Changes for You
One release this week could change the ordinary rhythm of a long coding task, provided its cloud trade-offs fit your work.
Kiro Web and cloud sessions are now generally available. A developer can start work in a Kiro cloud sandbox, disconnect, and let the session continue. The same session can later be resumed from the web, IDE, command line or mobile, so the local laptop and terminal no longer have to stay awake for the job to keep running.
Cloud configuration carries steering, custom agents, skills, powers and hooks between local and remote work. That continuity is the practical gain. A longer refactor or investigation can move into persistent execution without making the developer rebuild the agent’s working setup each time they change device.
Access is narrower than the phrase generally available might suggest. Cloud sessions are limited to paid Pro, Pro Plus, Pro Max and Power subscriptions, and they run in US East, in Northern Virginia. Sessions consume shared credits. For organisations, an administrator has to enable the capability, with supported sign-in options including AWS IAM Identity Center, Okta and Microsoft Entra ID.
There’s no independent evidence in the briefing about reliability, security or code-quality outcomes. The vendor documentation establishes availability and the stated behaviour. It doesn’t show that a background session will make better engineering decisions than a local one.
The immediate difference is persistence. Developers who already use Kiro can hand off a long task, close the machine and return from another surface. The limitation is that source code, configuration and execution move into a vendor-controlled service in one US region, with subscription and credit costs attached.
For an individual developer, that can be a genuine convenience. For an organisation, the decision belongs with whoever owns source-code handling, identity and cloud-region policy. My judgement is that durable background execution is becoming a meaningful feature of coding agents, but the value comes from continuity, not autonomy by itself. The agent still has to do sound work when nobody is watching.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- datacenterdynamics.com/en/news/anthropic-signed-517bn-in-compute-agreements-in-past-11-months
- questions-statements.parliament.uk/written-statements/detail/2026-09-07/hcws314
- openai.com/index/an-alien-mind
- bloomberglinea-bloomberglinea-prod.cdn.arcpublishing.com/negocios/tsmc-eleva-casi-al-doble-su-prevision-de-compras-de-herramientas-para-chips-de-ia
- arxiv.org/abs/2609.04611
- aws.amazon.com/security/security-bulletins/2026-101-aws
- github.blog/changelog/2026-09-03-multiple-trusted-publishing-configurations-for-npm
- kiro.dev/blog/agentic-engineering-in-the-cloud