AI & Tech Daily
Huawei Rethinks the Connections Inside Giant AI Systems
Huawei puts optical interconnects at the centre of its next large AI systems, and resident expert Maya Chen explains what buyers should watch before the first Ascend 960 chips arrive in 2027. Jesse also covers Gemini 3.8 Live's concurrent tool use, Australia's consultation on AI training and data centres, the proposed EU KIDS Act, an actively exploited Cisco firewall-management flaw, Google's anomaly detection for enterprise agents, Anthropic's biomolecular-model optimisations, and an open-source SDK generation stack from Google and Speakeasy.
Full transcript
Read the episode.
I'm Jesse Owen. This is AI and Tech Daily.
Huawei's Optical AI Bet
Huawei says one new AI system could connect 4,096 accelerators while cutting more than 550 kilowatts from the interconnect's power demand. The catch is that its central chips aren't due until 2027.
That gap between a striking system design and hardware that buyers can deploy is why this development deserves a closer look. At its September 17 event, Huawei unveiled the Atlas 960E SuperPoD, a design that brings processors, high-bandwidth memory and storage into one very large AI system. A SuperPoD is a cluster engineered to behave more like one computing unit than a collection of separate servers.
Huawei says a single Atlas 960E can contain up to 4,096 Ascend neural processing units, delivering eight exaflops of FP8 compute and one petabyte of high-bandwidth memory. FP8 is an eight-bit number format used to reduce the work and memory needed for many AI calculations. The company also described a million-NPU architecture built by linking these systems at a much larger scale. The Associated Press independently confirmed the launch and Huawei's positioning of Atlas 960 as a faster successor for training and inference.
The distinctive part is Hi-ONE, Huawei's near-packaged-optics design. Instead of relying on tens of thousands of conventional optical modules spread through the system, it moves more of that optical connection closer to the computing hardware. Huawei claims this removes those modules and saves more than 550 kilowatts of power per SuperPoD. Those performance and power figures are company claims, not independent test results.
Timing matters as much as architecture. The Ascend 960DT training chip is scheduled for the first quarter of 2027, with the 960PR inference chip scheduled for the third quarter. Production availability, software compatibility and reliability still have to be demonstrated.
For infrastructure buyers, my read is that Huawei has presented a credible alternative design to assess, not a finished capacity plan to buy today. Optical links could make giant clusters more efficient, but procurement decisions need independent results from complete systems, using real workloads, after the roadmap becomes hardware.
Jesse: Joining me is Maya Chen, a Senior AI Analyst and one of the show's AI resident experts. Maya, welcome.
Maya Chen: Thanks, Jesse. It’s lovely to be back, and good to be on the show.
Jesse: It’s good to have you back. The AI news cycle has clearly ignored our request for a quiet week.
Maya Chen: Completely ignored it. At this rate, a quiet afternoon would feel ambitious.
Jesse: Let’s get into Huawei’s new Atlas 960E SuperPoD. It brings processors, memory and storage into a very large AI system, with optical links doing more of the connecting. When you look past the sheer scale of the announcement, what stands out to you first?
Maya Chen: The optical interconnect is the interesting move. Once you’re joining thousands of accelerators, the connections aren’t background plumbing; they shape how much useful computing the whole system can deliver and how much power it consumes. Huawei says a single Atlas 960E can hold up to 4,096 Ascend NPUs, with eight EFLOPS of FP8 compute and one petabyte of high-bandwidth memory. Bringing that scale together as one integrated design could offer another route for building giant AI clusters.
Jesse: That shift from plumbing to a core part of the machine is striking. Huawei says Hi-ONE replaces tens of thousands of conventional optical modules and cuts power use by more than 550 kilowatts per SuperPoD. In practical terms, how could changing the interconnect make a system that large easier to run?
Maya Chen: It could reduce both the electrical load and the complexity sitting between the chips. Conventional optical modules add power demand and a huge number of separate components to manage. Huawei’s near-packaged-optics approach moves those connections closer to the computing hardware. If production systems deliver the claimed saving, operators would have more of their power budget available for computation and fewer conventional modules in the design. But that benefit still needs to show up in complete systems running real training and inference workloads.
Jesse: And for a buyer, efficiency on a slide only becomes useful when the system is reliable and the software fits what they already run. With the main chips still due next year, what should businesses do with this roadmap now, and when might ordinary people notice any difference?
Maya Chen: Businesses can add it to their future-capacity assessments, but the central chips aren’t deployable today. The Ascend 960DT training chip is scheduled for the first quarter of 2027, and the 960PR inference chip for the third. Buyers should watch whether those dates hold, then look for independent results on performance, power use, reliability and software compatibility. For ordinary people, the effect would come later through the AI services built on this infrastructure. If the systems prove efficient and dependable, they could influence the cost, availability and responsiveness of those services, rather than changing a phone or app overnight.
Jesse: Maya, thanks for making a very large system feel understandable. I’ve appreciated your time and insight, and I hope you’ll return soon.
Maya Chen: Any time, Jesse. I’ll come back as soon as the news cycle grants us that quiet afternoon.
Jesse: Well, that was Maya Chen. I hope you found that insightful. And now back to the news that's changing the world today: Gemini 3.8 Live brings concurrent tool use to Google's voice AI.
What Changes for You
That brings us neatly from the hardware underneath AI to a voice model developers can start testing now.
Google began rolling out Gemini 3.8 Live and a separate Gemini 3.8 Live Extended Thinking model on September 15. Both are available to developers through the Gemini API and Google AI Studio. The practical change is concurrent tool use: the standard model can continue a voice conversation while it executes tools or API calls in the background. The user no longer has to sit through dead air every time the agent checks another system.
Google says the standard model can also use visual context and switch automatically among 97 supported languages. Consumer access is more fragmented. Standard Live is appearing in Search Live, while Extended Thinking reaches Gemini Live and selected Workspace surfaces for Google AI subscribers. The rollout is staggered, so the product and subscription determine what someone can use. Google also says audio produced by its AI products carries SynthID watermarking.
For voice-agent developers, this makes a less rigid interaction possible today through the API and AI Studio. It also raises a design problem. A smooth conversation can hide a slow, mistaken or unsafe action happening behind it. My practical judgement is that status, cancellation and confirmation controls now become more important, not less. Google hasn't published independent production-reliability evidence across the claimed languages and tool workflows, so builders still need to test the messy cases before treating continuous speech as proof that the work underneath is going well.
Australia Consults on AI Infrastructure
Back on the ground, Australia is asking what the physical growth of AI should require from the companies building it.
The Australian Government opened a public consultation at 6 am AEST on September 18 about possible standards for AI model training and large data-centre infrastructure. It closes at 5 pm AEDT on October 9. That's a short, defined window for operators, infrastructure providers and affected communities to put evidence into the policy process.
The questions reach beyond computing performance. The consultation covers energy and renewable-energy requirements, water use, where projects are located, their effects on communities, workforce skills and the thresholds that could determine which projects face requirements. Submissions will inform government advice and later decisions. The consultation itself creates no new legal obligation, and the government hasn't committed to final rules, thresholds or an implementation date.
That distinction is important for projects already being planned. A developer doesn't have a new compliance checklist this morning, but it does have a clearer view of the issues government is considering. Water, power, location and local workforce effects could become part of the approval or operating environment rather than side questions dealt with after a facility is designed. The threshold question is especially consequential: it could determine whether requirements apply only to the largest campuses or reach a wider set of training facilities. The consultation asks the question but doesn't answer it.
Communities also have a formal route to describe effects that a national infrastructure plan can otherwise flatten into capacity figures. Their submissions can address local water, project location and workforce concerns alongside industry evidence about what is technically feasible. That doesn't guarantee a particular policy outcome, but it gives the later advice a broader evidence base.
For data-centre and AI companies, there's value in putting real project data into the consultation before expensive designs are locked in. My assessment is that the immediate cost is uncertainty: nobody yet knows which proposals will become standards or regulation. Still, early uncertainty is more manageable than discovering the requirements after land, grid capacity and cooling systems have already been committed. Organisations with Australian projects now have until October 9 to explain where practical thresholds should sit and what the trade-offs look like in operation.
Europe Proposes Age Rules for Platforms
A different policy question is taking shape in Europe: who should carry the burden when children use digital services?
The European Commission adopted its proposed EU KIDS Act on September 17. The proposal covers online platforms as well as risky digital services and AI systems used by minors. It would prevent children under 13 from accessing social media and set 15 as the EU-wide minimum age for opening an account independently. A child between those ages couldn't independently open the account under the proposal, while the under-13 bar for social media would be stricter.
The other substantial shift is responsibility. Service providers would have to show that their products are age-appropriate and safe by design. That moves the work towards the platform: assessing likely use by children, producing safety evidence and, where necessary, checking age. Today, however, this remains a legislative proposal. It doesn't change the current rules, and both the final obligations and the date they might apply can change during the legislative process.
A common EU threshold could make expectations more consistent than a patchwork of national age limits. It could also affect more than account-registration screens because the proposed scope includes risky digital services and AI systems used by minors. Providers would need to understand which parts of a product children encounter and what evidence supports the safety of that experience.
The hard part is implementation. Stronger age assurance can ask people to disclose more information, exclude users who struggle with the chosen method, or add cost and friction to ordinary account creation. Those aren't arguments against protecting children; they're engineering and privacy constraints legislators and platforms will have to resolve. A system can be strict on paper and still be poor if it gathers excessive identity data or fails for legitimate users.
Platforms serving European users shouldn't present the proposal as settled law, but they can sensibly map where their existing account, age-assurance and child-safety evidence would fall short. My read is that waiting for the final vote before examining those systems would make adaptation harder. The useful work now is identifying which safety claims can be demonstrated without collecting more personal data than the service needs.
Cisco Firewall Flaw Is Under Attack
Now for an issue where waiting is the wrong response. Cisco has confirmed that attackers are exploiting a critical firewall-management flaw.
Cisco updated its advisory for CVE-2026-20079 on September 16. The vulnerability affects Secure Firewall Management Center systems and carries the maximum CVSS score of 10. An unauthenticated remote attacker can bypass authentication and execute scripts or commands with root privileges. Because the target is the management plane, successful access can put the system used to control firewalls into an attacker's hands.
Cisco says there is no workaround. Earlier preventive hotfix guidance has been replaced by fixed, hardened software releases, and affected operators should move to one of the versions listed in Cisco's advisory. Restricting public access to the management interface can reduce exposure, but it doesn't remove the vulnerability. It is a useful exposure-reduction step, not a substitute for installing a fixed release.
An upgrade is only half the job because Cisco has confirmed active exploitation. The company provides a log indicator that may show compromise and tells customers to contact its technical assistance centre if they suspect exploitation. Cisco hasn't said how many customers were compromised or identified them, so the public evidence doesn't tell us how broad the campaign is. Nor does the absence of a public customer count reduce the consequence for an operator whose system was reachable.
The sequence matters. Operators need to identify affected systems, install the appropriate fixed release, and review the available evidence for activity that occurred before the upgrade. If exploitation is suspected, Cisco's technical assistance route is part of the response. Treating the update time as the start of the incident window could miss earlier access.
For security and network operators, I would treat this as a potential incident rather than a difficult maintenance window. Patch promptly, then examine the available logs and follow the compromise path if the indicator or other evidence appears. A fixed release prevents the known route from remaining open; it doesn't undo commands an attacker may already have run as root. That combination of remote access, no authentication and management-level control is what makes the forensic check necessary.
Monitoring Agents After They Act
The security problem changes shape when software agents can call tools on their own. Google is testing a monitor built for that behaviour.
Agent Anomaly Detection entered private preview on September 16 for customers using the Gemini Enterprise Agent Platform with Agent Development Kit version 1.2 or later. It examines OpenTelemetry traces, tool calls and the path an agent took through a task. OpenTelemetry is a common format for recording what software did, how long it took and where failures occurred.
Google describes the service as an out-of-band layer. It reviews activity alongside or after agent sessions rather than inserting another model into the live request path. A statistical first pass filters sessions, then an LLM analyses the selected activity for unusual behaviour. That architecture avoids adding latency to every interaction. Findings are available through an API, and a separate callback or plugin can use them to block later tool calls.
This is different from watching only for crashes or failed API requests. An agent can complete a tool call successfully while the sequence of calls or its route through the session still gives the detector something unusual to examine. Those traces provide more behavioural context than a single error code. Google hasn't provided independent evidence showing how reliably that distinction works in production.
The timing creates the central limitation. A monitor may discover suspicious behaviour only after an agent has completed the damaging action. Developers have to connect findings to a useful response, such as restricting later tools or changing access, for detection to become enforcement. That also means deciding how much confidence is enough to limit an agent without burying operators in false alarms. Google hasn't published independent figures for detection quality, false positives or operating cost, and custom detectors for particular business rules are still described as forthcoming.
For eligible enterprise developers, my judgement is that this could reveal misuse that ordinary error logs miss without slowing the user-facing path. It isn't a safety switch by itself. Teams in the preview still need to decide which actions demand live approval, which can be reviewed asynchronously, and what an anomaly finding should disable before the next session begins.
Claude Optimises Biomolecular Models
There is also a useful example of coding agents working well beyond ordinary application code.
Anthropic has released optimised implementations for more than 30 open-source biomolecular models. The company says an internal Claude research model produced the changes under the supervision of two Anthropic technical staff over just under four weeks. These are models used in biomolecular research, where inference speed and GPU memory can limit the size or number of experiments a researcher can run.
Anthropic reports about a fourfold average speed-up with a small reduction in precision, and nearly a twofold improvement for optimisations whose outputs remained identical. It also says a new low-memory mode handled biomolecular systems longer than 10,000 tokens on a single Nvidia GPU node. The released code means researchers can inspect the changes rather than relying only on a description of the work.
The qualifications are substantial. Anthropic measured the results itself, the Claude model involved is internal, and the precision trade-off depends on the workload. In scientific computing, a faster answer isn't automatically an acceptable answer if a changed numerical result affects a validated pipeline.
My take for biomolecular researchers is straightforward: this release is worth benchmarking against their own models, inputs and accuracy tolerances. If the reported gains reproduce, larger or faster experiments may need less compute. The open code makes that claim testable, but local validation decides whether an optimisation belongs in research work where numerical behaviour already has an established baseline.
Open-Source SDKs for Agents
One last developer story deals with a quieter dependency that agents increasingly rely on: the code connecting them to APIs.
Google says Speakeasy's OpenAPI client-generation suite is now open source under the AGPL version 3 licence. The companies used the suite while migrating Google's generative-AI SDK pipeline. It can generate typed software-development kits for Python, TypeScript, Go, Java, C sharp, PHP and Ruby, with support for streaming, retries and pagination.
The agent-specific pieces are notable. The stack can produce command-line tools that coding agents can operate and MCP servers that expose API documentation and schemas. For an API owner, that creates a route from an OpenAPI definition to the SDKs, tools and structured context an agent needs, all inside a build system the owner can run.
Licensing needs a careful read. Owners may choose the licence for SDKs generated by the software, but modifications to the generator itself remain subject to the AGPL. Running the stack also means taking responsibility for its upgrades, build integration and compatibility. Google and Speakeasy haven't yet demonstrated what long-term community governance or maintenance will look like outside Google's tested pipeline.
For API teams, my assessment is that the open release can reduce dependence on a hosted proprietary generator at a layer that is becoming more operationally important. The trade is ownership: control over the pipeline comes with the work of governing it and understanding the copyleft terms. That exchange is most attractive where deterministic, repeatable agent tooling matters enough to justify maintaining the machinery.
You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.
Sources
Reporting behind this episode.
- huawei.com/en/news/2026/9/hc-wang-keynote
- apnews.com/article/26ab418df1339c518483918218ffbe57
- blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking
- pmc.gov.au/news/have-your-say-future-ai-training-and-infrastructure-australia
- digital-strategy.ec.europa.eu/en/news/eu-kids-act-restrict-social-media-platforms-access-children-eu
- sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-onprem-fmc-authbypass-5JPp45V2
- developers.googleblog.com/agent-anomaly-detection-now-in-private-preview-on-the-gemini-enterprise-agent-platform
- anthropic.com/research/claude-uplifts-biomolecular-modeling
- developers.googleblog.com/why-client-sdk-generation-belongs-in-the-open