Back to the show

AI & Tech Daily

When synthetic media enters the courtroom

17:53

An Arizona appeals court has vacated a manslaughter sentence after a judge treated an AI-generated video of the deceased victim as genuine. Jesse examines the line between an authentic record, advocacy and synthetic reconstruction, then covers Google's pause on one class of open-source bug-bounty reports, transactional queues in Spanner, CoreWeave's Forge platform, Aleph Alpha's Kolibri-1 model and Nvidia's new 64 GB DGX Spark. In What Changes for You: GitHub brings Copilot code-review requests to APIs, while Perplexity offers typed probabilities instead of chatbot prose through its Decisions API and an open-weight model.

Full transcript

Read the episode.

I'm Jesse Owen. This is AI and Tech Daily.

A Synthetic Victim in Court

A video made a deceased victim appear to speak at sentencing. The judge called it genuine, and an appeals court has now ruled that reliance made the proceeding unfair.

Our main story today looks at that Arizona ruling and where a court must draw the line between an authentic record, a family's advocacy and a generated reconstruction.

On September 30, the Arizona Court of Appeals affirmed Gabriel Horcasitas's manslaughter conviction but vacated his sentence. The case now goes back for resentencing. The conviction itself remains intact, which is an important boundary around what the court decided.

The sentencing presentation mixed two very different kinds of material. It contained genuine footage of the victim, and it contained synthetic speech and behaviour based on the victim's sister's interpretation of what he might have wanted to say. The appellate court found that the sentencing judge's description of the synthetic victim as genuine showed the generated material had affected the proceeding.

That distinction is the heart of the problem. A family member can tell a court about grief, loss and the effect of a crime. A lawyer can make an argument. Genuine recordings can show a person as they were. A generated likeness does something else: it gives someone else's interpretation the face, voice and apparent authority of a person who cannot confirm, reject or qualify it. The emotional force comes partly from presenting a reconstruction as if it were the source.

The difficult part isn't only detecting whether pixels or speech were generated. The court knew it was looking at a presentation assembled for sentencing. The failure identified on appeal was one of treatment: the synthetic performance was described as genuine and influenced the result. A perfectly labelled clip could still blur roles if the person deciding the case hears an advocate's interpretation in the victim's simulated voice. Provenance tells you how the material was made. It doesn't decide what weight the material deserves.

That gives judges a practical set of questions before such media is shown. Which moments are authentic recordings? Which words or actions were generated? Whose interpretation supplied them? And is the clip helping the court understand a living witness's statement, or presenting a new statement that the depicted person never made? Those questions don't require a court to reject every use of visual technology. They require the court to keep the source of each claim visible while the emotional presentation is doing its work.

This ruling doesn't create a nationwide ban on synthetic media in court. It comes from an intermediate appellate court in Arizona, and the briefing doesn't establish whether there will be further review. Nor did the court throw out the conviction. Its remedy was narrower and tied to the fairness of sentencing.

Even with those limits, courts and lawyers now have a concrete example of what can go wrong. Labelling generated material is necessary, but the deeper question is what role it is allowed to play. Is it evidence, an illustration of a relative's statement, or a persuasive device? Those categories carry different weight, particularly when a judge is deciding a person's sentence.

My read is that institutions need an explicit boundary before the next emotionally powerful synthetic reconstruction arrives. Keep authentic records distinct, attribute advocacy to the living person making it, and don't let generated speech borrow the authority of the dead. The technology can make an interpretation vivid. It can't make that interpretation genuine.

Bug Reports Meet the Proof Test

That courtroom dispute is unusually stark, but the same verification gap is showing up in much more routine technical work.

Google stopped accepting new product-vulnerability submissions through the product-flaw category of its Open Source Software Vulnerability Reward Program on October 1. Outstanding reports are still being handled, and supply-chain reports remain open. Some flaws in relevant Google Cloud repositories may also qualify under the separate Cloud vulnerability programme. So this is a pause in one category, not the closure of Google's whole bug-bounty operation.

Google says it is restructuring this part of the programme and plans to provide an update by the first quarter of 2027. That's an update window, not a promised reopening date. Researchers with a new finding need to check whether another eligible programme applies or use Google's route for submitting a patch.

The background explains why this is more than an administrative tidy-up. In March, Google said it was receiving a large volume of AI-generated reports containing incorrect or non-exploitable claims. It responded by tightening the rules around reproductions and patch evidence. Now the company has temporarily stopped the affected intake category altogether. It hasn't disclosed detailed volumes or false-positive rates, so we shouldn't manufacture a precise causal scorecard. The sequence still exposes the cost imbalance.

AI can produce plausible defect reports cheaply. A maintainer still has to inspect the code path, establish whether an attacker can reach it, reproduce the behaviour and judge the impact. A report that sounds technical but can't clear those steps creates work without reducing risk. Scale the generation side and the verification queue becomes the bottleneck.

For security researchers, the practical signal is that a possible defect is no longer much of a deliverable on its own. A minimal reproduction, reachability evidence or a sound patch carries more value because it moves the claim towards a decision. For maintainers, the pause offers some relief, though potentially at the cost of delaying valid reports that would have used this route. Automated discovery can widen the search. The scarce skill is still proving that a finding is real, exploitable and worth fixing.

Transactions Become the Queue

Now to a less dramatic problem that can still ruin an AI workflow: a request without its matching state change, or the other way around.

Google Cloud has made Spanner queues generally available. An application can update its database state and enqueue asynchronous work within the same Spanner transaction. That atomic step removes a familiar failure gap between writing to the database and publishing a message to a separate queue.

Picture an agent handing a case to a reviewer. The application records that the case is awaiting approval, then schedules a timeout or follow-up job. If those actions happen in separate systems, a crash between them can leave the database saying one thing while the queue says another. Developers often deal with that through an outbox table and reconciliation logic. Spanner queues bring the state change and the enqueue operation under one transaction.

The queue records support scheduled delivery, streaming SQL consumption, leases, lease renewal and transactional acknowledgement. Operators can inspect queue state with SQL instead of treating the messaging layer as a separate opaque store. That's useful for delayed retries, approval deadlines and agent hand-offs, where knowing why a task is waiting can matter as much as knowing that it exists.

There is still an important delivery detail. Google guarantees at-least-once delivery and at-most-once acknowledgement. A worker may therefore see the same task more than once. Exactly-once processing still depends on transaction design, idempotent handlers and careful treatment of external effects. If a worker charges a card or calls another service twice, the queue can't undo that by itself.

For organisations already on Spanner, this could remove a separate queue and some of the plumbing that keeps an outbox consistent with it. The trade-off is equally clear: the capability is tied to Spanner, so its cost and migration considerations remain. My practical take is that this is valuable infrastructure for reliable agents precisely because it isn't about model intelligence. Atomic state transitions, controlled retries and an auditable record are what turn a fluent response into dependable work. Independent production-scale comparisons with dedicated queue systems remain worth watching.

CoreWeave Moves Up the Stack

Infrastructure providers don't want to stop at renting the accelerators, and CoreWeave has made its next move quite explicit.

On September 30, CoreWeave launched Forge, a joined-up environment for model serving, production tracing, data curation, training and evaluation. It combines CoreWeave services with technology from Weights & Biases, OpenPipe and marimo. Two of the new pieces are Agent Lens, for tracing and monitoring agents, and Registry, for keeping track of models, agents and datasets.

The pitch is a closed improvement loop. An organisation observes how a model or agent behaves in production, finds weak cases in the traces, turns those cases into evaluation or training data, makes a change and measures the result. Many AI teams already do those jobs, but they move artefacts and context among separate systems. Forge is CoreWeave's attempt to keep the chain connected. The company named Canva and MasterClass as early users and opened both free and paid access at launch.

CoreWeave says Forge can connect with workloads using other models, frameworks and clouds. Its own cloud is still the native operating environment, and there isn't yet independent evidence showing how portable complete workflows are across providers. The claim that integration shortens improvement cycles is also the company's claim, not an independently measured result.

For an AI organisation, the useful question isn't whether a unified console looks tidier. It's whether production traces remain linked to the dataset, evaluation, model version and deployment decision that followed. That lineage can reduce the manual work of diagnosing a failure and establishing whether a change fixed it.

The strategic consequence is a new kind of cloud dependence. Competition is moving above raw GPU supply into the evaluation and improvement loop, where a platform can become embedded in how a team learns from production. Forge may reduce operational friction for organisations that want that integration. Early export and cross-cloud tests can expose the hardest lock-in to unwind: the history connecting traces, datasets and decisions, not the compute job itself.

Kolibri Targets Bilingual Control

For developers who want more control over the model itself, a new open-weight release is aimed squarely at German-English work.

Aleph Alpha released Kolibri-1 on October 3. It's a bilingual German-English reasoning model with tool calling, downloadable weights and an Apache 2.0 licence covering the published weights and configuration files. The licence does not extend to the company's architecture, training methods or other intellectual property.

Kolibri uses a mixture-of-experts design. It has 78.1 billion parameters in total, but activates about 3.46 billion for each token. Sparse activation is meant to provide more model capacity without using every parameter on every step. That doesn't make the deployment small. The FP8 release occupies about 78 gigabytes, and Aleph Alpha lists at least two 80-gigabyte A100s, two H100s, or one H200, B200 or B300 among the supported hardware configurations. This is data-centre or serious lab equipment, not a casual laptop download.

The context figures also need a careful reading. Aleph Alpha reports validation up to 1,048,576 tokens, while recommending no more than 262,144 tokens for efficient serving and complex work. A maximum validated window doesn't tell you the latency, memory use or reasoning quality you'll get across that whole span. The practical limit for a real application may be much lower.

Kolibri's more credible point of difference is deployable bilingual control. A company handling German and English material can run the weights in an environment it controls and add tool use without sending each request to a frontier service. The sparse design could make that large total parameter count more manageable at inference time, though the published capability and efficiency figures still come from Aleph Alpha and need broader independent evaluation.

I wouldn't treat this release as a settled performance winner. For developers with the hardware and a genuine sovereignty requirement, though, it adds a technically interesting option where local control, German-English coverage and tool calling matter more than a generic benchmark crown.

A Smaller DGX Spark

Local control also depends on what can fit beside a developer's desk, and Nvidia is lowering the memory entry point rather than the price expectation.

Nvidia announced a 64-gigabyte unified-memory version of its DGX Spark desktop system on October 2. Manufacturing partners are due to begin offering it on October 23, with a starting price of 4,999 US dollars. It keeps the Grace Blackwell GB10, ConnectX-7 networking, DGX OS and Nvidia's AI software stack, but carries half the memory of the existing 128-gigabyte configuration. It is not shipping yet, and partner-specific Australian pricing and availability haven't been announced.

Nvidia says one unit can support models up to 100 billion parameters. That headline needs the usual memory arithmetic. Whether a model fits and runs usefully depends on its precision, context length, cache, runtime overhead and the workload around it. Fitting weights into memory isn't the same as delivering comfortable interactive performance.

Two 64-gigabyte systems can pool 128 gigabytes through Nvidia's Sync Cluster Assistant. The hardware bill starts at 9,998 US dollars before networking accessories. Nvidia reports up to 1.7 times the performance of one system in its own Qwen 3.8 27B test, rather than the two-times result somebody might assume from buying two boxes. That figure is vendor testing, and one workload doesn't settle performance across models.

The release does reflect a real shift in local AI: smaller and more efficient models can make 64 gigabytes useful for serious development. But memory capacity remains a hard constraint, and Nvidia's managed desktop experience carries a professional price.

For a developer or small lab already committed to Nvidia's software stack, the new configuration offers a lower-cost way into that environment and a path to pairing two machines later. For everyone else, the purchase decision should start with the exact model, precision and context they need to run. The label saying 100 billion parameters is a ceiling claim, not a description of the experience they'll get.

What Changes for You

Two developer releases are useful now because they replace awkward manual steps with interfaces ordinary application code can call.

GitHub has made Copilot code-review requests available through its REST and GraphQL APIs on Copilot Pro, Pro+, Max, Business and Enterprise plans. Developers can trigger a review from a script, workflow or internal tool and choose the effort level for that request. Separately, Balanced became the default effort on September 28 for repositories and organisations that had left the setting on Default. An explicitly selected Lite setting was preserved, and lower-level settings can override an organisation or enterprise choice.

That makes Copilot review a workflow primitive rather than something that begins only in GitHub's interface. Administrators should check the effective setting for important repositories, because GitHub hasn't quantified how Balanced changes review time, consumption or defect detection compared with Lite. Automating the request is useful; a developer still owns the merge decision.

Perplexity has taken the opposite approach to model scope. Its Decisions API accepts text, JSON or images and returns probabilities for yes-or-no questions, fixed choices or rubric scores. It doesn't draft prose, explain an answer or generate code. Application logic interprets the probabilities and sets the threshold. Up to 128 questions can go in one request, with a limit of ten requests per second per organisation. The hosted price is 0.04 US dollars per million input tokens, with no output-token charge, and the sole hosted model accepts fewer than 262,144 input tokens.

Builders can use those typed probabilities for classification, routing and grading without parsing a chatbot's sentence into a decision. Perplexity also released the underlying 27-billion-parameter weights under Apache 2.0. Self-hosting needs Python 3.12 or newer and a CUDA GPU with roughly 49 gibibytes available for the weights and working memory. The limitation is the point: there is no per-decision explanation, and Perplexity's comparisons are vendor-run. A team would need to test calibration on its own data and give uncertain cases somewhere sensible to go.

You'll find the sources and full transcript at owenonthenet.com. Thanks for listening.

Sources

Reporting behind this episode.

  1. coa1.azcourts.gov/Portals/1/OpinionFiles/Div1/2026/State%20v.%20Horcasitas%20-%201%20CA-CR%2025-0191%20-%20Opinion.pdf
  2. apnews.com/article/cd1ca553c7fa80c6698d7f97b51b1edd
  3. tomshardware.com/tech-industry/artificial-intelligence/google-suspends-part-of-the-oss-vrp-bug-bounty-program-due-to-an-influx-of-invalid-ai-submissions-product-vulnerability-submissions-ended-october-1
  4. bughunters.google.com/blog/ossvrp-rule-updates-2026
  5. cloud.google.com/blog/products/databases/spanner-queues-provide-native-transactional-messaging
  6. coreweave.com/news/coreweave-forge-launches-turning-the-ai-loop-production-run-into-a-better-model-and-agent
  7. aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model
  8. huggingface.co/Aleph-Alpha/Kolibri-1
  9. blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync
  10. github.blog/changelog/2026-10-02-copilot-code-review-api-support-and-new-default-effort-level
  11. docs.perplexity.ai/docs/decisions/quickstart
  12. huggingface.co/perplexity-ai/pplx-decider-v1-27b
  13. x.com/perplexitydevs/status/2105725598882832414