---
title: "Nobody Has Asked a Judge to Read the System Card"
description: "In Raine v. OpenAI, the wrongful death case now sitting in California's coordinated AI litigation, both sides have built their case on documents called system cards. These documents describe how OpenAI's models perform, what risks they carry, and what testing"
url: https://kaynemcgladrey.com/blog/nobody-has-asked-a-judge-to-read-the-system-card/
date: 2026-10-05
modified: 2026-10-05
author: "Kayne"
image: https://kaynemcgladrey.com/wp-content/uploads/2026/10/Judge-Ethan-P.-Schulman.webp
categories: ["Blog"]
type: post
lang: en-US
---

# Nobody Has Asked a Judge to Read the System Card

In *Raine v. OpenAI*, the wrongful death case now sitting in California’s coordinated AI litigation, both sides have built their case on documents called system cards. These documents describe how OpenAI’s models perform, what risks they carry, and what testing preceded release. The plaintiffs cite them as proof that OpenAI knew its safety testing was inadequate, while OpenAI cites the same documents as proof that it tested responsibly and shipped safeguards.

More than a year into this litigation, no judge has interpreted a single word of any of it. The case was coordinated into [Judicial Council Coordinated Proceeding](https://www.sdcourt.ca.gov/sdcourt/civil2/jccp2) No. 5431, known as [JCCP 5431](https://techjusticelaw.org/wp-content/uploads/2026/03/OAI_JCCP_Order_re_Petition_for_Coordination_-_5431.pdf), stayed, and parked. Everything we know about how courts will read AI safety documentation comes from two groups of lawyers shouting past each other in pleadings.

If you build or deploy AI, this is the case to watch, because the eventual ruling decides whether your transparency documentation protects you or exposes you.

## What’s In The Cards

A system card describes a deployed AI system. A model card describes the weights. The distinction sounds pedantic until you realize that most of the safety machinery a user encounters, the routing, the refusal behaviors, the moderation layers, lives in the wrapper around the model rather than in the model itself. [SQ Magazine’s guide to reading system cards](https://sqmagazine.co.uk/glossary/system-card/) explains that a card deserves the name only when it describes the wrapper. OpenAI’s GPT-5 system card passes by naming its fast model (the cheap default that answers most queries), its reasoning model, and the real-time router that decides which one answers you.

Three structural weaknesses matter legally.

- The evidence is self-generated, since the labs grade their own homework.
- No shared template exists, so two cards from different labs are hard to compare, and regulators haven’t imposed one.
- Publication is voluntary everywhere; no jurisdiction requires a *public* system card. The closest analogue is the EU AI Act, which requires general-purpose AI providers to keep technical documentation on hand for regulators, not for the public.

Those weaknesses define the interpretive problem a court inherits. There’s no standard to measure a card against, no third-party certification behind the claims, and no legal duty that forced its creation. A company wrote the document because it wanted to, and everyone is now reading it as if it were a regulatory filing based on an established format.

## One Document, Two Verdicts

*Raine v. OpenAI*, case number CGC-25-628528 in San Francisco Superior Court, involves the death of a 16-year-old California teenager who used ChatGPT for roughly eight months before his death in April 2025. The facts of the underlying case are upsetting. They’re not, however, what the system-card question turns on.

The plaintiffs’ theory, laid out in the [original complaint and the October 2025 amended complaint](https://sf.courts.ca.gov/online-services/case-information), runs in a straight line. GPT-4o was evaluated primarily with single-prompt tests, while internal multi-turn testing caught the same instructions 73.5 percent of the time, against a publicly claimed 100 percent. The Model Spec, OpenAI’s published rulebook for model behavior, contained directives that pulled in opposite directions on exactly this kind of content. Then came the removals, with a categorical self-harm refusal protocol allegedly removed on May 8, 2024, and suicide and self-harm allegedly deleted from the “disallowed content” category on February 12, 2025. Each date maps to a period before or during the teen’s use of the product.

OpenAI reads the same corpus as exoneration. Its November 2025 answer describes structured safety testing, crisis-resource referrals in the conversation history, and usage policies prohibiting exactly the conduct at issue. OpenAI’s misuse defense treats the published policies as a fence, with the user on the wrong side of it: the company documented the rules, and in its telling, the teenager crossed them anyway.

> “Our safeguards work more reliably in common, short exchanges. We have learned over time that these safeguards can sometimes be less reliable in long interactions.”

That admission comes from an [OpenAI blog post](https://openai.com/index/helping-people-when-they-need-it-most/) published August 26, 2025, the same day the first complaint was filed, and it’s quoted in the amended complaint. It’s a public acknowledgment that testing conditions and deployment conditions diverge, which is precisely the gap the litigation is about.

## A Wave, Not One Case

*Raine* isn’t the only complaint leaning on this material. The JCCP 5431 proceeding swept up a cluster of related cases, and two of them lean on the documents harder than Raine’s lawyers did. In *Knowlton*, the complaint alleges that the GPT-5 system card itself disclosed that GPT-4o had been evaluated with single-prompt testing, and contrasts that with the 73.5 percent multi-turn figure discussed above. *Gray* attacks GPT-4o’s evaluation methodology through the GPT-5 card as well. Seven additional cases filed in November 2025 across California courts share the compressed-testing theory, according to [Nolo’s litigation tracker](https://www.nolo.com/legal-encyclopedia/can-ai-companies-be-held-liable-for-user-suicide.html).

The evidentiary posture is worth pausing on: a later product’s safety disclosure is being used as retrospective testimony about an earlier product’s testing. The company’s newest transparency document becomes a witness against its oldest shipping decisions.

That multiple plaintiffs are making this argument is a real signal about where the litigation is heading. It is not, however, a finding.

## What the Court Has Done

| Date | What happened |
| --- | --- |
| Aug. 26, 2025 | Complaint filed in San Francisco Superior Court, jury demanded |
| Oct. 2, 2025 | Defendants’ complex litigation designation denied |
| Oct. 22, 2025 | First amended complaint filed |
| Nov. 26, 2025 | Answer filed, exhibits lodged under seal |
| Feb. 2026 | Case stayed; coordination order entered |
| Mar. 2, 2026 | Coordination judge enters stay; JCCP 5431 assigned to Hon. Ethan P. Schulman |
| Sept. 23, 2026 | Discovery in the coordinated proceeding begins, per the case management order |

Source: San Francisco Superior Court docket, CGC-25-628528

Everything else on the docket since March is continuances and lawyers renewing their permission to appear. The two discovery motions that did get briefed were OpenAI demanding document responses from the plaintiffs, not the reverse. No motion targeting OpenAI’s internal testing, red-team, or launch-review records existed on the public docket as of early October 2026.

### What’s Sealed, and What Isn’t

OpenAI’s answer was filed with exhibits lodged under seal covering the individual user’s chat history and factual material supporting the causation defenses. The GPT-4o and GPT-5 system cards are fully public. What’s hidden is evidence both sides would use to argue causation, not the documents central to this article.

## Three Questions a Judge Will Eventually Answer

When the stay lifts and the coordinated judge reaches the merits, the system-card fight turns on three unresolved questions.

### Which layer does the card describe, and where did the harm live?

A card scoped to model weights says nothing about the router, the moderation stack, or the deployment context that shaped the user experience. If the testing described the weights and the failure lived in the wrapper, both sides get to argue about whether the document covers the conduct at issue.

### What’s a fair reading of the test data?

Single-prompt results generated offline, sitting next to multi-turn internal figures that tell a different story. The 73.5 percent versus 100 percent spread is the opening salvo, and discovery into internal evaluations will decide whether that spread was disclosed carefully, buried, or misrepresented.

### Does a stale card create its own liability?

If a safety feature was disclosed publicly and then removed with no corresponding disclosure, the document stops describing the system. The removal timeline alleged in the amended complaint sits directly on top of that question, and California’s [Transparency in Frontier Artificial Intelligence Act](https://www.cdt.ca.gov/initiatives/transparency-in-frontier-ai-act/), effective January 1, 2026, now attaches civil penalties up to $1 million per violation to frontier AI disclosure failures.

## The Regulatory Environment

While the litigation is stayed, the same principle, public claims matching actual behavior, is moving through regulators.

| Authority | Action | Exposure or deadline |
| --- | --- | --- |
| Federal Trade Commission (FTC) | Special reports ordered from companion chatbot providers under 6(b) authority | Ongoing inquiry opened Sept. 2025 |
| Securities and Exchange Commission (SEC) Division of Examinations | AI-disclosure alignment flagged in exam priorities | Nov. 2025 priorities |
| California | Companion chatbot statutes, including one named for the decedent in *Raine* | Crisis protocols and disclosure duties, effective July 2027 |
| New York | Ideation-detection and referral requirements | Fines up to $15,000/day for noncompliance |

State legislatures are writing laws about what the plaintiffs are arguing, and that convergence is bad news for defendants.

## What’s Next

The cards in dispute were written to demonstrate transparency, and they’re now operating as evidence, intended or not. Both sides hold the same documents up to the light and see opposite conclusions, and no judge has told either of them how the judiciary interprets these voluntary cards.

Discovery in the coordinated proceeding began in late September 2026. The internal test records are coming, and the gap between what those records show and what the cards claimed is where this case gets decided. If your company publishes safety documentation, assume a future plaintiff has already bookmarked it, and assume the version you wrote three models ago is still in circulation too.

No judge has ruled on this matter yet. When one is, the free ride is over.

---

**If you or someone you know is in crisis:** You can call or text 988 in the United States to reach the Suicide and Crisis Lifeline anytime, or contact your local emergency services.
