---
title: "Guidelight Publishes Another AI Safety Report. Here&#8217;s Where It Fits In."
description: "Another AI safety report dropped on August 18, 2026. Guidelight , a startup founded by two former OpenAI employees, graded five frontier companies on six control practices. Everyone failed. Anthropic and OpenAI tied for the top score at C-plus. Google"
url: https://kaynemcgladrey.com/blog/guidelight-publishes-another-ai-safety-report-heres-where-it-fits-in/
date: 2026-08-20
modified: 2026-08-20
author: "Kayne"
image: https://kaynemcgladrey.com/wp-content/uploads/2026/08/pexels-mart-production-7605981.webp
categories: ["Blog"]
type: post
lang: en-US
---

# Guidelight Publishes Another AI Safety Report. Here&#8217;s Where It Fits In.

Another AI safety report dropped on August 18, 2026. [Guidelight](https://guidelight.ai/blog/control-assessment-august-2026), a startup founded by two former OpenAI employees, graded five frontier companies on six control practices. Everyone failed. Anthropic and OpenAI tied for the top score at C-plus. Google was D-plus, xAI D-minus, and Meta got an F.

Before rolling your eyes at one more watchdog publication, look at what’s already on the shelf. Stanford’s [Foundation Model Transparency Index (FMTI)](https://crfm.stanford.edu/fmti/December-2025/index.html) has tracked company disclosure on 100 indicators for three years. The UK’s [AI Security Institute (AISI)](https://www.aisi.gov.uk/) ran hands-on evaluations of 30+ systems, and the [OECD’s AI Incidents and Hazards Monitor](https://oecd.ai/en/incidents) has logged 17,104 realized incidents to date. So why does another report deserve attention?

| Report | Publisher | Org Type | Primary Focus | Latest Edition |
| --- | --- | --- | --- | --- |
| FMTI | Stanford CRFM | Academic | Company disclosure (100 indicators, 13 firms) | Dec 2025 |
| AISI | UK Government (DSIT) | Government | Model capabilities and safeguards (30+ systems) | Dec 2025 |
| OECD AIM | OECD | Intergovernmental | Realized incidents (media-derived) | Live (Aug 2026) |
| Guidelight | Private (ex-OpenAI founders) | Private | Organizational controls (6 practices, 5 firms) | Aug 2026 |

### What Guidelight Measures That Others Don’t

The six practices Guidelight evaluates fill a gap nobody else has addressed: logging, monitor efficacy, gated actions, circuit breaking, third-party review, and containment planning.

- FMTI grades what companies say.
- AISI tests what models can do.
- OECD counts what breaks.
- Guidelight asks whether the organization itself has the plumbing to catch a misbehaving model before it becomes an OECD statistic.

No company scored above 3.0 on the 5.0 scale, where 2.0 marks limited partial implementation and 3.0 marks substantial partial. Four and five went unclaimed.

![](https://kaynemcgladrey.com/wp-content/uploads/2026/08/image-1.webp)

The specifics carry more weight than the letter grades. Google published the most detailed forward-looking AI Control Roadmap in July 2026, but hasn’t implemented most of it. That’s not a footnote; it’s the report’s central pattern: published intent outpaces operational reality at every company evaluated. That gap between intent and execution is because nothing in the market rewards closing it.

OpenAI earned best-in-class on containment because it paused workloads after the July 2026 Hugging Face breach, but no formal written plan exists. Reactive pauses aren’t institutionalized safety; they’re damage control.

Meta’s F traces partly to what it told METR during February-March 2026 evaluations. xAI’s D-minus reflects non-participation in METR’s Frontier Risk Report and the absence of any third-party red-teaming. The grades punish opacity even when the underlying cause might be security concerns rather than negligence. FMTI documented the same industry-wide retreat, with engagement falling from 74 percent in 2024 to 30 percent in 2025.

### Where the Reports Reinforce Each Other

Put all four side by side and the trend lines point the same direction:

- **Transparency is shrinking.** FMTI’s mean score dropped 17 points to 41, and companies are participating less.
- **Capabilities are accelerating.** AISI found cyber task capability doubling roughly every eight months, self-replication success climbing from 5 percent to 60 percent, and universal jailbreaks in every system tested.
- **Incidents are rising.** The OECD count is up 89.8 percent year over year, though the OECD noted it’s declining as a share of total AI news.
- **Controls remain partially implemented.** Guidelight’s top score was 2.50 out of 5.0, which makes “substantial partial” the 2026 ceiling for best practices rather than complete implementation.

[Reuters](https://www.reuters.com/technology/artificial-intelligence/ai-firms-cant-yet-contain-what-theyve-built-study-finds-2026-08-19/) coverage of the new report reminded everyone that OpenAI and Anthropic agents [escaped testing environments](https://kaynemcgladrey.com/blog/writing-your-2027-security-budget-after-ai-vendors-set-the-house-on-fire/) and probed other companies’ defenses. OpenAI chief scientist Jakub Pachocki was quoted wondering whether a capable model could “figure out that it should evade any monitoring on its own” or “disable the monitors.” That’s the head of science at the company tied for best controls acknowledging the tech might outpace containment.

### The Founder Question

Steven Adler came from OpenAI’s safety team, and Page Hedley handled policy and ethics there; both co-founded Guidelight. OpenAI tied for the top grade. The report shared preliminary scores with company staff and invited corrections, which is decent process. But the field is too small for this to be neutral territory (independent verification will matter more if there’s a second edition). Stanford maintains distance through an academic advisory board, and the UK AISI operates under a government directorate. Guidelight has two people and a methodology document. Their insider credibility cuts both ways: they know how these systems work from the inside, and they have the most intimate knowledge of the company they graded highest.

### What’s Still Not Measured

The control standard covers inference-time practices, not training-phase safeguards or third-party testing modes, and recent incidents occurred during exactly those uncovered phases. Guidelight’s limitations section admits this explicitly. That’s honest, even if it undercuts the completeness of the grades.

There’s another gap nobody is measuring: commercial pressure to ship. OpenAI announced a teen-targeted model the same week it slowed development to reexamine safety after the Hugging Face incident. None of the four reports asks whether deployment timelines reflect actual risk findings.

Guidelight doesn’t replace the incumbents, but it fills one specific gap: the organizational control assessment nobody else was making. Whether those six practices reduce realized harm is a question none of these reports can answer. At least now the coordinate is plotted.
