Methods and limits
How red teams design tests and document failures without claiming more than a result supports.
We write about testing AI systems and responding to the results. The focus is what a test can actually prove and what the organization does once the report is delivered.
How red teams design tests and document failures without claiming more than a result supports.
Why technically sound work sometimes stalls after review, and what makes a finding usable to the people who can act.
How teams decide when a system needs another evaluation and compare the new result with the earlier one.
Essays and field notes for people who test AI systems or make decisions about their use.
It’s getting harder to know not just which AI you can trust, but whether you can trust AI at all.
Read the article →