How evidence is produced
Red teaming, model and agent evaluation, benchmark design, and the practical limits of a test result.
Notes and research on AI evaluation, red teaming, governance, and continuous assurance. Written for people responsible for what happens after a finding.
Red teaming, model and agent evaluation, benchmark design, and the practical limits of a test result.
Ownership, controls, review gates, escalation, and the decisions that determine whether a finding matters.
Regression testing, change review, continuous evaluation, and the institutional work of keeping evidence current.
Practical writing grounded in technical evaluation evidence and the decisions it should inform.
It’s getting harder to know not just which AI you can trust, but whether you can trust AI at all.
Read the entry →