01
Nonprofits & civil society
Evaluation and research for organizations working on public-interest questions, including rights, welfare, institutional practice, and the effects of AI on communities.
Common work- Benchmark design and validation
- Independent research and evidence reviews
- Evidence for client-led advocacy and policy work
- Organizational AI policies
- Datasets, prototypes, and public tools
Useful outputsA documented method and public evidence that nontechnical readers can examine. We state the assumptions, funding, review rights, and limits.
02
Academia & research institutes
We collaborate with research teams on rigorous studies and build the tooling needed to run them reproducibly.
Common work- Research and experimental design
- Benchmark and evaluator validation
- Replication and robustness testing
- Dataset and tooling development
- Analysis and publication support
Useful outputsTransparent methods, versioned code and data where appropriate, traceable results, uncertainty analysis, documentation, and publication-ready material.
03
Defense & government
Evaluation and governance for public institutions that buy, test, authorize, oversee, or deploy AI-enabled systems.
Common work- TEVV and mission-oriented evaluation
- AI acquisition and vendor evidence
- Red teaming and agent security
- Governance and authorization readiness
- Strategic risk and decision support
04
Enterprise & product teams
Evaluation and implementation support for teams building, buying, or already operating AI in products and internal workflows.
Common work- Pre-deployment evaluation
- Agent and workflow red teaming
- Vendor and model comparison
- Controls and review processes
- Continuous evaluation infrastructure
Useful outputsEvidence tied to a product or operating decision, practical mitigations, release or acceptance criteria, accountable owners, and tests that can run again.
05
AI labs & model developers
Independent or collaborative evaluation work for teams developing models, agents, alignment interventions, safeguards, and evaluation programs.
Common work- Model and agent evaluation
- Benchmark design and adversarial validation
- Red teaming and safeguards testing
- Training-intervention evaluation
- Evaluation harnesses and data pipelines
Useful outputsRobust methods, held-out tests, calibrated scoring, replayable traces, model and evaluator sensitivity analysis, and explicit claims and non-claims.