Who we work with

AI evaluation for different kinds of organizations

The same evaluation can serve very different decisions. We adapt the method and deliverables to the client’s authority and practical constraints. We state the limits and fit the work into the client’s review process.

Ways to work with Syntony

Syntony can act as an independent evaluator, collaborate on research, advise a governance team, or build evaluation tooling. We state the role, conflicts, and publication or review terms before work begins. We do not broker access or represent clients in external business development.

Useful work in each setting

A nonprofit may need a public benchmark; a research group, a reproducible study. A public agency may be preparing an acquisition or oversight review. A product team may need a release decision, while an AI lab may need an evaluation that remains valid across model versions.

In every case, the people responsible for the decision should be able to inspect the method, understand the result, and use the output.

Client groups

Client groups and common work

These descriptions show common starting points. We also work on projects that cross sectors or combine technical and policy teams.

01

Nonprofits & civil society

Evaluation and research for organizations working on public-interest questions, including rights, welfare, institutional practice, and the effects of AI on communities.

Common work
  • Benchmark design and validation
  • Independent research and evidence reviews
  • Evidence for client-led advocacy and policy work
  • Organizational AI policies
  • Datasets, prototypes, and public tools
Useful outputs

A documented method and public evidence that nontechnical readers can examine. We state the assumptions, funding, review rights, and limits.

02

Academia & research institutes

We collaborate with research teams on rigorous studies and build the tooling needed to run them reproducibly.

Common work
  • Research and experimental design
  • Benchmark and evaluator validation
  • Replication and robustness testing
  • Dataset and tooling development
  • Analysis and publication support
Useful outputs

Transparent methods, versioned code and data where appropriate, traceable results, uncertainty analysis, documentation, and publication-ready material.

03

Defense & government

Evaluation and governance for public institutions that buy, test, authorize, oversee, or deploy AI-enabled systems.

Common work
  • TEVV and mission-oriented evaluation
  • AI acquisition and vendor evidence
  • Red teaming and agent security
  • Governance and authorization readiness
  • Strategic risk and decision support
Useful outputs

Test protocols, replayable evidence, control and ownership maps, acceptance criteria, review inputs, regression tests, and decision briefs.

Explore the defense & government practice →
04

Enterprise & product teams

Evaluation and implementation support for teams building, buying, or already operating AI in products and internal workflows.

Common work
  • Pre-deployment evaluation
  • Agent and workflow red teaming
  • Vendor and model comparison
  • Controls and review processes
  • Continuous evaluation infrastructure
Useful outputs

Evidence tied to a product or operating decision, practical mitigations, release or acceptance criteria, accountable owners, and tests that can run again.

05

AI labs & model developers

Independent or collaborative evaluation work for teams developing models, agents, alignment interventions, safeguards, and evaluation programs.

Common work
  • Model and agent evaluation
  • Benchmark design and adversarial validation
  • Red teaming and safeguards testing
  • Training-intervention evaluation
  • Evaluation harnesses and data pipelines
Useful outputs

Robust methods, held-out tests, calibrated scoring, replayable traces, model and evaluator sensitivity analysis, and explicit claims and non-claims.

Ways of working

Roles Syntony can take

  1. 01Independent evaluator

    Test an existing system, benchmark, claim, or program against agreed criteria.

  2. 02Research partner

    Co-design and execute a study while making authorship, review, and publication terms explicit.

  3. 03Governance advisor

    Help the client use findings and policy commitments to build controls and decision processes people can use.

  4. 04Evaluation engineer

    Build the harness, pipeline, prototype, or internal tool needed to run the work repeatedly.

Project terms

Working terms by institution type

Public-interest and academic projects may require open methods, publication independence, accessible outputs, or funder disclosure. Commercial and public-sector work often requires confidentiality, security controls, information separation, or restricted disclosure.

We agree those terms in writing, along with the scope of any claims.

Plan the work

Tell us what your organization needs to test.

Include who will use the result and what decision it must support.