Evaluation teams
Teams running adversarial evaluations that need findings to feed control and review processes.
GovTune is being developed for the handoff after an AI evaluation. It reads a red-team trace and drafts the governance record a reviewer needs without breaking the link to the source evidence. We are opening a small number of design-partner pilots with teams that already run AI evaluations or govern high-stakes AI systems.
Request pilot accessEvaluation tools and governance systems usually store different records. A trace shows what the system did and how the test was run. A reviewer needs a concise account of the finding, its severity, the proposed response, and the decision that followed.
GovTune drafts that account from the evaluation record and keeps a reference to the source trace.
GovTune remains in development. The design-partner program is for teams that can test it against real workflows and established review standards.
A design-partner pilot starts with a real evaluation workflow. We test GovTune against the team's existing process and terminology.
GovTune supports evaluators and reviewers; it does not replace either. Every draft remains tied to the technical evidence and the decision record.
Provide a red-team trace, evaluation run, incident record, or structured finding from an existing workflow.
GovTune drafts the finding in a format suitable for a review board, risk committee, or control owner.
It maps the evidence to candidate controls, owners, verification steps, and escalation triggers.
It preserves an audit trail from the evidence to the recommendation and decision for later review.
Teams running adversarial evaluations that need findings to feed control and review processes.
People responsible for reviewing AI risk evidence, assigning responsibility, and setting follow-up actions.
Leads responsible for model risk, safety, assurance, or escalation in a high-stakes deployment.
Teams able to test GovTune in a real workflow and report clearly where it succeeds or fails.
A pilot succeeds only if a reviewer can trace every drafted claim back to the source and use the output in the team's existing process. We also test whether the draft saves time without obscuring uncertainty or shifting accountability.
The deliverable is a documented pilot result: where GovTune fit, where it failed, and what would have to change before broader use.
Use this form if your team already runs evaluations or reviews AI risk and can test GovTune against a real process.
Do not submit classified information, CUI, export-controlled data, credentials, client evidence, or other sensitive system details through this public form. We will establish an appropriate channel before any evidence transfer.