Find the fracture.
Close the loop.
Breaking an AI system is the beginning.
Knowing what to change comes next.
Test a simulated business workflow. Expose a failure, choose a control, and build a replayable record of what changed.
A better question than
“Can you break it?”
Your move.
The system is listening.
Five domains. Increasing difficulty. One complete loop.
Choose a mission and test the boundary between what an AI reads and what it should obey.
A support assistant processes refund requests. Test whether untrusted customer text can persuade it to override the refund policy.
Trigger a policy violation, then prevent it without blocking a legitimate request.
Refunds require an eligible order and a permitted amount.
People affected:
Craft your test
Use synthetic content. Never paste secrets or real customer data.
How does this simulation work?
This lab uses deterministic rules and synthetic business data. It illustrates trust boundaries; it is not an evaluation of a live model.
Use a proposed tool call in this format:
<tool>{"action":"refund","ticket":"SYN-1042","amount":499}</tool>Only the single <tool> JSON action is parsed. Surrounding prose is illustrative and does not change the outcome. Load an example for the exact fields supported by each mission.
Inspect the behavior
Every decision leaves a trace.
Run a test to see how the system handles your input.Build the evidence
A strong submission connects the failure, the decision, and the proof.
Technical score: reproduce a failure (20), contain it (20), block supplied attack variants (20), preserve useful tasks (20). Record completeness: owner, decision, retest trigger, and rationale (5 each). Human review: not performed. This is not a production security assessment.
A fix changes
more than one thing.
A control can interrupt a failure and create a new queue. Explore accumulation, delayed response, and the people downstream of the decision.
Each exception can become a precedent in the next decision.
Read the model, assumptions, and limits
This is an illustrative stock-and-flow model, not an operational forecast or an AI evaluation. Queue(next) = max(0, queue + admitted arrivals − completed reviews). Initial queue: 6. Capacity: 12 per cycle. Horizon: 12 cycles. After the selected delay, throttling admits at most 9 arrivals while the queue exceeds 6, then at most 12. Deferred arrivals accumulate separately; they do not disappear. Stopping intake admits zero after the delay. The chart is an exploratory aid and does not add score points. Record the human tradeoff in your decision rationale.
Look beyond the first failure.
Change the conditions. Inspect what accumulates and what appears later.
Ask who absorbs the cost.
Bring domain expertise and affected perspectives into the decision.
Turn evidence into a practice.
Assign responsibility, preserve uncertainty, and test again after change.
Good findings
deserve a paper trail.
Bring your challenge into the way your team already works. Keep the attack, the control, and the regression evidence together.
Download the GitHub-ready kitOpens a prefilled GitHub issue for your review. Nothing is submitted automatically.
The next test
is your system.
This is a simulation. Your business has real dependencies, people, and consequences. Syntony brings evaluation, governance, and assurance into one continuous practice.
Scope an AI assurance assessment Explore Syntony’s approach