Repeated evidence
One source can travel through several systems and return looking like consensus.
WardBench v1.5 · For trust and safety teams
WardBench tests whether safety tools can tell the difference between three independent warnings and one warning repeated three times.
In these controlled cases, Turnkeeper Ward keeps the source and history of each warning together so a reviewer can see one clear case. Your team still makes the final decision.
These controlled cases test whether Turnkeeper keeps warnings and their sources connected. They use no live customer data, do not measure results with real users, and never let Turnkeeper decide what action your team should take.
Results from WardBench v1.5
Tracking where each warning came from reached the expected result in all 12 cases.
These results come from controlled test cases. They do not show how the system performs with real users.
Why it matters
A safety reviewer needs to know whether several warnings confirm one another—or simply repeat the same underlying information.
One source can travel through several systems and return looking like consensus.
Conflicting evidence, changed timelines, and revoked signals can alter the case.
Reviewers need an explainable case, not a larger pile of disconnected alerts.
What WardBench tests
WardBench examines whether a system can preserve where evidence came from, how it changed, and what might point toward a different conclusion.
We are inviting safety practitioners to test whether these synthetic cases reflect the difficult situations reviewers encounter in practice.