AnnouncementTurnkeeper Ward brings scattered safety signals into cases people can review.

WardBench v1.5 · For trust and safety teams

Three warnings do not always mean three sources.

WardBench tests whether safety tools can tell the difference between three independent warnings and one warning repeated three times.

In these controlled cases, Turnkeeper Ward keeps the source and history of each warning together so a reviewer can see one clear case. Your team still makes the final decision.

12 controlled test cases4 with stated assumptionsNo live user data
Method and limits

These controlled cases test whether Turnkeeper keeps warnings and their sources connected. They use no live customer data, do not measure results with real users, and never let Turnkeeper decide what action your team should take.

Results from WardBench v1.5

What happened in the 12 test casesv1.5 · controlled tests

Tracking where each warning came from reached the expected result in all 12 cases.

Cases wrong when every warning is counted separately
9/12
Cases correct when each warning's source is tracked
12/12
Matched the expected result
12/12
Cases with stated assumptions
4/12

These results come from controlled test cases. They do not show how the system performs with real users.

Why it matters

Signal volume can create false confidence.

A safety reviewer needs to know whether several warnings confirm one another—or simply repeat the same underlying information.

01

Repeated evidence

One source can travel through several systems and return looking like consensus.

02

Missing context

Conflicting evidence, changed timelines, and revoked signals can alter the case.

03

Human decisions

Reviewers need an explainable case, not a larger pile of disconnected alerts.

What WardBench tests

A focused test for safer case reconstruction.

WardBench examines whether a system can preserve where evidence came from, how it changed, and what might point toward a different conclusion.

Available in this benchmarkv1.5 · implemented synthetic
  • Twelve synthetic cases (WB01–WB12), including four assumption-tagged v1.5 additions
  • A direct comparison between naive counting and lineage-aware reconstruction
Not claimedGated / roadmap
  • Practitioner-validated case set (next after sharp conversations)
  • Human-participant reviewer study (Labs research gate)
  • Model Lab persistence, live traffic, or prevalence claims

Help us challenge the benchmark.

We are inviting safety practitioners to test whether these synthetic cases reflect the difficult situations reviewers encounter in practice.