12 frozen cases
Synthetic metadata cases covering independence, derivation, duplication, counterevidence, and revocation.
Turnkeeper Labs · Paper 002
When do several warnings really come from just one source?
WardBench tests a simple but consequential mistake: treating several related safety signals as if they were independent support. Paper 002 explains what the benchmark checks now and what still needs human study.
WardBench v1.5 · 12 authored cases · No field validation
The problem
A system can show three warning signs even when all three trace back to the same observation, model, dataset, or report.
A system can display three signals even when all three inherit the same mistake. Paper 002 isolates that gap: the difference between the number of observations a reviewer can see and the number of independent sources those observations actually represent.
An observation, not a verdict.
Its origin and changes stay visible.
Related signals can be examined together.
No automated action follows from the benchmark.
Apparent agreement must not become stronger evidence unless the supporting sources are independent enough to justify the added weight.
What WardBench checks
The implemented benchmark asks whether two transparent reference comparators behave as specified on a frozen synthetic suite.
WardBench v1.5 contains 12 metadata-only cases with declared sources, parent relationships, counterevidence, revocations, and authored ground truth.
A naive comparator counts every visible supporting signal. A lineage-aware comparator resolves each active signal to its surviving root source.
The suite checks false corroboration, preservation of counterevidence, threshold behavior, and recovery when a parent source is revoked.
The scorecard establishes that the fixtures and comparators behave as specified. It does not estimate reviewer behavior or production accuracy.
Current synthetic result
These descriptive figures come from WardBench v1.5. They verify the benchmark implementation and expose its baseline failure mode.
WardBench v1.5 results
Synthetic metadata cases covering independence, derivation, duplication, counterevidence, and revocation.
The naive signal-count comparator overstates the number of independent sources in nine authored cases.
The lineage-aware reference comparator does not exceed the authored independent-source count in this suite.
This is consistency with co-authored fixtures and ground truth, not evidence of external or human validity.
Takeaway
Several warnings are not necessarily several independent sources. In these authored test cases, tracking where each signal came from prevented false corroboration.
This validates the benchmark's expected behavior, not real-world performance or a production decision method.
Next research phase · Gated
The synthetic suite shows whether the benchmark behaves as specified. It cannot show how people read evidence, make judgments, or use a provenance display in practice.
What remains untested
WardBench cannot tell us how people interpret the same evidence. The planned human study remains Gated and has not recruited participants.
Matched case variants preserve the same observations while changing whether their shared lineage is hidden, listed, or shown as a graph.
Qualified adult reviewers would receive one lineage representation without being told which interface is expected to perform better.
Confidence, escalation, perceived independence, correction recovery, time, and inspection steps remain separate outcomes.
Exclusions, contrasts, decision thresholds, sample-size justification, and stopping rules must be preregistered before participant data are collected.
How a future study would work
Signal count and source count are manipulated separately so multiplicity cannot stand in for independence.
One, two, or four visible signals while the number of independent sources is controlled separately.
Genuinely independent observations, transformations of one observation, or exact and near duplicates.
No source relationship, a flat source list, or an explicit derivation graph with confidence and timestamps.
Matched high-reliability, mixed-reliability, and unknown-reliability sources without averaging them together.
What a future study would measure
No composite safety score will hide which part of the decision changed.
How often multiple dependent signals are treated as support from multiple independent sources.
The confidence added by duplicate or derived signals beyond the matched single-source condition.
Whether apparent corroboration moves a case across a preregistered review threshold.
How much a decision changes when the same evidence is shown with its source relationships made explicit.
Whether removing or revoking a parent signal removes every conclusion that depended on it.
Time, inspection steps, and uncertainty introduced by each provenance representation.
What we would try to disprove
These must be finalized with the analysis plan before preregistration. They are not results and have not been tested with reviewers.
When lineage is hidden, dependent signals will produce more threshold crossings than the matched single-source condition.
A provenance graph will reduce double-counting more than a flat source list while preserving genuinely independent support.
Dependence mistakes will be largest when a low-quality parent produces several polished downstream signals.
Revoking a parent signal will reverse more dependent conclusions when lineage survives aggregation.
What evidence fusion must preserve
Each invariant is a property the benchmark will try to break.
Multiple observations may guide investigation, but their count does not establish a finding.
Every derived claim remains traceable to the source observations that support it.
A transformed or duplicated signal cannot silently add the weight of a new independent source.
A correction, expiry, or revocation reaches every conclusion that inherited the original signal.
Counterevidence is preserved beside supporting evidence instead of being flattened into one score.
No benchmark output or reviewer response grants authority to act against a person or account.
Built now, not validated beyond the suite
The benchmark implementation exists. External validation and the reviewer experiment do not.
Research question and claim boundary
Frozen WardBench synthetic suite (WB01–WB12, v1.5 assumption-driven)
Versioned naive and lineage-aware reference comparators
Descriptive implementation scorecard and focused regression tests
Candidate reviewer-study hypotheses, factors, and outcome ledger
Verified public literature basis
Externally validated or held-out benchmark cases
Preregistered reviewer-study analysis and decision thresholds
Held-out confirmatory run
Human-participant recruitment or data collection
Production-system or real-world safety claims
What this cannot prove
The current scorecard can catch implementation regressions. It cannot establish how people or production systems behave.
The cases and their ground truth were authored together; a perfect match is an implementation consistency check, not independent validation.
The lineage-aware comparator receives explicit parent links. It does not infer uncertain relationships from noisy observations.
The three-signal review threshold is a synthetic stress device, not a calibrated safety threshold.
WardBench does not currently test probabilistic or model-assisted reconstruction against the reference comparators.
Lineage-display labels are fixture metadata until a reviewer experiment actually presents matched interfaces to people.
No practitioner panel, design partner, or production distribution has validated the representativeness of WB01–WB12.
That dependent signals cause real-world reviewers to make a specific decision
That a provenance graph is the best interface across every workflow
That deterministic, probabilistic, or hybrid reconstruction is most reliable in production
The prevalence of duplicate or derived safety signals in production systems
The correctness of any live allegation, identity link, or enforcement decision
Causal effects on platform safety, people, or protected groups
Before involving reviewers
Human-participant work is a separate phase, not an automatic continuation of the synthetic benchmark.
A frozen synthetic benchmark establishes that the manipulation works as intended.
A documented ethics, privacy, consent, and data-retention review approves the protocol.
Recruitment avoids live accusations, active cases, customer content, and identifiable minors.
The study can stop without affecting a participant's work, access, or employment.
Why this matters
Prior work explains source dependence, correlated evidence, adversarial evaluation, and provenance. We are studying what happens when those ideas collide inside safety-review workflows.
AART shows how to generate systematic adversarial examples for evaluation. WardBench applies that discipline to synthetic *cases*—dependence, trajectory, and counterevidence—not jailbreak prompts.
Turnkeeper relevanceSynthetic cases can expose predictable failure modes before a product makes broader claims.
Why it matters →Strittmatter, Pilditch, and Lagnado report that people can perceive source dependence yet often fail to discount it fully.
Turnkeeper relevanceIndependent-looking signals can still trace back to one source, and people may not discount that relationship fully.
Why it matters →Tardiff, Kang, and Gold study how people weight and accumulate observations under experimentally controlled correlation.
Turnkeeper relevanceSignal count cannot stand in for independent support when observations are correlated.
Why it matters →STIX 2.1 defines relationships such as derived-from and duplicate-of that can carry lineage without deciding how a reviewer should weight it.
Turnkeeper relevanceA shared relationship vocabulary can preserve derivation without deciding how a reviewer must act.
Why it matters →The Lantern program describes exchanged signals as clues rather than proof and leaves investigation to the receiving company.
Turnkeeper relevanceShared signals should remain clues for local investigation, not automatic conclusions.
Why it matters →Challenge the synthetic cases, outcome measures, research design, or fit for safety-review workflows before any human-participant study.