AnnouncementTurnkeeper Ward brings scattered safety signals into cases people can review.

Turnkeeper Labs · Paper 002

Two Signals, One Source

When do several warnings really come from just one source?

WardBench tests a simple but consequential mistake: treating several related safety signals as if they were independent support. Paper 002 explains what the benchmark checks now and what still needs human study.

WardBench v1.5 · 12 authored cases · No field validation

Why more signals are not always more evidenceWardBench v1.5
SourceA
01Direct
02Derived
03Duplicate
Visible signals
3
Independent sources
1
9/12cases misread when signals are counted separately
12/12authored-ground-truth matches when sources are tracked
Three visible signals can still contain only one independent source. This is an authored synthetic suite, not a field result.
Current evidenceWardBench v1.5 · 12 authored cases
What it showsCounting signals can double-count one source
Not claimedNo field validation or human findings
Next stepGated reviewer-study design

The problem

Count independent support, not the number of boxes on screen.

A system can show three warning signs even when all three trace back to the same observation, model, dataset, or report.

A system can display three signals even when all three inherit the same mistake. Paper 002 isolates that gap: the difference between the number of observations a reviewer can see and the number of independent sources those observations actually represent.

  1. 01Safety signal

    An observation, not a verdict.

  2. 02Source relationship

    Its origin and changes stay visible.

  3. 03Reviewable case

    Related signals can be examined together.

  4. 04Human decision

    No automated action follows from the benchmark.

Apparent agreement must not become stronger evidence unless the supporting sources are independent enough to justify the added weight.

What WardBench checks

Can the benchmark catch double-counting?

The implemented benchmark asks whether two transparent reference comparators behave as specified on a frozen synthetic suite.

  1. Freeze an inspectable synthetic suite

    WardBench v1.5 contains 12 metadata-only cases with declared sources, parent relationships, counterevidence, revocations, and authored ground truth.

  2. Run two reference comparators

    A naive comparator counts every visible supporting signal. A lineage-aware comparator resolves each active signal to its surviving root source.

  3. Test explicit failure conditions

    The suite checks false corroboration, preservation of counterevidence, threshold behavior, and recovery when a parent source is revoked.

  4. Report verification without generalizing

    The scorecard establishes that the fixtures and comparators behave as specified. It does not estimate reviewer behavior or production accuracy.

Current synthetic result

In this authored suite, source tracking avoids double-counting.

These descriptive figures come from WardBench v1.5. They verify the benchmark implementation and expose its baseline failure mode.

WardBench v1.5 results

12 frozen cases

Synthetic metadata cases covering independence, derivation, duplication, counterevidence, and revocation.

9/12 false corroboration

The naive signal-count comparator overstates the number of independent sources in nine authored cases.

0% false corroboration

The lineage-aware reference comparator does not exceed the authored independent-source count in this suite.

12/12 specification matches

This is consistency with co-authored fixtures and ground truth, not evidence of external or human validity.

Takeaway

Several warnings are not necessarily several independent sources. In these authored test cases, tracking where each signal came from prevented false corroboration.

This validates the benchmark's expected behavior, not real-world performance or a production decision method.

Next research phase · Gated

What WardBench cannot answer on its own.

The synthetic suite shows whether the benchmark behaves as specified. It cannot show how people read evidence, make judgments, or use a provenance display in practice.

What remains untested

No reviewer study has started.

WardBench cannot tell us how people interpret the same evidence. The planned human study remains Gated and has not recruited participants.

  1. Hold evidence content constant

    Matched case variants preserve the same observations while changing whether their shared lineage is hidden, listed, or shown as a graph.

  2. Randomize the presentation condition

    Qualified adult reviewers would receive one lineage representation without being told which interface is expected to perform better.

  3. Keep judgments decomposed

    Confidence, escalation, perceived independence, correction recovery, time, and inspection steps remain separate outcomes.

  4. Freeze the analysis before recruitment

    Exclusions, contrasts, decision thresholds, sample-size justification, and stopping rules must be preregistered before participant data are collected.

How a future study would work

Four factors keep signal count separate from source count.

Signal count and source count are manipulated separately so multiplicity cannot stand in for independence.

Signal count

One, two, or four visible signals while the number of independent sources is controlled separately.

Dependence

Genuinely independent observations, transformations of one observation, or exact and near duplicates.

Lineage display

No source relationship, a flat source list, or an explicit derivation graph with confidence and timestamps.

Source quality

Matched high-reliability, mixed-reliability, and unknown-reliability sources without averaging them together.

What a future study would measure

Six outcomes, kept separate.

No composite safety score will hide which part of the decision changed.

  1. False corroboration rate

    How often multiple dependent signals are treated as support from multiple independent sources.

  2. Confidence inflation

    The confidence added by duplicate or derived signals beyond the matched single-source condition.

  3. Escalation flips

    Whether apparent corroboration moves a case across a preregistered review threshold.

  4. Lineage sensitivity

    How much a decision changes when the same evidence is shown with its source relationships made explicit.

  5. Correction recovery

    Whether removing or revoking a parent signal removes every conclusion that depended on it.

  6. Review cost

    Time, inspection steps, and uncertainty introduced by each provenance representation.

What we would try to disprove

What the study is designed to falsify.

These must be finalized with the analysis plan before preregistration. They are not results and have not been tested with reviewers.

  1. H1 · Multiplicity without independence inflates support

    When lineage is hidden, dependent signals will produce more threshold crossings than the matched single-source condition.

  2. H2 · Explicit lineage reduces false corroboration

    A provenance graph will reduce double-counting more than a flat source list while preserving genuinely independent support.

  3. H3 · Mixed source quality magnifies the error

    Dependence mistakes will be largest when a low-quality parent produces several polished downstream signals.

  4. H4 · Provenance-aware decisions recover more completely

    Revoking a parent signal will reverse more dependent conclusions when lineage survives aggregation.

What evidence fusion must preserve

What evidence fusion must preserve.

Each invariant is a property the benchmark will try to break.

  1. Signals are not conclusions

    Multiple observations may guide investigation, but their count does not establish a finding.

  2. Lineage survives aggregation

    Every derived claim remains traceable to the source observations that support it.

  3. Dependence is not hidden confidence

    A transformed or duplicated signal cannot silently add the weight of a new independent source.

  4. Correction travels downstream

    A correction, expiry, or revocation reaches every conclusion that inherited the original signal.

  5. Contradictions remain visible

    Counterevidence is preserved beside supporting evidence instead of being flattened into one score.

  6. No automatic enforcement

    No benchmark output or reviewer response grants authority to act against a person or account.

Built now, not validated beyond the suite

Separate what is built from what remains untested.

The benchmark implementation exists. External validation and the reviewer experiment do not.

Defined now

  • Research question and claim boundary

  • Frozen WardBench synthetic suite (WB01–WB12, v1.5 assumption-driven)

  • Versioned naive and lineage-aware reference comparators

  • Descriptive implementation scorecard and focused regression tests

  • Candidate reviewer-study hypotheses, factors, and outcome ledger

  • Verified public literature basis

Not started

  • Externally validated or held-out benchmark cases

  • Preregistered reviewer-study analysis and decision thresholds

  • Held-out confirmatory run

  • Human-participant recruitment or data collection

  • Production-system or real-world safety claims

What this cannot prove

What the synthetic result does not buy us.

The current scorecard can catch implementation regressions. It cannot establish how people or production systems behave.

  • The cases and their ground truth were authored together; a perfect match is an implementation consistency check, not independent validation.

  • The lineage-aware comparator receives explicit parent links. It does not infer uncertain relationships from noisy observations.

  • The three-signal review threshold is a synthetic stress device, not a calibrated safety threshold.

  • WardBench does not currently test probabilistic or model-assisted reconstruction against the reference comparators.

  • Lineage-display labels are fixture metadata until a reviewer experiment actually presents matched interfaces to people.

  • No practitioner panel, design partner, or production distribution has validated the representativeness of WB01–WB12.

  • That dependent signals cause real-world reviewers to make a specific decision

  • That a provenance graph is the best interface across every workflow

  • That deterministic, probabilistic, or hybrid reconstruction is most reliable in production

  • The prevalence of duplicate or derived safety signals in production systems

  • The correctness of any live allegation, identity link, or enforcement decision

  • Causal effects on platform safety, people, or protected groups

Before involving reviewers

No recruitment before these conditions are met.

Human-participant work is a separate phase, not an automatic continuation of the synthetic benchmark.

  • A frozen synthetic benchmark establishes that the manipulation works as intended.

  • A documented ethics, privacy, consent, and data-retention review approves the protocol.

  • Recruitment avoids live accusations, active cases, customer content, and identifiable minors.

  • The study can stop without affecting a participant's work, access, or employment.

Why this matters

The research exists. The system connecting it does not.

Prior work explains source dependence, correlated evidence, adversarial evaluation, and provenance. We are studying what happens when those ideas collide inside safety-review workflows.

EvaluationAART · arXiv 2023

Structured adversarial evaluation (inspiration)

AART shows how to generate systematic adversarial examples for evaluation. WardBench applies that discipline to synthetic *cases*—dependence, trajectory, and counterevidence—not jailbreak prompts.

Turnkeeper relevanceSynthetic cases can expose predictable failure modes before a product makes broader claims.

Why it matters →
Evidence scienceCogSci 2024

Source dependence in human evidence judgments

Strittmatter, Pilditch, and Lagnado report that people can perceive source dependence yet often fail to discount it fully.

Turnkeeper relevanceIndependent-looking signals can still trace back to one source, and people may not discount that relationship fully.

Why it matters →
Evidence scienceeLife 2025

Correlated evidence in perceptual decisions

Tardiff, Kang, and Gold study how people weight and accumulate observations under experimentally controlled correlation.

Turnkeeper relevanceSignal count cannot stand in for independent support when observations are correlated.

Why it matters →
StandardOASIS STIX 2.1

Machine-readable relationship vocabulary

STIX 2.1 defines relationships such as derived-from and duplicate-of that can carry lineage without deciding how a reviewer should weight it.

Turnkeeper relevanceA shared relationship vocabulary can preserve derivation without deciding how a reviewer must act.

Why it matters →
Safety practiceTechnology Coalition · Lantern

Shared signals require local investigation

The Lantern program describes exchanged signals as clues rather than proof and leaves investigation to the receiving company.

Turnkeeper relevanceShared signals should remain clues for local investigation, not automatic conclusions.

Why it matters →

Help make Paper 002 harder to fool.

Challenge the synthetic cases, outcome measures, research design, or fit for safety-review workflows before any human-participant study.