AnnouncementTurnkeeper Ward brings scattered safety signals into cases people can review.

Turnkeeper Labs · Paper 001

One Bad Link

What happens when one mistaken identity connection spreads?

A wrong identity link can make unrelated entities appear connected. This proposed synthetic study will test how far that mistake can travel through scores and review status—and whether correcting the link removes those effects.

Synthetic fixture only · Gated research · No run or findings

One link, three effective statesProposed study
Clean

Unrelated groups stay separate.

Mistaken link

One wrong connection changes effective state.

Corrected

Active outputs should return to the clean state.

The illustration shows effective state only. Any correction remains part of the immutable history.
Research statusGated · design not frozen
Built nowOne frozen synthetic contract pair
Not claimedNo study run or findings
Next stepIndependent methods review

The problem

Measure what the mistake changes—not only whether the link was wrong.

One incorrect connection can change the paths a system treats as related, even when the underlying entities are unrelated.

Conventional matching metrics tell us whether a link is correct. This study asks what that error changes after evidence begins moving through a system — and whether correction actually removes every dependent conclusion.

An uncertain identity association may guide investigation, but it must not silently become evidence that the linked person engaged in harmful behavior.

How we would test it

One edit. Paired counterfactuals. Inspectable policies.

  1. Generate a clean graph

    Create a seeded synthetic graph with known entities, observations, and provenance.

  2. Inject one false link

    Join an elevated-risk component to an unrelated benign component with one controlled edit.

  3. Compare downstream outputs

    Evaluate transparent propagation policies on identical clean and corrupted graph pairs.

  4. Correct the link and test recovery

    Compare full recompute, where active state should match the clean counterfactual, with deliberately incomplete link-record-only invalidation, where residual output divergence is measured; both preserve correction history.

What we would measure

Six outcomes, kept separate.

The study will keep the downstream effects separate instead of hiding them in one universal harm score.

  1. False-link blast radius

    Separate counts of how many entities, cases, scores, explanations, and review-status outputs change.

  2. Innocent-entity contamination

    How often a ground-truth benign entity is changed by the false link.

  3. Intervention flips

    Whether an unrelated entity crosses a preregistered human-review threshold.

  4. Transferred risk

    How much score moves to an unrelated entity across the injected link.

  5. Attribution concentration

    How much of an affected score depends on a path crossing the false link.

  6. Recovery completeness

    How much affected state returns to the clean counterfactual after correction.

What must stay true

What the proposed study is designed to pressure-test.

Each invariant is a property the study tries to break, not a property we claim as established.

  1. No silent transfer

    An uncertain association must not become an unmarked risk signal on a linked entity.

  2. Separate confidence

    Link confidence and risk conclusion stay distinct values, never collapsed into one score.

  3. Bounded influence

    A single link has a declared, limited ceiling on how far its effect can travel.

  4. Counterfactual traceability

    Any affected output can be traced to the paths that depend on the link in question.

  5. Complete reversibility

    Correcting a link invalidates every active conclusion that depended on it while preserving the assertion, correction, and historical outputs in the audit trail.

  6. No automatic enforcement

    Nothing in the study proposes autonomous action against any person or account.

What exists now

A tested contract kernel, not an executable study.

A frozen contract pair and truth-blind active-state projection are implemented and tested. There is no seeded generator, propagation benchmark, development run, or finding.

Implemented or documented now

  • Research question and synthetic-only scope boundary

  • Clean, corrupted, and corrected counterfactual design

  • One frozen, structure-matched false/true contract pair

  • Active-state policy input separated from oracle truth and audit history

  • Immutable assertion and correction graph-state history

  • Focused contract integrity and isolation tests

  • Targeted prior-work verification record

  • Draft factor matrix with position and true-link controls

  • Six proposed downstream-contamination outcomes

  • Design-freeze and no-enforcement gates

Not yet tested

  • Seeded synthetic graph generator and frozen factor-matrix fixtures

  • Propagation-policy implementations and metric computation

  • Full-recompute and link-record-only derived-output recovery comparison

  • Development-only analysis artifacts and bootstrap intervals

  • Held-out confirmatory hypotheses

  • Frozen peripheral, hub, and boundary endpoint-stratum definitions

  • Matched true-link positive-control / correct-link utility results

  • Real-world prevalence or base rates

  • Production-system performance

  • Demographic fairness or subgroup effects

What happens next

Phase 1 verdict: freeze pending.

The initial literature scan is recorded. The preregistration remains open and no study run may begin until these conditions are fixed.

  1. Operational factor definitions

    Mutually exclusive peripheral, hub, and boundary endpoint strata—plus graph size and link-confidence buckets—must be frozen before a run.

  2. Correct-link utility control

    A structure-matched true-link twin needs its own unmerged baseline, oracle, utility outcomes, and contrast so benefit is not inferred from the false-link cell.

  3. Policy and decision freeze

    Propagation policies, confidence buckets, human-review threshold, and recovery semantics need a fixed version before implementation or any run.

  4. Independent methods review

    The initial prior-work scan is recorded. A reproducible literature search and independent methods review must still resolve critical objections before freeze.

What this cannot prove

A synthetic study cannot answer real-world questions by itself.

  • The prevalence or magnitude of downstream harm in production systems

  • Which propagation policy is best across all systems and datasets

  • Causal effects on real-world people, moderators, or platform outcomes

  • Performance outside the declared synthetic graph families

  • Fairness across protected classes or sensitive attributes

  • A production prevention mechanism or real-world harm reduction

Progress so far

The work and the gates, in the open.

  1. Research direction selected

    Narrowed Paper 001 to one controlled false identity merge and its downstream effects.

    Aug 11, 2026 · Complete

  2. Methodology review

    Identified the required position factor and matched true-link utility control before design freeze.

    Aug 11, 2026 · Revise

  3. Targeted prior-work scan

    Recorded adjacent work on false aggregation, record linkage, downstream graph effects, and recovery; systematic search remains pending.

    Aug 23, 2026 · Initial scan

  4. Contract kernel implemented

    Added one frozen false/true pair, isolated oracle truth, immutable correction history, and a truth-blind policy-input projection without running a study.

    Aug 23, 2026 · Contract only

  5. Design freeze and benchmark implementation

    Freeze factor definitions and policies, complete methods review, then implement the seeded generator and output-level recovery checks.

    Next · Gated

Open research

Help make the study harder to fool.

We welcome methodological challenge, domain criticism, and implementation review before any study run.

Methods review

Challenge the draft design, propagation policies, and hypotheses before the preregistration is frozen.

Domain review

Tell us where the synthetic graph families diverge from the evidence you work with in practice.

Implementation review

Challenge whether the proposed generator, metrics, and recovery checks would implement the frozen design faithfully.

Challenge the One Bad Link study before we run it.

Review the method, challenge the assumptions, or pressure-test the implementation criteria before any study run.