For surgeons · protocol v2

What these scores claim, and how we are testing whether they are right.

The analysis plan below was written before the data it describes existed, and shipped in the same release as the code that collects them. It is not a claim that the scores will be found valid. It is the statement of what would count as evidence, made in advance so the judgement is not made afterwards.

Primary documents: the validation protocol and the design brief it implements. Section numbers below refer to the protocol.

§2 · Intended use

Four outputs, four different standards.

These answer different questions and are validated against different things. Conflating them is the error the design brief exists to prevent, so the protocol refuses it at the outset.

OutputIntended useValidated against
Joint ImpactDescribes current impairment.Nothing new. The instrument is published and validated; we check only that our scoring matches the published crosswalk.
Indication Evidence, and the claim the page leads withReferral triage: whether the usual replacement indication is established, and what would settle it.Blinded independent surgeon assessment (§5).
The Fit Score itselfA descriptive summary, shown with its parts. Not an outcome prediction and not an eligibility decision.Descriptive only under this protocol. Its relationship to outcomes is measured (§6); no probability claim is made or planned from these data alone.
Readiness Score and the status ladderPreparation progress, and whether the surgical team has confirmed it.Workflow and safety audit (§7). No prediction model.

The Fit Score will not be presented as a probability of benefit under this protocol.

The Fit Score · scoring version 2

A weighted mean, and the reason it is not allowed to speak for itself.

Six domains over 100 points. A domain with no answer contributes nothing and is not counted against the patient; a weakly evidenced domain contributes a fraction.

DomainWeightWhat it reads
Joint health34100 − HOOS JR / KOOS JR. The only part from a published, validated instrument.
X-ray severity18Kellgren-Lawrence, weighted by how well the grade is evidenced.
Joint-problem match18Whether this joint, rather than the back or the hip, explains the limitation.
Quality-of-life impact16Independence, work or caregiving, valued activities.
Pain pattern8Frequency and night pain — what the JR instrument does not ask.
Treatment history6Whether appropriate care was tried and did not hold.

A weighted mean lets a strong part cover for a weak one even where that makes no clinical sense: severe symptoms can outweigh a normal film or a doubtful pain source. The arithmetic is not corrected for this. Instead a presentation gate sits between the score and the page and decides what the result is allowed to say.

  1. 1
    Confirm the pain source first

    The joint-problem match is below the existing gate — another structure may explain part of the pain.

  2. 2
    A different pathway, not arthritis arithmetic

    The MRI answer names avascular necrosis. A replacement indication this score was not built for.

  3. 3
    Reviewed imaging shows minimal arthritis — reassess first

    A radiologist or a surgeon read the film as Kellgren-Lawrence 0–1 while the arithmetic would band.

  4. 4
    Provisional — with what is missing named

    A domain is unanswered or only partly answered, or a good band rests on a recalled grade.

  5. 5
    The arithmetic band, stated

    Nothing above applies.

The order is the brief’s. A doubtful pain source outranks everything because it is the finding that most changes what operation, if any, is being discussed. Bands are 50 and above, 40 to 49, and below 40 — and a band is stated only when no gate above it applies.

Imaging provenance scales the domain’s weight, never its value.

A Kellgren-Lawrence grade is worth what its provenance supports, and the grade itself is never altered to achieve that:

  • No imaging on fileweight ×0 · completeness +0
  • As the patient recalls itweight ×0.5 · completeness +10
  • From the radiology reportweight ×1 · completeness +20
  • Confirmed by the surgeon against the filmweight ×1 · completeness +25

Two things a reader should notice. No imaging drops the domain out of the average entirely rather than scoring it zero, because absent imaging must never read as a healthy joint. And a surgeon’s confirmation against the film does not outweigh a radiology report in the arithmetic — the two carry the same weight and differ only in recorded completeness. That is a deliberate choice and an arguable one; it is the kind of thing this page is published to have argued with.

The Readiness Score

Completion of preparation work. Not a judgement of the patient.

8 equal parts of 12.5. All eight parts are required. Every one is reachable by every patient, so readiness is completeness — not a mark to clear.

  • Medical optimization
  • Medical clearances
  • Medication review
  • Additional testing
  • Home and social support
  • Mobility and recovery support
  • Expectations
  • Education

Nothing in it subtracts points for the health a patient arrived with. A reported condition adds accountable work with a named owner — a clearance, a plan, a review — rather than deducting from a total. A record nobody has touched scores nothing at all rather than zero, because it has not failed, it has not run.

Beside the number is a written status read from the same requirement rows.

  1. 1
    Not Started

    Your readiness order has not been reviewed yet.

  2. 2
    In Progress

    Work is underway. Nothing here is a judgment — it is a count of what is still open.

  3. 3
    Needs Attention

    A required item is past due or will not be current on surgery day. It needs action, yours or your team's.

  4. 4
    Ready for Review

    Everything you own is done. Your surgical team still needs to review and accept the rest.

  5. 5
    Ready for Surgery

    Every requirement is complete and your surgical team has accepted it. This is completion of the preparation pathway, not medical clearance.

The top rung is never awarded by arithmetic. It requires a named acceptance by a person on the surgical team; a patient’s own attestation does not reach it, and a record that has arrived but has not been reviewed is not complete.

§1 · Shadow mode

Records written beside the product, read by nobody but the report.

The brief asks that the system run without authority to release a case for surgery. It never had that authority: a surgeon’s named acceptance is what the status ladder waits for, and the Fit Score has never been an eligibility decision. So shadow mode here is not a switch. It is a set of records.

  • Shadow mode: every candidate model is scored on every saved assessment and shown to nobody. The live score is the control. Candidates are pre-registered in docs/validation-plan.md; one added after outcomes are in is not a hypothesis.
  • A blinded read is a surgeon's own triage recorded before the Fit Score was shown. Agreement is counted only for reads that map to a claim; "cannot say" is reported, never scored.
  • A readiness audit is a person's finding on a record the arithmetic called complete. False completions are the defect this system exists to catch, and are reported as a rate of what was audited, never of what exists.
  • Comprehension answers are anonymous counters keyed by what the page claimed. The first answer is counted; the page then states the right reading, so the question teaches as it measures.
  • Workflow defects are read off the final OTIF each surgery day records: which requirements were still open when the patient arrived. Nothing here is a judgment of a patient or a surgeon.

Nothing in shadow mode changes a score, a claim, a readiness status, or any sentence a patient or a surgeon is shown.

§4 · Pre-registered candidates · models version 1

Five models scored on every save, shown to nobody.

The live score is the control. Each candidate is a proposal that was made and, in most cases, declined for display — scored silently so the decision can later be checked against outcomes rather than argued. A candidate added after outcome data exist is not a hypothesis and is reported as exploratory, never beside these five.

live_v2

Weighted mean over the answered domains, re-sharing a dropped domain's weight; imaging by recall at half weight.

The live score (FIT_SCORE_VERSION 2). The control every candidate is read against.

cap_unread_imaging

The live score, capped at 49 when the X-ray grade is patient recall and at 39 when no grade is on file.

External memo, fit-score-improvements.md — declined for display by the brief §3.9.

no_renormalize

Each answered domain contributes value × nominal weight over the full 100; a dropped domain contributes nothing and its weight is not re-shared.

The brief §1.3, 'unknown must not be rewarded', taken literally as arithmetic.

treatment_doubled

Treatment history at twice its weight, taken from joint health; otherwise the live combination.

External memo, readiness-and-fit-scores.md — declined by the brief §3.7 ('do not double the treatment-history weight').

joint_impact_only

100 minus the HOOS JR / KOOS JR interval score. No other domain.

The brief §3.1 — the one validated instrument, on its own.

§5 and §6 · What would settle it

Two counts, and neither is reached yet.

Triage agreement

81

comparable blinded reads per joint. Enough to estimate agreement to within ±10 percentage points at an expected 70%. Reported as percent agreement and Cohen’s κ, blinded and unblinded apart, hip and knee apart.

Outcome calibration

290

operated patients per joint with a twelve-month outcome measure. A calibration analysis for a binary outcome needs at least 100 events and 100 non-events per model per joint; at an expected 65% acceptable-state rate that is this number.

Outcomes are collected at the published windows, and a patient who does not proceed stays in the cohort — no surgery date 365 days after the assessment is itself an outcome, not an exclusion.

  • Before surgery (baseline)day -90 to 0 · required
  • About 90 days after surgeryday 61 to 120 · optional
  • One year after surgery (CMS window)day 300 to 425 · required

Per joint, for each candidate and for the claim gates: discrimination for reaching the published acceptable state at twelve months, calibration slope and intercept with the curve plotted, and decision-curve net benefit across a threshold range of 0.3 to 0.8. Missingness is reported by domain and by claim gate before any imputation. Subgroups are prespecified: joint always, then age band, sex, imaging tier, treatment-barrier answer, and whether the assessment came through the preparation front door.

No claim of benefit from surgery relative to continued nonoperative care is made from this observational cohort. No cut point, weight or band boundary moves until §5 and §6 have been run and the result is recorded in the protocol’s log.

§5 · What a surgeon actually does

One click, before the score is shown.

On a case in your workspace the Fit Score and its claim are held back until you have recorded your own read of the same records — or chosen to see the score first. Either choice is one click, both are recorded, and a read taken after looking is kept and marked unblinded rather than discarded.

  1. A replacement candidate, on what is here
  2. Not yet — the case for replacement is not made
  3. Confirm the pain source first
  4. A different pathway from primary osteoarthritis
  5. Cannot say from the records

Agreement is counted only for reads that map to a claim. “Cannot say from the records” is reported and never scored, because a surgeon declining to call it on the evidence given is a finding about the evidence, not a disagreement.

Low agreement is information about the gates before it is an argument about the cut points.

The first question asked of a disagreement is which read was right, and that is answered by the outcome data, not by moving a threshold.

§7 and §8 · Readiness, and whether any of it is understood

A false completion is the defect this exists to catch.

Readiness is not a prediction model and is not validated as one. Every record whose parts reach complete, or whose status reaches Ready for Review or Ready for Surgery, is to be audited by a person against the requirement rows and the record. The finding is confirmed, or a false completion with what was open. The false-completion rate is the primary safety measure, reported as a rate of what was audited and never of what exists.

The requirement definitions, the validity windows and the teach-back prompts are to be reviewed line by line with arthroplasty surgery, anaesthesia, nursing, physical therapy, infection prevention, primary care, operations and patients. Until the first review is recorded, the validity windows carry their existing label as illustrative local policy rather than anyone’s standard.

Separately, one anonymous question on each result page asks what the reader took the number to mean. The success criterion is that at least 80% give the intended reading, for every claim kind separately. A claim kind below 80% is a copy defect on that headline and is fixed as one.

Governance

A protocol change is a new version, never an edit.

The candidate list, the endpoints, the subgroups and the success criteria are locked with this version. Changing an endpoint, a threshold, a candidate model or an analysis after data are in requires a new numbered version stating the change and its reason; the superseded text stays in the document.

Read the full protocol. Questions about the design, or a reading of it you disagree with, are the point of publishing it.