live_v2Weighted mean over the answered domains, re-sharing a dropped domain's weight; imaging by recall at half weight.
The live score (FIT_SCORE_VERSION 2). The control every candidate is read against.
For surgeons · protocol v2
The analysis plan below was written before the data it describes existed, and shipped in the same release as the code that collects them. It is not a claim that the scores will be found valid. It is the statement of what would count as evidence, made in advance so the judgement is not made afterwards.
Primary documents: the validation protocol and the design brief it implements. Section numbers below refer to the protocol.
§2 · Intended use
These answer different questions and are validated against different things. Conflating them is the error the design brief exists to prevent, so the protocol refuses it at the outset.
The Fit Score will not be presented as a probability of benefit under this protocol.
The Fit Score · scoring version 2
Six domains over 100 points. A domain with no answer contributes nothing and is not counted against the patient; a weakly evidenced domain contributes a fraction.
A weighted mean lets a strong part cover for a weak one even where that makes no clinical sense: severe symptoms can outweigh a normal film or a doubtful pain source. The arithmetic is not corrected for this. Instead a presentation gate sits between the score and the page and decides what the result is allowed to say.
The joint-problem match is below the existing gate — another structure may explain part of the pain.
The MRI answer names avascular necrosis. A replacement indication this score was not built for.
A radiologist or a surgeon read the film as Kellgren-Lawrence 0–1 while the arithmetic would band.
A domain is unanswered or only partly answered, or a good band rests on a recalled grade.
Nothing above applies.
The order is the brief’s. A doubtful pain source outranks everything because it is the finding that most changes what operation, if any, is being discussed. Bands are 50 and above, 40 to 49, and below 40 — and a band is stated only when no gate above it applies.
A Kellgren-Lawrence grade is worth what its provenance supports, and the grade itself is never altered to achieve that:
Two things a reader should notice. No imaging drops the domain out of the average entirely rather than scoring it zero, because absent imaging must never read as a healthy joint. And a surgeon’s confirmation against the film does not outweigh a radiology report in the arithmetic — the two carry the same weight and differ only in recorded completeness. That is a deliberate choice and an arguable one; it is the kind of thing this page is published to have argued with.
The Readiness Score
8 equal parts of 12.5. All eight parts are required. Every one is reachable by every patient, so readiness is completeness — not a mark to clear.
Nothing in it subtracts points for the health a patient arrived with. A reported condition adds accountable work with a named owner — a clearance, a plan, a review — rather than deducting from a total. A record nobody has touched scores nothing at all rather than zero, because it has not failed, it has not run.
Beside the number is a written status read from the same requirement rows.
Your readiness order has not been reviewed yet.
Work is underway. Nothing here is a judgment — it is a count of what is still open.
A required item is past due or will not be current on surgery day. It needs action, yours or your team's.
Everything you own is done. Your surgical team still needs to review and accept the rest.
Every requirement is complete and your surgical team has accepted it. This is completion of the preparation pathway, not medical clearance.
The top rung is never awarded by arithmetic. It requires a named acceptance by a person on the surgical team; a patient’s own attestation does not reach it, and a record that has arrived but has not been reviewed is not complete.
§1 · Shadow mode
The brief asks that the system run without authority to release a case for surgery. It never had that authority: a surgeon’s named acceptance is what the status ladder waits for, and the Fit Score has never been an eligibility decision. So shadow mode here is not a switch. It is a set of records.
Nothing in shadow mode changes a score, a claim, a readiness status, or any sentence a patient or a surgeon is shown.
§4 · Pre-registered candidates · models version 1
The live score is the control. Each candidate is a proposal that was made and, in most cases, declined for display — scored silently so the decision can later be checked against outcomes rather than argued. A candidate added after outcome data exist is not a hypothesis and is reported as exploratory, never beside these five.
live_v2Weighted mean over the answered domains, re-sharing a dropped domain's weight; imaging by recall at half weight.
The live score (FIT_SCORE_VERSION 2). The control every candidate is read against.
cap_unread_imagingThe live score, capped at 49 when the X-ray grade is patient recall and at 39 when no grade is on file.
External memo, fit-score-improvements.md — declined for display by the brief §3.9.
no_renormalizeEach answered domain contributes value × nominal weight over the full 100; a dropped domain contributes nothing and its weight is not re-shared.
The brief §1.3, 'unknown must not be rewarded', taken literally as arithmetic.
treatment_doubledTreatment history at twice its weight, taken from joint health; otherwise the live combination.
External memo, readiness-and-fit-scores.md — declined by the brief §3.7 ('do not double the treatment-history weight').
joint_impact_only100 minus the HOOS JR / KOOS JR interval score. No other domain.
The brief §3.1 — the one validated instrument, on its own.
§5 and §6 · What would settle it
Triage agreement
81
comparable blinded reads per joint. Enough to estimate agreement to within ±10 percentage points at an expected 70%. Reported as percent agreement and Cohen’s κ, blinded and unblinded apart, hip and knee apart.
Outcome calibration
290
operated patients per joint with a twelve-month outcome measure. A calibration analysis for a binary outcome needs at least 100 events and 100 non-events per model per joint; at an expected 65% acceptable-state rate that is this number.
Outcomes are collected at the published windows, and a patient who does not proceed stays in the cohort — no surgery date 365 days after the assessment is itself an outcome, not an exclusion.
Per joint, for each candidate and for the claim gates: discrimination for reaching the published acceptable state at twelve months, calibration slope and intercept with the curve plotted, and decision-curve net benefit across a threshold range of 0.3 to 0.8. Missingness is reported by domain and by claim gate before any imputation. Subgroups are prespecified: joint always, then age band, sex, imaging tier, treatment-barrier answer, and whether the assessment came through the preparation front door.
No claim of benefit from surgery relative to continued nonoperative care is made from this observational cohort. No cut point, weight or band boundary moves until §5 and §6 have been run and the result is recorded in the protocol’s log.
§5 · What a surgeon actually does
On a case in your workspace the Fit Score and its claim are held back until you have recorded your own read of the same records — or chosen to see the score first. Either choice is one click, both are recorded, and a read taken after looking is kept and marked unblinded rather than discarded.
Agreement is counted only for reads that map to a claim. “Cannot say from the records” is reported and never scored, because a surgeon declining to call it on the evidence given is a finding about the evidence, not a disagreement.
The first question asked of a disagreement is which read was right, and that is answered by the outcome data, not by moving a threshold.
§7 and §8 · Readiness, and whether any of it is understood
Readiness is not a prediction model and is not validated as one. Every record whose parts reach complete, or whose status reaches Ready for Review or Ready for Surgery, is to be audited by a person against the requirement rows and the record. The finding is confirmed, or a false completion with what was open. The false-completion rate is the primary safety measure, reported as a rate of what was audited and never of what exists.
The requirement definitions, the validity windows and the teach-back prompts are to be reviewed line by line with arthroplasty surgery, anaesthesia, nursing, physical therapy, infection prevention, primary care, operations and patients. Until the first review is recorded, the validity windows carry their existing label as illustrative local policy rather than anyone’s standard.
Separately, one anonymous question on each result page asks what the reader took the number to mean. The success criterion is that at least 80% give the intended reading, for every claim kind separately. A claim kind below 80% is a copy defect on that headline and is fixed as one.
Governance
The candidate list, the endpoints, the subgroups and the success criteria are locked with this version. Changing an endpoint, a threshold, a candidate model or an analysis after data are in requires a new numbered version stating the change and its reason; the superseded text stays in the document.
Read the full protocol. Questions about the design, or a reading of it you disagree with, are the point of publishing it.