When Better Models Meet Missing Overlap

When Better Models Meet Missing Overlap

What the Itvisma JCA says about designing evidence for questions someone else will define.

 

The new Joint Clinical Assessment of intrathecal onasemnogene abeparvovec is worth reading for reasons that go well beyond spinal muscular atrophy. At first glance, it looks like another story of PICO complexity. But the more interesting lesson is about when evidence architecture becomes fixed, and how little can be repaired afterwards.

The assessment target itself moved during the procedure. After confirmation of the dossier, the claimed indication was narrowed from patients aged six months and older to those aged two years and older. The assessment scope was revised accordingly, removing two of the original PICO questions. Under the HTA Regulation the claimed indication is not a fixed input to the assessment scope. It can move, and the scope moves with it.

Twelve remained. Five could not be addressed with comparative evidence at all, including comparisons between the intrathecal and intravenous formulations because the available populations did not overlap sufficiently.

The problem is fit

The available comparative evidence covers a substantially narrower patient population than the claimed indication: patients who could sit but had not walked independently. Patients unable to sit and independently ambulant patients are not represented.

Once indirect comparisons were needed, that lack of fit became increasingly visible: against nusinersen, matching reduced the effective sample size by around 40%. Further adjustment for important covariates could reduce the effective sample size to 21 or below. Against risdiplam, effective sample size fell by roughly 63-67%, leaving the analysis increasingly dependent on a small number of heavily weighted observations.

Important treatment effect modifiers remained difficult to balance. In the base-case MAICs, SMN2 copy number, a major genetic driver of disease severity, was not adjusted for, contributing to residual imbalance and uncertainty. And the scrutiny does not stop with the intervention trial. CHERISH, used to anchor the comparison with nusinersen, carries high risk-of-bias ratings related particularly to attrition and selective reporting. In other words, the comparator trial is not simply a convenient fixed point in the network. It is evidence too, and it is assessed as such.

The same data can answer one question and fail another

A smaller detail makes the same point rather elegantly: for treatment-experienced patients, the developer submitted a pooled comparator group containing patients treated with either nusinersen or risdiplam.

For the nusinersen PICO, the assessors accepted the analysis pragmatically because around 80% of the pooled comparator group were receiving nusinersen. For the risdiplam PICO, they did not accept the same pool as adequate comparative evidence because only around 20% were receiving risdiplam.

That is not inconsistency. The same dataset does not have the same evidentiary value for every question asked of it.

The report makes another related point repeatedly: failure to demonstrate a statistically significant difference is not evidence of equivalence. Wide confidence intervals can reflect uncertainty, not comparability.

Better mathematics cannot redesign the evidence

Perhaps the most important passage sits deep in the appendix: as supportive analysis, the developer submitted a multilevel network meta-regression, allowing more extensive adjustment across a connected evidence network. The assessors explicitly acknowledge the methodological sophistication.

But they then make the decisive point: where covariate overlap between trial populations is poor, the model is forced to extrapolate beyond the observed data. A mathematically superior framework cannot remove the underlying bias created by insufficient overlap. Adjustment is not a substitute for design.

So what should change?

For developers approaching JCA, the planning question therefore has to move upstream. Not simply: Will our pivotal study support regulatory approval?

But: Which comparative questions will this evidence still be able to answer when populations, prior treatment status and comparators are defined by an assessment process we do not fully control? Where might exchangeability fail? Which treatment effect modifiers have to be captured? Which comparator strategies will remain defensible? And which gaps will be impossible to reconstruct retrospectively?

There was no Joint Scientific Consultation under the HTA Regulation in this case. We cannot know whether one would have solved these problems. But that is almost beside the point. JSC can be one part of the answer. As is PICO simulation. The larger task starts earlier: building an evidence architecture around the questions the evidence may eventually be asked to answer.

Because JCA may be a new assessment process. But by the time it starts, many of its most important answers have already been designed into the evidence.

Our expertise:

Related articles

Get in touch

Contact us to discover more about Kintiga and how we can help shape a future of impactful, patient-focused solutions across Europe.

Scroll to Top