how much of what we believe about endpoint actually comes from STEP threads
how much of what we believe about endpoint actually comes from STEP threads. I would rather ask a basic question now than get this wrong quietly for two months.
Spent an evening with the appendix tables and found the subgroup detail that the entire thread had been speculating about.
Quoted a figure here confidently, got asked whether it was ITT, went and checked, and it was not. Learned something.
Read a press release and the publication three months apart. The hedging in the second one was substantial.
If two or three other people have done the same thing we might actually learn something. Alone it is an anecdote.
best — the order this archive was captured in
Why comparing across trials almost never works, with the specific failure modes.
Different populations: an obesity programme and a diabetes programme enrol different people with different baseline characteristics. Different endpoints: body weight change, glycaemic control and cardiovascular events are not convertible. Different durations: 68 weeks and 72 weeks are not the same, and the curves have not flattened by either.
Different analysis populations: one paper reports intention-to-treat, another emphasises completers. Different support: some trial designs include structured lifestyle contact that no member of this board receives.
Stack those and the "X beats Y" tables that circulate here are comparing five things at once and attributing the difference to the molecule.
On means, which this board treats as targets and which are nothing of the sort.
A reported mean body weight change is the centre of a distribution that in these trials is very wide. Substantial numbers of participants did much better, and substantial numbers did considerably worse while remaining on the drug and in the analysis.
Quoting the mean as an expectation therefore misleads in both directions: it makes ordinary results look like failures and it makes exceptional results look normal. If a paper publishes the distribution — and several do, in the appendix — look at that instead. It is far more informative than the number in the abstract.
open-label extensions are not the same evidence as the randomised phase
Went looking for the registered protocol to see whether the endpoint had changed. It had not, which was reassuring and worth checking.
The discontinuation numbers were the most useful thing in the paper for me and they were in a supplementary table.
Started keeping the trial identifiers straight in a note file because I kept mixing up two programmes in the same sentence.
Added the trial identifier to the title so this thread is findable in two years.
Argued for a week about a result and then read the limitations section, which conceded most of my opponent’s point.
Cardiovascular outcome trials are powered for events, not for weight, and are typically run in a different population with different inclusion criteria. Reading a weight number out of one is reading a secondary endpoint.
Cardiovascular outcome trials are powered for events, not for weight, and are typically run in a different population with different inclusion criteri
nayeli_fonseca is right about the programme names. They are different populations with different endpoints.
The registered protocol is public. Comparing the registered primary endpoint with the reported one is a two-minute check and it is how outcome switching gets caught.
How long was the randomised phase before any extension?
intention to treat versus completers changes the number substantially
Absolute or relative risk reduction?
intention to treat versus completers changes the number substantially
This is the distinction that would end about half the arguments on this board.
Which trial, and which arm?
Compared myself to a trial mean for about six months before realising the trial arm had dietitian contact every fortnight.
Compared myself to a trial mean for about six months before realising the trial arm had dietitian contact every fortnight.
Agreed. And the interval, not the point estimate, is what the trial actually established.
SURPASS is the diabetes programme and reports glycaemic endpoints
Discontinuation rates are a tolerability result. A trial with a strong efficacy number and heavy discontinuation is telling you two things and people only quote one.
This. Intention-to-treat versus completer analysis routinely moves the headline by several points.
Small fix — that was the cardiovascular outcomes trial, so weight was a secondary endpoint and the population was different.
registry entry, protocol, publication — three different documents
check who the comparator was before you compare anything
That is a relative risk reduction. Quoting it without the absolute numbers overstates the case considerably.
A confidence interval is the range of effects compatible with the data. Two trials with overlapping intervals have not disagreed, whatever their point estimates look like next to each other.
A confidence interval is the range of effects compatible with the data.
Adding the check nobody runs — the registered protocol is public and takes two minutes to compare.
- 1How long was the randomised phase before any extension?7 comments in this branch · started by u/nikhil_lindqvist
- 2Why comparing across trials almost never works, with the specific failure…6 comments in this branch · started by u/plateau_patrol