clinical
Cardiovascular outcomes: what SELECT showed
The trial that changed how clinicians talk about these drugs — who was in it, what it actually found, how big the effect was in plain terms, and the claims stacked on top of it.
A peer support community, independent and not for sale. Since February 2024.
clinical · how-to
The method our reading group uses on every paper — five questions, in order, that let anyone tell what a study actually showed from what people are claiming it showed.
11 min read1.4k wordsUpdated 3 July 2026Reviewed by Sunil
Sunil started the research reading group after watching the same thing happen four times in a year: a decent trial is published, a press release compresses it, a headline compresses the press release, and by the time it reaches anybody it has become a claim the researchers never made and would not endorse.
You do not need statistical training to catch most of that. You need a fixed order of questions and the willingness to stop reading when you have the answer. The group meets fortnightly, works through one paper, and nobody is expected to have read it all — most people read the abstract, the tables, and the limitations paragraph, which is genuinely most of the value.
A caution before the method. Being able to read a paper does not make you your own clinician, and we have watched people go badly wrong by deciding they have out-reasoned their prescriber from an armchair. What this skill actually buys you is better questions and a stronger immunity to nonsense. Both are worth having.
Always first, and it settles more arguments than everything else combined.
Find the inclusion and exclusion criteria — usually a paragraph in the methods, and always in the protocol if the paper is vague. Then compare that list with yourself, honestly.
SELECT (Lincoff, New England Journal of Medicine, 2023) required participants to be at least 45, to have a body mass index of 27 or above, to have established cardiovascular disease, and to not have diabetes. That single sentence disqualifies most of the internet arguments about SELECT, because most of them are about people who would not have been enrolled.
FLOW required type 2 diabetes and chronic kidney disease. SURPASS-2 required type 2 diabetes on metformin. STEP 1 excluded diabetes. These populations are not interchangeable and results do not automatically travel between them.
Also look at who was not there. Age ranges, the proportion of women, ethnicity, people with severe kidney or liver impairment, people with a history of an eating disorder, people already on other treatments. Every exclusion is a group the trial cannot speak for. Noor, who reads with us, calls this the shadow of the study and says it is usually more informative than the conclusion.
A drug is never good in the abstract; it is better or worse than something, at something, over some period.
The comparator. Placebo tells you whether a drug does anything. An active comparator tells you whether it beats the alternative, which is usually the question you actually care about. SURPASS-2 compared tirzepatide with semaglutide 1 mg — a real head-to-head — but note that it used the diabetes dose, so it does not tell you about the higher weight-management dose. Comparator details like that are where most of the meaning hides.
The endpoint. Was it something that happened to people — a heart attack, kidney failure, a death — or a number that moved? Both are legitimate, but they answer different questions. A surrogate endpoint is a bet that moving the number moves the outcome, and medicine has a long history of that bet failing.
Was it pre-specified? The primary endpoint is declared before the trial starts. Anything found afterwards by looking around in the data is hypothesis-generating, however striking it looks. If a result is described as exploratory, post-hoc, or a subgroup analysis, it is a suggestion, not a finding.
How long? STEP 1 ran 68 weeks. SURMOUNT-1 ran 72. SELECT followed people for a mean of a little over three years. Anything you want to know about year five is outside all of them.
This is the question that defuses most headlines, and it is arithmetic anyone can do.
A relative risk reduction says how much the risk shrank in proportion. An absolute risk reduction says how many fewer people had the event. In SELECT, 6.5 per cent of the semaglutide group had a major cardiovascular event against 8.0 per cent on placebo. That is a hazard ratio of 0.80 — a 20 per cent relative reduction — and an absolute difference of about 1.5 percentage points over roughly three years.
Identical result, two framings, wildly different emotional weight. Press releases overwhelmingly choose the relative one. When a paper only gives you relative figures, dig the raw event counts out of the tables and work it out.
Then look at the confidence interval, which is the honest expression of how uncertain the estimate is. SELECT reported 0.72 to 0.90 — comfortably away from 1, so the direction is secure. An interval that scrapes past 1, or that spans anything from trivial to enormous, means the trial has not pinned the effect down, whatever the p-value says. Sunil’s line: the point estimate is the story the trial tells, the interval is how much of it it can back up.
Trial results are reported as means, and means conceal people.
STEP 1 (Wilding, NEJM, 2021) reported a mean weight change of about 14.9 per cent with semaglutide 2.4 mg against about 2.4 per cent with placebo over 68 weeks. SURMOUNT-1 (Jastreboff, NEJM, 2022) reported means ranging from around 15 to around 21 per cent across tirzepatide doses over 72 weeks. Those averages are correct and they are also the source of enormous unnecessary distress, because the distribution around them is wide in both trials. Real participants sat well above and well below.
Look for the responder analyses and the distribution charts, which most papers of this kind include. They tell you what proportion reached various thresholds, and they make plain that a substantial minority did not respond much. No reliable way to predict who will be where exists.
Which is why nothing in a trial is a target, and why we do not repeat those percentages as goals in this community. They are descriptions of what happened to a group of strangers under trial conditions. Ada puts it more bluntly in the plateau circle: a mean is a statement about a population and never a promise made to a person.
The most under-read part of any trial, and often the most honest.
Find the flow diagram. How many were randomised, how many completed, how many withdrew, and why? Discontinuation for gastrointestinal adverse events is reported across the weight-management trials and is a real feature of these drugs, and a trial with a research nurse on the end of the phone will always retain people better than ordinary care does.
Then ask how the analysis handled them. An intention-to-treat analysis keeps everyone in the group they were assigned to and is generally the conservative, trustworthy choice. Analyses of people who completed treatment as intended flatter the drug, because they quietly remove the people it did not suit. Papers often report both — the gap between them tells you something.
Finally, read the limitations paragraph, which authors write honestly far more often than they are given credit for, and skim the funding and conflict-of-interest statements. Nearly all of these trials were funded by the manufacturer. That is normal and it is not disqualifying; it is context, and it belongs alongside everything else rather than instead of it.
Once you have the method, the failure modes become easy to spot. The recurring ones in this field:
If you want to practise, the group works through one paper a fortnight and newcomers are explicitly welcome to say they did not understand a table. That sentence gets said in most sessions, usually by someone who has been coming for a year.
Your product leaflet and the person who prescribed for you override anything on this page.
clinical
The trial that changed how clinicians talk about these drugs — who was in it, what it actually found, how big the effect was in plain terms, and the claims stacked on top of it.
clinical
eGFR, creatinine and albumin explained without jargon, what the FLOW trial found in people with type 2 diabetes and kidney disease, and the everyday kidney risk that matters more.
clinical
What glycated haemoglobin actually measures, the units confusion nobody warns you about, the things that make it lie, and why it is a poor scoreboard for a fortnight.
clinical
The rodent finding behind the warning label, what human data has and has not shown since, what it means if you already take levothyroxine, and where the honest uncertainty sits.
starting
The mechanism in plain language, which drug is which, what the major trials found, and an honest accounting of where the evidence runs out.
clinical
Why change slows or stops, what the trial curves actually look like, what is worth checking, and why a plateau is a physiological event rather than a verdict on you.