Key takeaways
- Ask four questions: what kind of study, how many people, who they were, and do others agree.
- Treat one small study as a hint. Trust several good studies that point the same way.
- Study type is a starting point. A meta-analysis is only as good as the studies inside it.
- "Linked to" is not "caused by". Studies that only watch people usually can't show cause.
- Most people in sport research studies were men. So there's less data on women. That doesn't show women respond differently.
To judge a fitness claim, ask four questions. What kind of study is behind it? How many people took part, and who were they? Did it show cause, or only a link? And have other studies found the same thing? One small study is a hint, not a finding. Several good studies that agree are evidence.
That sounds obvious. It gets hard when "studies show" is everywhere and nobody names the study.
What actually matters when you start lifting, the first article in this path, uses the citations and labels explained here.
What skipping the questions costs
A claim with a study attached feels settled. But the study might be a test on cells or one trial with 12 people. Act on every one and you'll switch programs monthly and pay for things that do little.
The opposite mistake costs too. Decide all research is junk and you're left with whoever sounds most confident.
What the research says about judging evidence
Some study types are stronger than others
The classic "evidence pyramid" ranks study types from weakest to strongest:[1]
- Lab studies and case reports. Lab studies use cells or animals. A case report describes one person or a few.
- Observational studies. Researchers watch groups of people without changing what they do.
- Randomised controlled trials (RCTs). People are split by chance into groups that get different treatments.
- Systematic reviews and meta-analyses. A systematic review collects every study on one question in a planned way. A meta-analysis pools their results into one estimate.
The authors of a 2016 paper on the pyramid call this order "intuitive and likely correct in many instances".[1] Expert consensus. This is the authors' view, not a tested finding.
Study type is a starting point, not the verdict
The same paper argues the classic pyramid needs revising. Study design alone, the authors write, appears to be too little on its own to judge how biased a study is.[1] A meta-analysis is only as good as the studies inside it. One that pools well-run RCTs can't be equated with one that pools weaker observational studies.[1]
A widely used system called GRADE rates evidence this way. Under GRADE, evidence from RCTs starts as high quality and evidence from observational studies starts as low.[2] Then it can be marked down for five reasons:[2]
- flaws in how the studies were run;
- results that disagree with each other;
- studies that tested a different group or question;
- imprecise results;
- signs that unfavourable results went unpublished.
At GRADE's top level, more research is "very unlikely" to change how confident we are in the result. At the bottom, any estimate is "very uncertain".[2] Both papers come from medicine, not sport science.
The five labels this blog uses
Our labels borrow one idea from GRADE: how sure you can be depends on more than study type.[1][2] But they are our own scale, not GRADE ratings.
- Strong evidence: several meta-analyses or a position stand agree, the results are consistent, and the studies included people like you.
- Moderate evidence: one good meta-analysis, or several RCTs that mostly agree, with some gaps (short studies, small samples).
- Mixed evidence: the studies disagree, or the effect depends on the setup.
- Limited evidence: a few small studies, short trials, or results borrowed from a different group of people.
- Expert consensus: practice recommended by position stands or coaches, without direct trials. We add the line that fits the source: "This is the panel's (or authors') view, not a tested finding." or "This is coaching practice, not a tested finding." The second may later shorten to "Coaching practice."
Sport science runs on small studies
Sample size is the number of people in a study. In 2020, the editors of the Journal of Sports Sciences checked 120 papers sent to them over three years. The typical study had 19 participants. Only 13 of the 120 (about 1 in 10) had worked out in advance how many people they needed.[3] Limited evidence: one journal's check of papers sent to it, not of published papers.
Small studies give fuzzy answers. In the editors' example, a study of 19 people could place a modest effect anywhere from a small negative effect to a large positive one.[3] And when a small study does find a "statistically significant" result (one that passes the standard test for chance), it will likely overestimate the effect.[3]
Results often shrink when repeated
A replication repeats a study to see if the result holds.
Limited evidence: one project of 25 studies from applied sport and exercise science. It isn't a measure of how much fitness research is wrong.
A well-known 2005 paper by John Ioannidis used statistical models to make a related argument. A finding is less likely to be true, it said, when studies and effects are small.[5] The same goes when the analysis is flexible, or money or strong beliefs are involved.[5] Despite its title, it's a general argument about research, with most of its examples from medicine. It's a modelling argument, not a count of false studies. Its advice: "What matters is the totality of the evidence."[5] Expert consensus. This is the author's view, not a tested finding.
Most of the people studied were men
Two audits counted who takes part. One covered 1,382 papers in three major sports-medicine journals: 39% of the participants were women.[6] A later one covered about 5,261 papers in six journals from 2014 to 2020. Women were 34% of participants, 31% of papers studied only men, and 6% studied only women.[7] Moderate evidence.
"Linked to" isn't "caused by"
An association means two things show up together, not that one causes the other. One analysis of 15 studies followed about 47,000 adults. People who took more steps a day were less likely to die during the follow-up.[9] Moderate evidence for that link. The authors call it an association, which is the honest word.[9] People who walk more may also be healthier to begin with.[9] Under GRADE, observational evidence like this starts at low quality.[2] More on steps in daily steps.
Expert summaries describe groups, not you
A position stand is an expert panel's summary of the evidence, not a new experiment. The American College of Sports Medicine's 2026 stand on lifting drew on 137 systematic reviews covering more than 30,000 people.[10] Even so, the panel calls it group-level evidence that individuals will sometimes need to adjust.[10] Expert consensus. This is the panel's view, not a tested finding.
A checklist for any claim
When a post says "studies show", find the study on PubMed, a free search engine for medical research. Read the abstract, the short summary at the top. Then ask:
- What kind of study is it? Cells, animals and one person's story sit at the bottom. A meta-analysis of good trials sits near the top.
- How many people? In one journal's check, the typical study had 19.
- Who were they? Check sex, age, whether they already trained, and how long it ran.
- Cause or link? If researchers only watched people, it shows a link. "Associated with" turned into "causes" is a red flag.
- Do other studies agree? Look for a review or meta-analysis on the same question.
A made-up example: a post says "Study shows 3 sets build as much muscle as 10." The abstract shows one trial: 14 trained young men, 8 weeks. That's a hint. Look for a meta-analysis on weekly sets first.
Common mistakes
- Treating the newest study as the answer. New feels like news. But results often shrink when repeated. Wait for a second study.
- Trusting anything labelled "meta-analysis". It sounds final. Check what went into it.
- Assuming women need a different plan. It sounds logical when most studies used men. But less data isn't proof of a different response. Check who was studied.
How Gymbers fits
The Gymbers app estimates nutrition targets with the Mifflin-St Jeor equation times an activity factor, and your one-rep max with the Epley formula. How we source lists the methods the app uses.
References
- 1.
- 2.
- 3.
- 4.
- 5.
- 6.
- 7.
Cowley ES, Olenick AA, McNulty KL, Ross EZ (2021). "Invisible Sportswomen": The Sex Data Gap in Sport and Exercise Science Research. Women in Sport and Physical Activity Journal.
ObservationalDOI
- 8.
- 9.
- 10.
Currier BS, D'Souza AC, Singh MAF, et al. (2026). American College of Sports Medicine Position Stand. Resistance Training Prescription for Muscle Function, Hypertrophy, and Physical Performance in Healthy Adults: An Overview of Reviews. Medicine & Science in Sports & Exercise.