You had a rough session this week. The reps came slower, the bar felt heavier than usual, and the pump never really showed up. Now you're staring at your training log, wondering if the program just stopped working.
Here's the direct answer: almost every lift carries real, measurable noise, often several percent, from one test to the next, even when your actual strength hasn't changed at all. A single session that lands a few percent below your best sits comfortably inside that normal wobble. It isn't proof of anything. A real plateau isn't one bad day. It's a pattern that holds up across several weeks, not a single number that dipped.
This article explains where that noise comes from, how big it usually is, and how many weeks of data actually separates a real stall from ordinary bounce. The fastest way to check it for yourself is to pull up your own logged sets from the past several weeks and look for a pattern, not a single low number.
Key Takeaways
- 1RM strength tests show a coefficient of variation of roughly 0.5-12% (median about 4.2%), even under controlled lab conditions, across 32 studies (Grgic, Lazinica, Schoenfeld & Pedisic, 2020).
- A single session a few percent below your best isn't necessarily a real decline. It can be fully explained by normal test-to-test noise.
- Sports scientists call this the "typical error," and the standard rule is that a change has to clearly clear that noise floor before you treat it as real (Hopkins, 2000).
- Competitive lifters gain roughly 18-25 kg a year on squat and deadlift, and roughly 10-13 kg a year on bench press, which is a small fraction of a kilogram in a normal week, smaller than the noise itself.
- Look for a trend across 6-8 weeks and multiple sessions before you decide a plateau is real, not one session.
Every Strength Number You Log Has Built-In Wobble
Even in a research lab, with a standardized warm-up and the same tester every time, the same person retested on the same lift rarely produces the exact same number twice. That's not a flaw in your training. It's baked into how strength testing works.
A 2020 systematic review pooled 32 studies and more than 1,595 lifters and found that one-rep max test-retest reliability, meaning how closely two tests of the same person's max actually match, varies from 0.5% to 12.1%, with a typical (median) swing of about 4.2% (Grgic, Lazinica, Schoenfeld & Pedisic, Sports Medicine - Open, 2020). That held up across training experience, sex, age, and whether the lift used one joint or several.
Put a number on that: at a 100 kg squat, a 4.2% swing is about 4 kg either direction. A session that lands 3 kg light of your best isn't a red flag. It's ordinary.
Think of it like a bathroom scale. Step on it twice in a row and you might see two slightly different numbers, even though your actual weight didn't change. Your lift has that same kind of built-in noise: your nervous system, your grip that day, how warmed up you happen to feel.
Why Sports Scientists Use "Typical Error," Not Raw Numbers
Sports scientists have a name for exactly this problem: the typical error, the normal amount a measurement bounces around even when nothing real has changed. A result has to clearly clear that noise floor before you trust it as a real change, rather than reacting to it as one.
This framework comes from sports scientist Will Hopkins, whose 2000 paper on reliability in sports medicine and science is still the standard reference for how researchers separate a genuine change in performance from ordinary measurement bounce (Hopkins, Sports Medicine, 2000). It's the model behind the idea that a result should clearly exceed the smallest change worth caring about, relative to that typical error, before anyone treats it as meaningful.
It's like judging whether a plant grew by measuring it with a ruler that has slightly blurry lines. If the plant only grew a millimeter, you can't tell that apart from where the blurry line happens to fall. You need it to grow enough to clearly pass the blur before you trust the reading.
For the shorter version of this same idea, applied directly to a squat number, see the noise section of the full diagnostic guide to training plateaus. This article is the expanded version of that same argument.
What a Normal Week of Progress Actually Looks Like
Real week-to-week strength gains, especially once you're past the beginner stage, are usually smaller than the measurement noise itself. That's exactly why one session can't tell you much.
A 2022 analysis of long-term strength gains in competitive powerlifters found squat strength climbing roughly 20.2 to 25.4 kg a year, deadlift roughly 18.1 to 21.1 kg a year, and bench press roughly 10.5 to 12.8 kg a year (Latella, Owen, Davies, Spathis, Mallard & Van Den Hoek, Medicine & Science in Sports & Exercise, 2022).
Spread across 52 weeks, the top end of that squat range works out to under half a kilogram a week. A normal week of real progress can be smaller than a single kilogram, while the noise in the test itself runs several kilograms wide. Expecting a visibly higher number every single session sets a bar that real progress, even good progress, usually can't clear on that timeline. It tends to show up as a trend over a month or two, not a jump every Tuesday.
One Bad Session Has a Lot of Innocent Explanations
A single rough session is far more likely to come from an ordinary, temporary cause than from your training suddenly failing.
A stressful day at work, a missed or rushed meal beforehand, a slightly off night of sleep, being a few days into a new gym or schedule, or simply the test-to-test noise covered above: any of these can turn a normal session into a rough one that has nothing to do with your program.
Sleep is worth being precise about, because it's easy to overstate. A 2024 study found that a chronic, sustained sleep reduction of one to two hours did not blunt resistance-training gains over a full training block (Borba et al., Sleep Science, 2024). If a sustained sleep deficit didn't derail progress in that study, one rough night before a single session is a weak explanation for a plateau on its own. For the fuller picture of where that line actually sits, see sleep and strength performance: one bad night versus a weekly pattern.
None of this is permission to ignore a real pattern. It's a caution against building a plateau story out of a single data point that has an obvious, boring explanation.
How Many Weeks of Data You Actually Need
A reasonable rule of thumb: look for 3 to 4 comparable data points spread across 6 to 8 weeks on the same exercise, at a similar effort level, before you conclude a plateau is real.
That window exists for a specific reason. Real progress, a fraction of a kilogram a week, needs time to add up to something bigger than the test's own noise, which runs several percent wide. Compress the window and you can't tell a real stall from a normal dip. Stretch it out to 6-8 weeks and a genuinely flat trend has had enough time to separate itself from the noise.
In practice, that means comparing top sets in the same rep range, logged at a similar RIR, across that window, not comparing random personal-record attempts against each other. If fatigue has also been building for a while, a flat trend across that same window can be a signal to check whether a deload is overdue; see deloads: when to take one and what should actually change if that sounds familiar.
What to Do Once You've Confirmed It's Real
Once a flat trend has actually cleared the noise bar across several weeks, it's time to go back to the diagnostic checklist, not before.
Start with the full diagnostic guide to training plateaus for the muscle-growth version of this question, or with workout plateau: seven causes to rule out before you change your program for a tighter, checklist version once you already know the stall is real.
This exact comparison, same lift, same rep range, similar RIR, tracked across weeks, is what a consistent training log is actually for. It's the basis of how myoxin's weekly review looks at a trend across your logged sets instead of reacting to any single session.
Conclusion
Noise is real, and it's usually bigger than the progress you're trying to see. Session-to-session swings of several percent are normal for a strength test, real week-to-week gains are usually smaller than that noise, and one bad day almost always has a boring, ordinary explanation. A real plateau needs weeks of pattern, not one rough session.
Before you rebuild anything, give the trend time to separate itself from the noise: 3-4 comparable data points across 6-8 weeks, on the same exercise, at a similar effort level. Consistent logging is the only real way to tell noise from signal, and it turns a guessing game into something you can actually check against your own numbers.
References
- Grgic J, Lazinica B, Schoenfeld BJ, Pedisic Z. Test-Retest Reliability of the One-Repetition Maximum (1RM) Strength Assessment: A Systematic Review. Sports Medicine - Open. 2020;6:31. pmc.ncbi.nlm.nih.gov/articles/PMC7367986 32 studies, 1,595 lifters: 1RM test-retest CV of 0.5-12.1% (median 4.2%), the core noise figure this article is built around.
- Hopkins WG. Measures of Reliability in Sports Medicine and Science. Sports Medicine. 2000;30(1):1-15. pubmed.ncbi.nlm.nih.gov/10907753 Establishes the "typical error" and "smallest worthwhile change" framework for interpreting repeated performance measurements.
- Latella C, Owen PJ, Davies T, Spathis J, Mallard A, Van Den Hoek D. Long-Term Adaptations in the Squat, Bench Press, and Deadlift: Assessing Strength Gain in Powerlifting Athletes. Medicine & Science in Sports & Exercise. 2022;54(5):841-850. pubmed.ncbi.nlm.nih.gov/35019902 Typical annual strength-gain rates by lift, showing real weekly progress is smaller than test noise.
- Borba VV, et al. Could a Habitual Sleep Restriction of One-Two Hours Be Detrimental to the Benefits of Resistance Training? Sleep Science. 2024. pmc.ncbi.nlm.nih.gov/articles/PMC11390164 Chronic mild sleep restriction did not blunt resistance-training gains, supporting "one rough night isn't a plateau."
Download