Blog Autoregulation

How to Track RIR Without Overthinking Every Set

8 min read

Quick answer

Track RIR (reps in reserve) with a deliberately cheap protocol: the moment you rack a working set, log one whole number from 0 to 5 for how many clean reps you had left, and move on. Skip warm-ups, timed holds, and stretches. Do not chase half-rep precision; your rating carries an error of a rep or two even near failure, so whole numbers already capture everything the measurement can honestly say. Log every working set rather than just the hard ones, and re-anchor the scale every few weeks with one set taken to technical failure on a machine or cable. The value is in the trend across sets and sessions.

Reps in reserve has a failure mode of its own: taking it too seriously. You rack a set, stare at the ceiling, and litigate whether that was a 2 or a 3 while your rest timer runs. Multiply by twenty sets and effort tracking has cost more effort than the training. The fix is a protocol built around what the measurement can deliver.

What you are actually estimating

RIR counts what you left behind: how many more good reps were available at the moment you stopped. End a set of 8 with two still in you, and that set goes in the log as 2 RIR. It matters because proximity to failure is the best practical proxy for how stimulating and how fatiguing a set was, and because a log of effort over time is what turns your training history into something that can steer decisions. For the scale end to end, and its RPE equivalents, see RIR explained; this post is about the habit of logging it.

How accurate your ratings really are (and why that is fine)

The numbers first. In the research where lifters call their RIR and then rep out to actual failure, the average miss is about 1-2 reps even when they call a set at 1 RIR, almost always with more left in the tank than called, and it grows to roughly 5 reps for sets called at 5 RIR. Accuracy also worsens in longer sets, and newer lifters tend to rate soft, believing they are closer to failure than they are, though that experience effect is the less consistent finding. The pattern to internalize: ratings are usable near failure and blurry far from it.

This is why chasing precision is the wrong game. A rating with a built-in error of a rep or two cannot support half-rep bookkeeping; logging "2.5 RIR" records confidence the measurement does not have. Whole numbers, 0 through 5 with 5 meaning "5 or more", are the honest resolution, and they are all a trend needs. Individual ratings are noisy; averages across a session and across weeks are surprisingly stable, and every useful decision RIR feeds runs off those.

The three-second habit

Rate the set the moment you rack it, before you sit down. The feeling of proximity to failure decays fast, and the read you have at three seconds is better than the one you can reconstruct at ninety. One whole number, logged, done. If you are torn between two numbers, log your gut read and stop litigating: the single set is noise either way, and the trend it feeds washes it out. What matters is applying the same internal ruler every time, because a consistent ruler makes your trend comparable week to week even if your absolute calibration is off by a rep.

Re-anchor the scale every few weeks

Ratings drift, especially when you never visit the end of the scale. The fix costs one set: every 3 weeks or so, take a single set to technical failure, the point where the next clean rep will not come, and compare where it ended against what you would have called it. Run it on a machine or cable movement, where failing is safe and the rep count is unambiguous, and rotate it across your main movement patterns (a press, a pull, a squat or leg press, a hinge pattern) rather than repeating one exercise. Keep these anchors off heavy barbell compounds, and keep them occasional: every set to failure is maximal fatigue for one data point, and a weekly habit of it quietly becomes its own fatigue problem.

Rate every working set, even the easy ones

The tempting shortcut is rating only the top sets, the ones with a story. The problem is statistical: any average computed from your log now runs on a sample of mostly-hard sets, so it reads grimmer than your training actually was. This bites hardest in tools that act on the average. Anneal, for example, reads average logged RIR across recent sessions against a grinding threshold as a deload signal; feed it only your hardest sets and a perfectly normal week can look like chronic grinding. Selective logging turns a good detector into a nag. Rating every working set costs three seconds each and keeps every downstream read honest.

Warm-ups stay out of the log for the mirror-image reason: they are submaximal by design, so they would all read 5+ and only dilute the signal. Timed holds and stretches stay out too; there is no discrete rep to count back from.

Which sets actually count toward growth

Once the log exists, one question follows fast: are the easier sets doing anything? The current evidence reads as a gradient rather than a cliff. Hypertrophy stimulus per set rises as you approach failure, with the steepest returns in the last couple of reps; sets left at 4-5 RIR still contribute, just somewhat less per set, and added volume can partly make up the difference. Strength outcomes stay nearly flat across a wide RIR range, which is why heavy work at 1-3 RIR builds strength while sparing fatigue. Most productive hypertrophy programs concentrate their working sets at roughly 0-3 RIR.

One popular way to picture the gradient is the "effective reps" model: the idea that roughly the last five reps before failure are the strongly stimulating ones. Hold it loosely. It is a modeling convenience that makes the gradient easy to reason about, not an established finding. The calculator below applies it to your own numbers so you can see how set effort and weekly volume interact.

If you want the inverse direction, turning a target RIR into a concrete working weight, the RIR-to-load calculator does that from a recent hard set.

What the log unlocks

A consistent RIR column turns a workout log into an early-warning system. Effort rising while loads stay flat is accumulated fatigue announcing itself weeks before a stall shows up in the weights; the decision rules for acting on that are in how to know when to deload, and the broader practice of steering load and volume off the readings is autoregulation.

In Anneal the column works without ceremony: tapping an RIR value logs the set in one gesture, and the number feeds everything downstream. Prefills adjust the next session's weight using your last reps and RIR together, a session of near-failure grinding or a multi-week slide in average RIR each count among the deload signals, and the effort distribution shows up in analytics so you can see a block getting heavy before you feel it clearly.

Common questions

Should I log RIR on warm-up sets?

No. Warm-ups are intentionally submaximal, so there is no meaningful proximity to failure to rate; a warm-up rated honestly would just read 5 or more every time. The same goes for timed holds and stretches, where there is no discrete rep-failure point to count from. Rate working sets only. That keeps the log clean and keeps the habit cheap enough to sustain.

What if I genuinely cannot tell my RIR?

Log your gut read as a whole number and move on; a single set's rating barely matters, because the useful signal is the trend across sets and sessions. Ratings are genuinely hard to make far from failure, where the error can reach several reps, and much easier near it. Two things sharpen the read: training experience, and occasionally taking a set to technical failure on a machine or cable so you re-learn what zero actually feels like.

Is half-rep RIR precision worth logging?

No. Even close to failure your rating carries an error of roughly one to two reps, and it grows the further from failure you are. Half-rep precision sits below that noise floor, so it records confidence the measurement does not have. Whole numbers from 0 to 5, where 5 means 5 or more, capture everything the data can honestly say.

Does tracking RIR matter if I train close to failure anyway?

Yes, arguably more. If your sets genuinely live at 0-1 RIR, the log confirms it and shows how long you can sustain it, because a string of sessions grinding at the bottom of the scale is the classic signature of accumulated fatigue and a leading indicator that a deload is due. And if your "to failure" sets turn out to log at 2-3 RIR once you calibrate, that is worth knowing too.

Posted by the founder of Anneal, a workout tracker that knows when to deload.