task-ops

Fibonacci Story Point Scales: Why the Gaps Get Wider on Purpose

October 3, 2026 ・ Pinateca Editorial

A team that has been estimating in fibonacci story points for a few months usually arrives at the same two questions. Why can a story be a 5 or an 8 but never a 6, and what is the team supposed to do when half the room says 3 and half says 13. Both questions have the same answer, and it has nothing to do with the mathematical properties of the sequence.

The scale is a constraint, not a measurement system. Understanding what it is constraining is what makes the practice useful rather than ceremonial.

The scale, and the version most teams actually use

The sequence itself is familiar: each number is the sum of the two before it, giving 1, 2, 3, 5, 8, 13, 21, 34 and onward. For estimation, teams typically stop somewhere around 13 or 21 and treat anything larger as a signal to split the work.

Most planning poker decks do not use the pure sequence. They use a modified version that rounds the upper end, so that after 13 the values become 20, 40 and 100 rather than 21, 34 and 55. Many decks also add a 0 for work already effectively done, a half point for something trivial, and a question mark for a story nobody understands well enough to size. The rounding is deliberate. At the top of the scale the difference between 21 and 20 carries no information, and round numbers are easier to talk about.

What matters is the ratio between adjacent values, not the values themselves. From 5 to 8 is a jump of about sixty percent. From 8 to 13 is about sixty percent again. The sequence produces a roughly constant proportional step, which converges toward the golden ratio as the numbers grow. A fibonacci scale is a geometric scale wearing a familiar costume.

That is why substituting a different geometric progression, such as 1, 2, 4, 8, 16, changes very little in practice. Teams that switch usually report the conversations feel the same. The important property is that the steps grow, and any sequence with that property will do the same job.

Why the gaps widen on purpose

The widening gaps exist to stop the team from claiming precision it does not have.

A small, well understood piece of work can be judged closely. The difference between a one point story and a two point story is visible, because both are concrete enough to picture in full. A large piece of work is different. Once something is a multi week effort with unknowns in it, the honest range is wide. Saying it is a 13 rather than a 12 or a 14 is not modesty, it is accuracy. The scale removes options that would only produce fake distinctions.

The rationale most often cited for this comes from the study of perception, specifically the observation associated with Weber and Fechner that people judge differences in proportion to the size of what they are comparing. The gap between one kilogram and two is obvious by hand. The gap between twenty kilograms and twenty-one is not. Estimating work behaves the same way, so the scale is built to match how the estimators actually perceive difference.

There is a second effect, and it may be the more valuable one. Because the gaps are large, the act of choosing between two adjacent values forces a conversation about scope. When one person says 5 and another says 8, they cannot split the difference into 6 and move on. They have to say out loud what they think is in the story. Most of the time the disagreement turns out to be about the work itself: one person is including data migration and the other is not. The estimate is a side effect. The shared understanding is the product.

A wide gap also gives the team somewhere to put uncertainty. If a story might be simple or might be a nightmare depending on what is found in the existing code, the larger value is the correct one, and the size itself becomes a flag that something needs investigation first.

What a story point is not

Several habits quietly destroy the value of the practice, and each one comes from treating points as a unit of measurement.

A story point is not an hour. Publishing a conversion rate, such as one point equals four hours, removes every advantage of relative sizing. The team is now estimating in hours with extra steps, and every estimate becomes a commitment that can be held against them. The moment a conversion exists, points inherit all the problems that relative sizing was introduced to avoid.

A story point is not comparable across teams. Points are calibrated against one team's own reference stories and its own working context. Two teams can both average a similar total per iteration while sizing the same work very differently. Using points to compare teams is the fastest way to teach both of them to inflate.

A story point is not a measure of individual speed. The number describes the work, not who picks it up. If a story gets larger when a particular person is likely to take it, the scale is being used as a performance review.

A story point is not a promise. It feeds a forecast, and forecasts have ranges. A team whose recent iterations delivered totals spread across a wide band should quote a range, not a single figure.

Running the estimate without it taking an hour

The mechanics are simple and the failure modes are predictable.

Start with reference stories. Pick two or three pieces of work the team finished recently and agree on their sizes, for example a small one at 2 and a mid-sized one at 5. Every later estimate is made by comparison to those. Without references, the scale drifts within weeks and old numbers stop meaning anything.

Estimate simultaneously, not sequentially. Everyone reveals a value at the same time, whether with cards, fingers or a chat message sent on a count of three. The reason is anchoring: if the most senior person speaks first, the rest of the numbers cluster around that figure and the exercise produces one person's opinion with a quorum attached.

Talk only about spread. When the values are close, take the higher one and move on. Discussion is only worth the time when the spread is large, and then the goal is to surface what the outliers are each including. Two rounds is usually enough. A third round rarely changes the number and always costs ten minutes.

Split anything above the team's ceiling. Pick a value, commonly 13 or 20, and treat it as a rule: stories larger than this do not enter an iteration, they get broken down first. This single rule prevents most of the forecasting problems teams blame on the scale.

Decide what counts as a split. The most common way to break a large story down is by scenario rather than by layer. Splitting into a database task, a backend task and a frontend task produces three items that cannot be finished independently, so the total still lands at the end of the iteration. Splitting into the simple case first and the edge cases second produces two items that each deliver something. The scale rewards this, because a story split into independent slices tends to size lower in total than the original monolith.

Timebox the session. A backlog refinement meeting that runs long produces worse estimates at the end than at the start, because attention is gone. If the team cannot size an item in a couple of minutes, that is information: the item needs a spike, not a number.

Fibonacci, t-shirts, and the alternatives

The choice between sizing approaches is narrower than it looks. The comparison below covers what the differences actually produce.

Approach What it produces Best suited to
Fibonacci points A number that can be summed for forecasting Teams tracking a total per iteration
Modified fibonacci with 20, 40, 100 The same, with a coarser top end Backlogs that contain large unrefined items
T-shirt sizes A shared sense of scale, no arithmetic Early roadmap work, stakeholder conversations
Powers of two A geometric scale, simpler to explain Teams who find the sequence itself distracting
Ideal days An estimate in time, easily misread as a commitment Work with firm external deadlines
Right-sizing without points A yes or no on whether an item fits an iteration Teams with steady flow who forecast by count

T-shirt sizes are worth a closer look, since they solve a real problem. When the audience for the estimate is outside the team, small, medium and large communicate scale without implying a schedule. The cost is that letters do not add up, so any forecast requires converting them to numbers, at which point the team has fibonacci points with an extra translation step.

Right-sizing deserves the same honesty. Teams with a stable flow of similarly sized items often find that counting items forecasts as well as summing points, and it costs no meeting time at all. This works when the items really are similar. When one item in four is five times the size of the others, counting breaks and the geometric scale earns its place.

Whatever scale is chosen, the numbers have to live where the work lives. Points recorded in a separate spreadsheet stop matching the board within two iterations, because a story gets split and only one of the two places gets updated. Keeping the estimate on the card itself, as a field alongside the owner and the status, is what makes the total trustworthy at the end of the iteration. Most board tools support that with a custom field, and it is worth confirming on the features list of any candidate rather than assuming it.

Reading the numbers afterwards

The point of estimating is the forecast, and the forecast is only as good as the record kept after the fact.

Track the completed total per iteration and look at the last several, not the last one. A single iteration total is noise. A band across recent iterations is a forecast, and the width of that band is the honest uncertainty in any date the team gives.

Do not re-estimate finished work to match what it actually took. It is tempting, and it destroys the data. Points describe the size judged before starting, and the gap between that judgment and reality is exactly the information the forecast needs.

Watch for inflation. If the average total climbs steadily while output feels unchanged, the scale has drifted and the reference stories need to be re-agreed. This happens naturally when new people join and calibrate against recent numbers rather than the original references.

Finally, keep the history visible. A team that can see which stories were sized 8 and took three days, and which were sized 8 and took three weeks, gets better at sizing without any process change. Teams moving between tools often lose that history, which is worth checking against the import path before a migration.

What to change first

Pick a ceiling value and enforce it for one iteration: nothing larger than a 13 enters the sprint without being split. That one rule fixes more forecasting pain than any change to the scale. If the estimates then live in a spreadsheet separate from the board, move them onto the cards themselves and see whether the totals start matching reality, starting from Pinateca.

Q1. Why is there no 6 or 7 on the scale?

Because at that size the team cannot reliably tell the difference between a 6 and a 7, and offering both invites a pointless debate. The gaps widen so that every available value represents a genuinely different size. When two people disagree between 5 and 8, the absence of a middle option forces them to say what scope they are each assuming, which is the useful part.

Q2. How many story points equal one hour?

None, and publishing a conversion rate removes the reason for using points at all. A point is a relative size compared to the team's own reference stories, calibrated by that team, in its own context. If hours are what stakeholders need, forecast from the recent range of completed totals per iteration instead of converting individual estimates.

Q3. What should a team do when estimates are split between 3 and 13?

Treat the spread as the signal. Ask the lowest and highest voters to describe what they believe the story includes, because a gap that large is nearly always a difference in assumed scope rather than a difference in skill assessment. After one round of discussion, re-vote. If the spread stays wide, the story is not understood well enough to size and needs investigation first.

Q4. Can story points be compared between two teams?

No, in any way that supports a decision. Points are calibrated against each team's own reference stories, so identical work can be sized differently by two teams that are both estimating honestly. Using points to compare output teaches both teams to inflate their numbers, which destroys the forecasting value for everyone.

Q5. Is it worth estimating at all if the team has steady flow?

Possibly not. Teams whose backlog items are similar in size often forecast just as well by counting items completed per iteration, which costs no meeting time. The test is variance: if one item in four is several times larger than the rest, counting breaks down and a geometric scale like fibonacci earns its keep.

Back to the blog