timesheet
An estimate is almost always requested at the worst possible moment. The work has not been broken down, the team is not confirmed, and nobody knows yet which pieces block which. A number gets said out loud anyway, and from that point on it is treated as a promise. Most of the pain that follows has nothing to do with the arithmetic. It comes from the fact that a rough figure and a committed schedule were never separated, and there was no mechanism to keep them apart.
An estimate is a statement about uncertainty. A commitment is a statement about accountability. They travel together in conversation and get flattened into the same cell of the same spreadsheet, which is where the damage starts.
Consider the usual sequence. Someone asks how long a piece of work will take. The answer is "roughly three weeks". By the following week that has become "delivery in three weeks" in a status update, and by the week after it is a date on a plan that other teams have built dependencies around. Nothing dishonest happened at any step. The qualifier simply had nowhere to live. Spreadsheets and slide decks have a cell for the number and no cell for the confidence attached to it.
The practical fix is to force the two to be recorded differently. An estimate belongs to the task and should carry a range and the assumptions behind it. A commitment belongs to the schedule and should carry a date, an owner, and the buffer that was deliberately added. When both live on the same card in the same system, the qualifier survives the trip. When the estimate lives in a spreadsheet and the schedule lives somewhere else, the qualifier is lost in the copy.
This also changes who can be held to what. A team that gave a range of eight to fourteen days and was scheduled at nine has been set up to miss. That is a planning decision, not an estimating failure, and it is only visible if the range was written down in the first place.
Estimating literature lists many techniques. In small and mid sized teams, four do nearly all of the work, and the choice between them is mostly a question of what information already exists.
Find a past piece of work that resembles this one and start from what it actually took. Fast, requires almost no breakdown, and only as good as the record of the past project. Teams that never captured actual effort cannot use this method at all, which is the main reason it is dismissed as unreliable.
Build a rate and multiply. Hours per page, days per screen, cost per endpoint. This works when the work is genuinely repetitive and the unit is well defined. It fails quietly when the unit hides variation, for example when "one screen" covers both a static page and a form with fourteen validation rules.
Break the work down, estimate each leaf, and sum. The most accurate of the four when the breakdown is complete, and the most expensive to produce. Its real weakness is that the summing step hides optimism: fifty tasks each shaved by ten percent produce a total that is wrong by ten percent and looks meticulous.
Give an optimistic, a most likely, and a pessimistic figure for each item, then combine them. The common weighting is the PERT formula, which is the optimistic figure plus four times the most likely plus the pessimistic, divided by six. The output is a number with a visible spread, which is exactly the property the other three methods lack.
| Method | What it needs | Time to produce | Best use |
|---|---|---|---|
| Analogous | A record of past actuals | Hours | Early sizing, go or no go decisions |
| Parametric | A stable unit and a known rate | Hours | Repetitive, well understood work |
| Bottom-up | A full task breakdown | Days | Committed schedules and budgets |
| Three-point | A breakdown plus a view on risk | Days | Anything with a hard external date |
In practice these are used in sequence rather than chosen once. An analogous figure answers whether the work is worth scoping. A bottom-up pass produces the figure that goes into the schedule. Three-point is applied to the handful of items that are genuinely uncertain, not to all of them, because asking a team for three numbers on two hundred tasks produces three hundred tasks worth of guesses.
An estimate of effort is not a date. Turning one into the other requires three pieces of information that are frequently skipped.
The first is dependency order. A hundred hours of work that must happen in strict sequence and a hundred hours that can run in parallel produce completely different end dates. This is where a Gantt or timeline view earns its place, not as a presentation artifact but as the only readable way to see whether the chain of blockers is longer than the calendar allows.
The second is available capacity, which is never full time. A person assigned to a project is also in meetings, supporting last quarter's release, and taking leave. Teams that plan against a five day week and then measure against reality tend to find that four days of project work per person per week is closer to the truth. Whatever the real figure is for a given team, it should be measured rather than assumed, and the same figure should be used for every plan so that comparisons between plans mean something.
The third is the calendar itself. Public holidays, company shutdowns, and the fact that a two week task starting on a Thursday does not finish on a Thursday. This is arithmetic, and it is also the most common source of a plan that is wrong by a week before any work starts.
Once those three are applied, the range converts into a range of dates. Publishing the earlier date as the commitment and holding the later one privately is the standard way this goes wrong. Committing to the later date and delivering early is unpopular in the moment and is the only version that survives contact with a second project.
Every method above except pure guesswork depends on knowing what past work actually took. That record is the asset, and most teams do not have it.
The reason is rarely disagreement about its value. It is that capturing actual effort requires people to enter something, and every additional place to enter it reduces the number of people who do. A separate timesheet tool that nobody opens produces no data. A spreadsheet filled in retroactively on the last day of the month produces data that is a reconstruction rather than a record.
What tends to work is keeping the recording next to the work. If the task already lives on a board, and the estimate is a field on that task, then the actual figure belongs in the next field along. The comparison then costs nothing to produce, and after two or three projects a team has its own parametric rates instead of borrowing industry averages that were measured on a different kind of work.
Two numbers are worth the effort of collecting. The first is the ratio of actual to estimated effort, by person and by type of work, which tells a team whether it is optimistic by twenty percent or by two hundred. The second is the count of items that were not in the original breakdown at all, which is usually the larger error and is invisible if the breakdown is edited in place rather than kept as a baseline.
A small number of patterns account for most estimating disasters, and all of them are recognisable early.
Padding that is hidden rather than declared is the first. Every estimate gets a private margin, and margins compound up the hierarchy until the total is unusable and nobody can say which part is real. Declared buffer at the project level is defensible. Silent buffer at every leaf is not.
Estimating by desired outcome is the second. The date is set, the estimate is fitted to it, and the fitting is done by trimming the tasks that are hardest to defend, which are typically testing, review, and handover. The plan then looks achievable and the last third of it is fiction.
Estimating without the person who will do the work is the third. A figure produced by a lead on behalf of an engineer is a forecast about someone else's speed. It can be right, and it removes the one person who would have noticed the missing dependency.
The fourth is treating the first estimate as final. An estimate produced before discovery should be replaced after discovery. Teams that never re-estimate carry their least informed number all the way to delivery and then discuss why it was wrong.
The mechanical cause of most of the above is that the estimate, the schedule, and the record of actual effort live in three different places. The breakdown happens in a spreadsheet, the schedule is redrawn in a chart tool for the steering meeting, and the actual hours are either in a separate timesheet or nowhere.
Three artifacts in three places have to be reconciled by hand, so they are reconciled once a week at best, and the version people look at is the one that was easiest to open. Comparing this quarter's estimates against last quarter's actuals becomes a project in itself, so it does not happen, and the team estimates the next piece of work from memory.
Keeping them on one surface removes the reconciliation step rather than making it faster. An estimate field and an actual field on the same task, a timeline view generated from those same dates rather than redrawn, and hours recorded against the task they belong to. Tools differ in how much of this they include before a paid tier; comparing what is available in each free plan is a more useful exercise than comparing feature counts, because the tier boundary is what decides whether the whole team participates or only the people with licences. The feature list is where to check whether timeline views and custom fields are gated.
Pick the last three completed pieces of work and write down what each one was estimated at and what it actually took. That single table will tell you more about your estimating than any method in this article. If the actual figures are not recoverable, the first change is to start recording effort against tasks rather than in a separate place, which is what a board with estimate and actual fields plus Pinateca timesheets is for.
Accuracy expectations should scale with how much is known. An early analogous estimate that lands within fifty percent of actual is doing its job, which is to decide whether to proceed. A bottom-up estimate produced after discovery is usually held to plus or minus ten to twenty percent. Quoting a single precise number before discovery implies an accuracy that does not exist.
Hours are necessary when the output is a budget or an invoice, because those are denominated in money and money is denominated in time. Points work when the only question is relative size and the team has enough history to convert points to a delivery rate. Running both adds a conversion step and rarely adds information, so pick the one that matches what the estimate is used for.
The people who will do the work, with a reviewer who has seen similar projects finish. The doers catch missing tasks and technical dependencies. The reviewer catches the optimism that comes from estimating only the happy path, and the items that are always forgotten, such as review cycles, environment setup, and handover.
Estimate the discovery work instead, with a fixed time box, and commit only to that. State plainly that the figure for the delivery phase will be produced at the end of discovery. This is more defensible than a wide range on undefined scope, because a wide range is read as its lower bound by whoever is planning around it.