Showing posts with label Software Estimation. Show all posts
Showing posts with label Software Estimation. Show all posts

Thursday, May 22, 2014

The Mythical Story Point

I fairly recently became embroiled in an argument about whether Story Points or hours are better for estimating the effort associated with Software Engineering tasks. Having helped a lot of teams adopt Scrum and other agile practices, this is not the first time I have danced this dance, and my experience has left me with a strong preference for using hours rather than an abstraction.

The motivations for using Story Points (or any other abstraction for that matter, e.g. T-shirt Sizes) to estimate effort, seem very reasonable and arguments for their use are initially very compelling. Consistent high-accuracy software estimation is probably beyond current human cognitive capability, so anything that results in improvements, even small ones, constitutes a good thing.

The primary motivation for using Story Points is that they represent a unit of work that is invariant in the face of differing levels of technical skill, differing levels of familiarity with the code or domain, the time of day or year, or how much caffeine either the estimator or developer doing the work has ingested. They also provide a typically more course-grained unit of estimation than hours, which by necessity will result in more apparently accurate estimates. By combining this course-grained unit of work, with mandatory refactoring of Stories (or Epics, or Product Backlog Items, or whatever nomenclature you choose to use) larger than a particular effort size, a team is bound to improve the accuracy of their estimates.

The use of estimation abstractions also seem to be beneficial when a team follows the Principle of Transparency, which is espoused by most Agile philosophies. When a team follows this principle, they make the team’s velocity, estimates, actuals and other data available to all stakeholders (e.g. Sales, Marketing, Support and Management), who almost invariably care a great deal about the work items that they have a stake in, and particularly when those work items will be DONE. By using Story Points for estimates one initially avoids setting unrealistic expectations with stakeholders, who may not necessarily understand the prevalence of emergent complexity in the creation of software.

I would imagine that the human brain has a lot of deep circuitry designed exclusively to deal with time. It was clearly highly adaptive for an early hominid to be able to predict how far in the future a critical event might occur; whether that was knowing when the sun might go down and nocturnal predators appear, or when a particular migratory species would be in the neighbourhood. We are clearly genetically hard-wired for understanding course-grained time, e.g. circadian rhythms, synodic months, the 4 seasons, and the solar year. And human cultural evolution has yielded many powerful and ubiquitous time-related memes, which have added a deep and fine-grained temporal awareness to the human condition, measured in seconds, minutes and hours. Almost every modern electronic device’s default screen displays the time and date, including phones, microwaves, computers, thermostats etc. Time is so ubiquitously available in our modern digital lives that the site of an analog wall clock will require a Tweet or post to Instagram. And everyone has a calendar of some sort that they use to manage their futures. We have clearly become the Children of Time.

Unfortunately, being the Children of Time has obviously not made us capable of even vaguely accurate estimation of the time any task of significant complexity will take. However, we are also terrible at estimating pretty much everything else, so I suspect this is not indicative of a specific limitation of our time-related neural circuitry. 

It is also our aforementioned parentage that limits the usefulness of Story Points and similar abstractions for estimating effort in general.  After some, typically short, period of time everyone on the team and all the stakeholders unconsciously construct a model in their minds that maps the unit of the abstraction back to time in hours or days. And as soon as this  model has been constructed they ostensibly go back to thinking in hours or days, though they now require an extra cognitive step to apply the necessary transformation.

So why bother with using the abstraction in the first place?

I have experimented with the use of estimation abstractions with teams in the past and I can confidently say that using abstractions has proven to be a distraction in the long run. I have settled on an approach that uses course-grained time buckets for initial estimates, e.g. 1, 5, 10, 25, 50 and 100 hours. The key performance indicator for estimation should a be a steady improvement in estimation accuracy over time, rather than the absolute accuracy of the estimate.

Accuracy in the software estimation process is emergent and improvements are predicated on iteration, increased familiarity with the code and the domain, and visibility into the historical data (and analysis thereof). Showing team members how far their estimates have been off over time, just before they estimate other work, is a good way to prime them, and give them an appropriate anchor.

I suspect that I will dance this dance again in the future.

Monday, October 1, 2012

Emergent Complexity Management

In a previous post I asserted that accurate estimation of even moderately-complex software projects is impossible. This stems from our inability to predict the probability and impact of emergent complexity, due to our cognitive biases and the current human-centric method of generating software. Our inability to accurately predict and estimate is evident in almost every domain; the Great Recession being a perfect and painful example. Despite the fact that every decade or so a Black Swan Event occurs in the financial markets, Black-Scholes is still being used to predict the future price of options, and the markets still seem to be in denial about our inability to make accurate long-term predictions (even when we use sophisticated mathematical models). The difference between the markets and your average medium- to high-complexity software project is that in software projects you are almost guaranteed to have one, and typically multiple, Black Swan Events!

Despite the current significant failure rate of software projects, Professional Project Managers still believe that they can mitigate the risks associated with emergent complexity by building in a “buffer” or “contingency”. The data show unequivocally that this is, at best, pure self-delusion, and at worst criminal negligence; no matter how large this buffer is made, there is no way to ascertain upfront whether or not it is big enough to cover overruns that may arise due to emergent complexity. And obviously making it too large will cause project failure for purely financial reasons.

It is fairly obvious that the emergence of Agile Practices is an attempt to address the obvious shortcomings in software estimation specifically, and in our software development processes in general. Recursive refactoring of Epics into relatively-low-complexity User Stories and Tasks, and using techniques like Planning Poker, do make vaguely-accurate estimation possible, but I suspect that the fact that many Agile Gurus suggest abandoning estimation (in hours as opposed to an abstraction like Story Points) entirely, is an indication that this does not substantially improve the overall temporal predictability of Software projects.     

Selling Agile Practices to developers is like selling Gatorade to people who have been lost in the desert for days without a canteen. They get it instantly. The problem however is selling Agile to other “stakeholders”; whether they are internal Management, Sales, Product Management or Marketing teams, or a client for whom one is building software. This is where the real Agile battle is being fought.

The status quo is our, i.e. software engineers, own fault; for nearly half a century now we have been telling our customers that we know how long it is going to take, and implying that we have high confidence in our estimates. One has to wonder why, after we have been shown to be so utterly untrustworthy in this regard, that our customers and managers have continued to let us near computer keyboards.

So why do customers and managers keep buying the snake oil? Obviously they believe that they are limiting their exposure to the risks associated with emergent complexity (and sometimes incompetence unfortunately),  by passing that risk on to the development team or company responsible for the software development. Sure, if they make the penalty clauses punitive enough, and the developers agree to those clauses, it may partially mitigate the risks, but it does not address the fact that, in this day and age of rapidly-and-always changing competitive landscapes, a missed software release may very well mean the demise of an organization. So instead of mitigating the risk, it is simply being hidden and made binary.

That said, and despite what the Gurus say, it is not practical to simply leave software project costs and schedules open-ended. We need a way to roughly size software projects, but we also need a way to manage emergent complexity and minimize it’s negative impact  to the schedule and ultimately to the value of the software to the user.

I am obviously one of a multitude of software professionals grappling with this, and I have yet to find a solution that is totally satisfying (and sellable). I suspect that the answer includes motivating stakeholders to abandon the illusion of certainty for real transparency and a daily opportunity to inflect the engineering process; and a more scientific approach to product management along the lines of The Lean Startup.

Friday, January 20, 2012

The Emperor Will Never Have Clothes!

Half a decade ago I did some fairly exhaustive research into software time and cost estimation techniques for Microsoft Services. I evaluated the effectiveness of both formal and informal techniques used within Microsoft, across the software development industry, and even in other industries. I looked at everything from Estimation by Analogy to COCOMO II.

After months of research I discovered that the most commonly used estimation techniques were Estimation by Analogy and Wideband Delphi, often done in combination. However, most of the teams using these techniques were doing so informally, and had never heard of either of these terms.

I also discovered that there were no software estimation techniques that consistently produced accurate estimates, other than for small projects, even those sophisticated parametric estimation techniques like COCOMO II. 

At the time I was doing the research the failure rate for IT projects was around 70%, and many of those failures were due to catastrophic budget and schedule overruns. It quickly became evident to me that the way the industry had historically thought about estimating software projects was deeply broken; The Emperor had no clothes! 

My final recommendation to Microsoft Services was that they adopt Estimation by Analogy and Wideband Delphi, and formally train their staff on the use and limitations of these techniques. My other recommendation was that they also invest in maturing their risk management processes, given the historical inaccuracy of software estimation in general. It was on this latter area that I focused until I left Microsoft.

Our obviously apparent inability to do accurate estimation bugged me for years after I completed this research. This itch that I could not seem to scratch motivated me to do some reading about predication in general, and in particular Quantitative Analysis and its use of Stochastic models; one would imagine that if any industry had worked out how to predict the future it would be the Financial Industry given their huge monetary incentive to do so. Obviously recent events have shown that not to be the case, but this was before all that nastiness.

And then I read “The Black Swan: The Impact of the Highly Improbable” by Nassim Nicholas Taleb, and it all made sense. It is not that we have yet to find an accurate technique to do software estimation; it is that accurate estimation of large, complex software development projects is simply NOT POSSIBLE! (given our current understanding of the laws of physics anyway). Not only does the Emperor not have clothes, but he will remain naked for the foreseeable future.

I suspect that the move to  Agile software development practices is a natural response to our inability to do accurate estimation, but I am continually amazed at how many people in the industry and how many of its customers are still in denial about this, what should now be self-evident, fact. I still encounter many projects where development teams are held to early estimates, which are typically done by non-technical project managers. And this is particularly true on fixed-cost\budget projects.

Surely it is time for the entire industry and its customers to acknowledge that this whole approach is simply broken? The illusion is that fixed-cost\budget projects shift ownership of the risk from the customer to the organization who is developing the software; but given the percentage of IT projects that still fail today, this is obviously just that: an illusion; and a pernicious one at that.

We need to start from the premise that accurate software estimation is currently, and probably forever, impossible, and go from there. Agile practices have made a good start but we obviously need to do more. And most importantly we need to educate our customers!

Here is an interview with Nassim Nicholas Taleb wherein he talks specifically about our inability to estimate IT projects:

Tuesday, December 20, 2011

Why Software Development is like Ironing a Thneed

 I’m being quite useful. This thing is a Thneed.
A Thneed's a Fine-Something-That-All-People-Need!
It's a shirt. It's a sock. It's a glove. It's a hat.
But it has OTHER uses. Yes, far beyond that.
You can use it for carpets. For pillows! For sheets!
Or curtains! Or covers for bicycle seats!
"

                                         – from “The Lorax” by Dr. Seuss

During my time in the military two decades ago I become highly skilled at ironing. And not only shirts and pants, but beds, hats, sheets and other items one would not normally consider “ironable”. I continue to this day to do my own ironing. I actually find it rather therapeutic.

Since having children I have also become reacquainted with the works of Theodor Geisel, more affectionately known as Dr. Seuss. It struck me the other day while ironing a particularly pesky shirt that software development is very much like ironing a Thneed.

So let’s make some assumptions about a Thneed based on the description above. Some poetic licence and imagination will be required.

  1. Nobody is entirely sure what a Thneed is, not even the manufacturer.
  2. When customers buy a Thneed they have only a vague idea of what they need it for and how it is going to make their lives better.
  3. No two Thneed’s are exactly alike; they are the snowflakes of garments.
  4. Thneed producers create new and improved Thneed’s all the time.
  5. A Thneed is too big and awkwardly shaped to fit on your ironing board and Dr. Seuss makes no mention of a Thneed-press.
  6. It is hard, if not impossible, to estimate how long it will take you or a highly trained team of Thneed-ironers to iron a Thneed, with any level of accuracy or confidence in your estimate.
  7. The market for Thneed-ironing accessories is confusing in its profusion, pace and super-competitiveness.
  8. Thneed-ironing methodologies were derived from shirt and pants ironing methodologies, but they are actually poorly suited.
  9. A Thneed is like most other garments in that it occasionally requires ironing.
  10. Thneed manufactures’ own Thneed’s are usually the worst ironed.

So based on the assumptions above, how does one go about ironing this Thneed thing? Well the trick is known by every person who has ever had to iron a shirt.

Manipulate the garment so that a small piece of it is flat on the ironing board and then iron that piece. Then get another piece flat and iron that piece, and so on and so on, until you have ironed the entire garment. If you are foolish enough to try to iron large sections of the garment at the beginning, you will become very frustrated and will probably run out of steam before you are done. You have to “divide and conquer” when it comes to ironing a Thneed; that is the winning strategy.

I don’t think that I need to actually spell out why this is like software development; if you have done any software development in your life you will know that I am right (or mostly so).

Maybe someday I will post “Why Architects are like the Lorax, and Users like the Once-ler”. On second thoughts maybe I won’t.

Happy Ironing!

Friday, September 16, 2011

Good Software Design is Honest

Ten Principles of Good Design Redux | Part 6

The Software Industry has been lying to itself and its customers since its emergence. At some point in the early evolution of the Software Industry someone must have noticed that the emerging development lifecycle model, originally proposed described by Winston Royce and now known as the “Waterfall” or “Big Design Up Front” model, was better suited to Aeronautical Engineering than it was to software development. Designing and building software is not the same as designing and building aircraft.

[Update (2011-10-07): Winston Royce did not propose Waterfall; he was the first to formally describe it, and used it as an example of how not to do software development. He was the “someone” I was referring to in the previous paragraph. Though this error makes my assertion that Aeronautical Engineering is better suited to the Waterfall Model seem a little arbitrary, it does not invalidate the point of the post. Mea culpa for poor fact checking.]

I am not so naïve that I would suggest that they are different merely because of the complexity of the problem domain or because of the amount of uncertainty innate to the process. Though I can only imagine that having the Laws of Physics as a significant source of constraints must reduce the uncertainty somewhat in the design and implementation of an aircraft, I would assert that they are most different in the degrees to which they are subject to the more capricious aspects of Human Nature.

When an aircraft manufacturer designs a new aircraft they typically know precisely what the majority of the requirements are from the outset, e.g. carry this many Wi-Fi-connected, grumpy adults and screaming infants from this point on the globe to this other point, using this much fuel, and oh, don’t crash. The tolerances and constraints are well understood; the laws of physics are essentially immutable, and you can only “comfortably” cram so many humans into a given space (though as someone who is over six foot and flies regularly I have to say that I am not so convinced of my last point).

The aircraft manufacturer can plug all these requirements and constraints into a model and calculate whether or not, given existing and emerging technology, they can feasibly build the aircraft, or if and where the invention of new technology is going to be required. Much of the problem domain is known and is relatively constant, i.e. the Laws of Physics and the current state of the Material Sciences. They also know what they don’t know, and they can mostly quantify the risk\cost associated with that uncertainty thanks to good historical data. A lifecycle model that is dominated by a discrete design phase is clearly effective in bringing a new aircraft to market, though obviously prototyping and innovation are also required during this initial phase. The cost associated with this protracted design phase is accepted as necessary by the, now mature, aviation industry, probably because the cost of failure is so high.

So why is Software development different? One of the most significant differences between Aeronautical and Software Engineering is that users’ of software products and systems are very rarely able to articulate detailed requirements at the outset. The constraints and requirements have to be extracted out of the minds of the intended users, in the best case, or out of the mind of one or more analysts who thinks they understand the users’ requirements, in the worst. And though the Laws of Physics are at play in the hardware that the software runs on, they almost never have to be considered as constraints in the design of the software. Typically the requirements gathering process is a bootstrapping exercise that continues well into the actual development of the system. Users have to have usable examples of what they don’t want to lead them to understand what they actually want, and the engineers have to attempt to solve some of the known hard technical problems in order to reveal the initially-unknown, usually harder, ones.

And it is impossible to estimate how long it is going to take to extract the real requirements and then develop the software to meet those requirements, at the beginning of the project. That is not to say that a waterfall model, that included exhaustive prototyping during the design phase, would not work for software, but it would require that everyone acknowledge that the requirements gathering phase would need to be completed before a fixed-cost time and effort estimate of the development could be provided, and that the duration of that initial phase could only be roughly estimated. There are simply more unknown unknowns in Software Development than in classical Aeronautical Engineering. I say “classical” because modern aircraft probably require as much Software Engineering as they do Aeronautical Engineering.

Another difference is that because so much of the software that powered the growth of the industry was developed by under- or un-paid geeks, it has created a mass hallucination about how much [paid for] time and effort it actually takes to develop high-quality software. Everyone has now come to accept this hallucination despite the number of software projects it has caused to run over time and budget, or to fail completely.

But thankfully all of this is changing; Agile software development practices are an attempt to address this innate dishonesty, and acknowledge that we as software developers simply don’t know what we don’t know. And though Agile has become mainstream, there are still those who doggedly cling to the old dishonest and delusional ways.

It is every software architect’s and developer’s responsibility to promote and champion Agile as an honest approach to the design and development of great software, particularly in the face of grumblings from old-school project managers and customers who want the illusion of fixed risk.

Note: I could write an entire book on the topic above, and was well on the way to doing so before I reminded myself that this was just a blog post destined to be read by 10s of my friends. I am sure though that the above makes the point I was trying to make.