Showing posts with label technical debt. Show all posts
Showing posts with label technical debt. Show all posts

Monday, August 08, 2016

Single Point of Failure (Airlines)

Large business-critical systems can be brought down by power failure. Who knew?

In July 2016, Southwest Airlines suffered a major disruption to service, which lasted several days. It blamed the failure on "lingering disruptions following performance issues across multiple technology systems", apparently triggered by a power outage.
In August 2016 it was Delta's turn.

Then there were major problems at British Airways (Sept 2016) and United (Oct 2016).



The concept of "single point of failure" is widely known and understood. And the airline industry is rightly obsessed by safety. They wouldn't fly a plane without backup power for all systems. So what idiot runs a whole company without backup power?

We might speculate what degree of complacency or technical debt can account for this pattern of adverse incidents. I haven't worked with any of these organizations myself. However, my guess is that some people within the organization were aware of the vulnerability, but this awareness didn't somehow didn't penetrate the management hierarchy. (In terms of orgintelligence, a short-sighted board of directors becomes the single point of failure!) I'm also guessing it's not quite as simple and straightforward as the press reports and public statements imply, but that's no excuse. Management is paid (among other things) to manage complexity. (Hopefully with the help of system architects.)

If you are the boss of one of the many airlines not mentioned in this post, you might want to schedule a conversation with a system architect. Just a suggestion.


American Airlines Gradually Restores Service After Yesterday's Power Outage (PR Newswire, 15 August 2003)

British Airways computer outage causes flight delays (Guardian, 6 Sept 2016)

Delta: ‘Large-scale cancellations’ after crippling power outage (CNN Wire, 8 August 2016)

Gatwick Airport Christmas Eve chaos a 'wake-up call' (BBC News, 11 April 2014)

Simon Calder, Dozens of flights worldwide delayed by computer systems meltdown (Independent, 14 October 2016)

Jon Cox, Ask the Captain: Do vital functions on planes have backup power? (USA Today, 6 May 2013)

Jad Mouawad, American Airlines Resumes Flights After a Computer Problem (New York Times, 16 April 2013)

 Marni Pyke, Southwest Airlines apologizes for delays as it rebounds from outage (Daily Herald, 20 July 2016)

Alexandra Zaslow, Outdated Technology Likely Culprit in Southwest Airlines Outage (NBC News, Oct 12 2015)


Related post Single Point of Failure (Comms) (September 2016), The Cruel World of Paper (September 2016), When the Single Version of Truth Kills People (April 2019)


Updated 14 October 2016. Link added 26 April 2019

Wednesday, October 13, 2010

Brownfield Requirements Engineering

Some partial and selective notes from the RESG meeting on Brownfield Requirements Engineering, held at BCS headquarters on October 12th 2010. By Richard Veryard.


Chairman's Introduction

Ian Alexander opened the meeting by pointing out the gulf between the demands of real engineering projects (95% brownfield) and the body of knowledge about requirements engineering offered by academics and authors (95% greenfield), and sketching some of the characteristic features of brownfield requirements. The importance of this topic is demonstrated by an excellent turnout for this meeting.

Brownfield Systems Engineering

Ian Gallagher of Altran Praxis presented some experience from several large systems projects, both military and civilian. The typical scenario is a major upgrade of existing systems to improve performance and deal with obsolescence; the upgrade may therefore include different types of requirement

  • doing new things
  • doing old things in new ways
  • migrating away from ageing technologies, proprietary systems or restricted technologies (e.g. those subject to export licenses)
The upgrade may be regarded as a single large increment (from AS-IS to TO-BE) but is often managed as a series of smaller increments.

With greenfield engineering, the requirements of the engineering project are pretty much the same as the requirements of the TO-BE system, but brownfield engineering introduces a critical choice: whether to write total requirements for the TO-BE system or to write requirements for the project (or separate increments) based on the difference(s) between AS-IS and TO-BE. IanG expressed a strong preference for the former, but acknowledged that this isn't always acceptable to key stakeholders, especially if the effort would be perceived as excessive. But in any case it's obviously important to be clear what kind of requirements we are talking about.

Beyond the scope of the TO-BE system, there may be a larger system-of-systems context, in which other equally large systems are the subject of equally large brownfield engineering activity. This raises questions of the interoperability of the increments, and the coordination between autonomous brownfield engineering projects, which is another important issue for brownfield requirements engineering.

A key issue is the pressure to reuse and repurpose existing assets. There are three possible reasons for this.
  1. Rational economics - genuine cost-saving based on lifetime exploitation of assets
  2. Irrational inertia - desire to justify past investment decisions, resistance to change
  3. Commercial - vested interests from key stakeholders (such as contractors) to maintain proprietary dependencies
Requirements such as migrating towards more "open" architectures are therefore not only relevant for an immediate brownfield project but also have important implications for the economics of future brownfield projects.

Managing Brownfield Financial Software

Phil Cantor of Smartstream then talked about his past and present experience managing and maintaining software products in the financial sector. (Most of his examples came from previous companies he had worked for.) The requirements of such products don't come exclusively (or even primarily) from the "users" but from a body of knowledge that the package vendor is expected to master. Phil gave the example of an obscure piece of financial calculation that is now only required in the Philippines: users elsewhere in the world may not know about this, but a package that doesn't cater for this requirement will be unacceptable to a global bank.

The package vendor then has to balance a wishlist of enhancement requests and bug reports from hundreds of user organizations, with the practical constraints of maintaining - and hopefully improving the quality of - an existing body of software.

For me, one of the most interesting issues raised by Phil was when he talked about the platform and its requirements. There is an internal platform supporting the user-facing components of the product suite, containing common services and suchlike, and the requirements for this platform cannot be derived purely from the requirements of the package as a whole. Furthermore, the package as a whole serves as a platform for the business processes of the user organizations (Phil's customers), so the same question arises at that level as well. However, this is not particularly a brownfield issue.

Phil also mentioned the challenges of training. From a sociotechnical perspective this is a major brownfield issue, since the performance of the TO-BE system will be affected by the user habits and expectations (e.g. knowing where to find things) carried forward from the AS-IS system. Training is a solution (and often not a very good one) to some set of sociotechnical requirements, and there may be better ways of conveying a new set of habits and expectations to the users and technical staff, but clearly we need to have an understanding of these requirements as well.

Discussion


Lots of further issues came up in the discussion, and I didn't manage to capture half of them. Someone raised the question of scale - what do you do with a brownfield situation with a complex network of interdependent pieces of hardware and software, with the need for upgrades constrained by small budgets and massive complexity. This led to the question of timing - how do you decide whether to upgrade something now or to defer the upgrade until next year. (Nobody actually mentioned the concept of technical debt, but that is clearly relevant here.)


Finally, my memory and notes can be supplemented by a few choice tweets.

@rmonhem much debate over how much to document requirements with significant re-use of existing applications

@rmonhem most difficult challenges have often been posed by cultural, commercial and legal factors.

@jiludvik Bit too much focus on techniques documenting and signing off detailed reqs. This is not specific for brownfield projects!

@jiludvik Phil Cantor's pres has been very entertaining and spot on for brownfield req's eng. I can certainly relate to many challenges he mentioned

Tuesday, February 24, 2009

Architecture or Firefighting?

For many organizations, the main worry today is Business Survival.

ZDNet blogger Ian Finley reckons "companies will need architects when the crisis is over, but today, they need fire fighters". (When the building’s on fire, who calls an architect?)

In response, fellow ZDNet blogger Joe McKendrick said that "the problem is we don’t have enough architecture as it is, and even when times are good, we’re always firefighting". (SOA 2009: Do we need architects or firefighters?) Mike Kavis agrees: "The reason there are so many fires is that there is so little architecture!"

Comparing the IT crisis with the financial crisis, my colleague David Sprott talks about a toxic mess of IT assets that can’t be easily cleared up. As he says, "although we may be cancelling projects and reducing activity, unless we get smart on architecture we are increasing the toxic mess". Also known as technical debt or technology debt (HT @jpmorgenthal).

As far as I can tell from Joe's later post, Ian and Joe have found a middle point of agreement: focus-focus-focus on the essentials. (2009: a year of extreme focus)

These are not the only voices arguing that Little SOA is more practical/pragmatic than Big SOA. See for example Jordan Braunstein.



At the bottom of these debates is an important cultural clash.

In many organizations, there are strong cultural divisions within IT - developers versus architects versus project managers. (In previous posts I have used the metaphor of planets: Mars, Venus, Saturn, or Uranus, while of course Enterprise Architects are from Pluto.) The current climate maybe seems to favour the project managers and the developers over the architects. But that doesn't mean that the architects should abandon their principles and expertise, and pretend to be project managers or developers. As if that's what anyone needs.

What architects do need to do is find a new way of engaging with the business and the rest of IT, in a way that achieves a better balance between the different cultures.

Some organizations have always been strongly focused on delivery - following the JFDI management style. In other organizations, the emphasis has often been on careful planning and inclusive consensus-building.

There are problems with both extremes. As a consultant, I once had the experience of working with two organizations at the opposite ends of this spectrum at the same time - trying to get the first organization to slow down and the second organization to speed up. As a result of this experience, I came up with a novel approach to working with organizational culture, based on the five elements of ancient Chinese thought. See my Slideshare presentation: Elements of Change.


Further reading

Martin Fowler, Technical Debt (Oct 2003)
Ruth Malan, Technical Debt (March 2013)
Dharmesh Shah, Understanding Technology Debt (August 2005)


Updated 14 March 2017