Showing posts with label risk. Show all posts
Showing posts with label risk. Show all posts

Monday, May 23, 2022

Risk Algebra

In this post, I want to explore some important synergies between architectural thinking and risk management. 


The first point is that if we want to have an enterprise-wide understanding of risk, then it helps to have an enterprise-wide view of how the business is configured to deliver against its strategy. Enterprise architecture should provide a unified set of answers to the following questions.

  • What capabilities delivering what business outcomes?
  • Delivering what services to what customers?
  • What information, process and resources to support these?
  • What organizations, systems, technologies and external partnerships to support these?
  • Who is accountable for what?
  • And how is all of this monitored, controlled and governed?

Enterprise architecture should also provide an understanding of the dependencies between these, and which ones are business-critical or time-critical. For example, there may be some components of the business that are important in the long-term, but could easily be unavailable for a few weeks before anyone really noticed. But there are other components (people, systems, processes) where any failure would have an immediate impact on the business and its customers, so major issues have to be fixed urgently to maintain business continuity. For some critical elements of the business, appropriate contingency plans and backup arrangements will need to be in place.

Risk assessment can then look systematically across this landscape, reviewing the risks associated with assets of various kinds, activities of various kinds (processes, projects, etc), as well as other intangibles (motivation, brand image, reputation). Risk assessment can also review how risks are shared between the organization and its business partners, both officially (as embedded in contractual agreements) and in actual practice.

 

Architects have an important concept, which should also be of great interest for enterprise risk management - the idea of a single point of failure (SPOF). When this exists, it is often the result of a poor design, or an over-zealous attempt to standardize and strip out complexity. But sometimes this is the result of what I call Creeping Business Dependency - in other words, not noticing that we have become increasingly reliant on something outside our control.


There are also important questions of scale and aggregation. Some years ago, I did some risk management consultancy for a large supermarket chain. One of the topics we were looking at was fridge failure. Obviously the supermarket had thousands and thousands of fridges, some in the customer-facing parts of the stores, some at the back, and some in the warehouses.

Fridges fail all the time, and there is a constant processes of inspecting, maintaining and replacing fridges. So a single fridge failing is not regarded as a business risk. But if several thousand fridges were all to fail at the same time, presumably for the same reason, that would cause a significant disruption to the business.

So this raised some interesting questions. Could we define a cut-off point? How many fridges would have to fail before we found ourselves outside business-as-usual territory? What kind of management signals or dashboard could be put in place to get early warning of such problems, or to trigger a switch to a "safety" mode of operation.

Obviously these questions aren't only relevant to fridges, but can apply to any category of resource, including people. During the pandemic, some organizations had similar issues in relation to staff absences.


Aggregation is also relevant when we look beyond a single firm to the whole ecosystem. Suppose we have a market or ecosystem with n players, and the risk carried by each player is R(n). Then what is the aggregate risk of the market or ecosystem as a whole? 

If we assume complete independence between the risks of each player, then we may assume that there is a significant probability of a few players failing, but a very small probability of a large number of players failing - the so-called Black Swan event. Unfortunately, the assumption of independence can be flawed, as we have seen in financial markets, where there may be a tangled knot of interdependence between players. Regulators often think they can regulate markets by imposing rules on individual players. And while this might sometimes work, it is easy to see why it doesn’t always work. In some cases, the regulator draws a line in the sand (for example defining a minimum capital ratio) and then checks that nobody crosses the line. But then if everyone trades as close as possible to this line, how much capacity does the market as a whole have for absorbing unexpected shocks?

 

Both here and in the fridge example, there is a question of standardization versus diversity. On the one hand, it's a lot simpler for the supermarket if all the fridges are the same type, with a common set of spare parts. But on the other hand, having more than one type of fridge helps to mitigate the risk of them all failing at the same time. It also gives some space for experimentation, thus addressing the longer term risk of getting stuck with an out-of-date fridge estate. The fridge example also highlights the importance of redundancy - in other words, having spare fridges. 

So there are some important trade-offs here between pure economic optimization and a more balanced approach to enterprise risk.

Tuesday, April 19, 2022

Data-Driven Reasoning (COVID)

Following the data is all very well, but how should we decide which data to follow?

 

Nate Silver argues that the total data are more important than the marginal data. There is more virus transmission in restaurants than in aeroplanes.

 

Several people have challenged this, including Carl Bergstrom.


The first point I want to make is about risk. When calculating the risk of transmission in aeroplanes, the level of risk somewhere else is not really relevant. As Bergstrom points out, the calculation should be based simply on the costs and benefits of wearing masks in that particular setting.

The problem here, of course, is determining the cost of wearing masks. If some people regard mask-wearing as a minor inconvenience, while others regard it as a major infringement of their human rights, what sort of data would be relevant to this calculation? Or if mask-wearing impedes communication to some extent, can we quantify this effect?

The broader question is whether we should be looking at total costs and benefits, or marginal costs and benefits? In the case of mask-wearing, if we assume that most people possess a usable mask, then the marginal cost of wearing one might simply include the increased frequency of washing or replacing it, plus whatever psychological, physiological or social costs we can determine. 

But for some people, opposition to mask-wearing is much more fundamental than that. It is not about the costs and benefits of mask-wearing in a particular setting, but an overall objection to living in the kind of society where mask-wearing is mandated at all. Some people even object (violently) to the sight of other people wearing masks. So any kind of mask-wearing may cause social division, therefore incurring a sociopolitical cost, and this needs to be set against the overall benefits of controlling the transmission of disease.

Nate Silver acknowledges that a mask-wearing mandate may be useful at the margin, but questions its overall value. Presumably people aren't going to wear masks in restaurants while they are eating. So his tweet seems to be asking what is the point of making them wear masks on planes?

This then brings us onto a question about policy, and the extent to which policy can be evidence-based. Is it better to have a consistent policy on mask-wearing across different settings, or to focus policy on those areas that are regarded as the highest risk? And should policies be optimized for effectiveness (what will have the greatest effect on suppressing transmission of the virus) or social acceptability (what will most people accept as reasonable)? Nate Silver's data may be relevant to this calculation, at least to the extent that they influence people's opinions.


While there are some particularly emotive features of this example, it raises some important general points about data-driven reasoning, especially in relation to cost-benefit calculations and evidence-based policy.



Sunday, January 01, 2017

The Unexpected Happens

When Complex Event Processing (CEP) emerged around ten years ago, one of the early applications was real-time risk management. In the financial sector, there was growing recognition for the need for real-time visibility - continuous calibration of positions – in order to keep pace with the emerging importance of algorithmic trading. This is now relatively well-established in banking and trading sectors; Chemitiganti argues that the insurance industry now faces similar requirements.

In 2008, Chris Martins, then Marketing Director for CEP firm Apama, suggested considering CEP as a prospective "dog whisperer" that can help manage the risk of the technology "dog" biting its master.

But "dog bites master" works in both directions. In the case of Eliot Spitzer, the dog that bit its master was the anti money-laundering software that he had used against others.

And in the case of algorithmic trading, it seems we can no longer be sure who is master - whether black swan events are the inevitable and emergent result of excessive complexity, or whether hostile agents are engaged in a black swan breeding programme.  One of the first CEP insiders to raise this concern was John Bates, first as CTO at Apama and subsequently with Software AG. (He now works for a subsidiary of SAP.)

from Dark Pools by Scott Patterson

And in 2015, Bates wrote that "high-speed trading algorithms are an alluring target for cyber thieves".

So if technology is capable of both generating unexpected events and amplifying hostile attacks, are we being naive to imagine we use the same technology to protect ourselves?

Perhaps, but I believe there are some productive lines of development, as I've discussed previously on this blog and elsewhere.


1. Organizational intelligence - not relying either on human intelligence alone or on artificial intelligence alone, but looking for establishing sociotechnical systems that allow people and algorithms to collaborate effectively.

2. Algorithmic biodiversity - maintaining multiple algorithms, developed by different teams using different datasets, in order to detect additional weak signals and generate "second opinions".





John Bates, Algorithmic Terrorism (Apama, 4 August 2010). To Catch an Algo Thief (Huffington Post, 26 Feb 2015)

John Borland, The Technology That Toppled Eliot Spitzer (MIT Technology Review, 19 March 2008) via Adam Shostack, Algorithms for the War on the Unexpected (19 March 2008)

Vamsi Chemitiganti, Why the Insurance Industry Needs to Learn from Banking’s Risk Management Nightmares.. (10 September 2016)

Theo Hildyard, Pillar #6 of Market Surveillance 2.0: Known and unknown threats (Trading Mesh, 2 April 2015)

Neil Johnson et al, Financial black swans driven by ultrafast machine ecology (arXiv:1202.1448 [physics.soc-ph], 7 Feb 2012)

Chris Martins, CEP and Real-Time Risk – “The Dog Whisperer” (Apama, 21 March 2008)

Scott Patterson, Dark Pools - The Rise of A. I. Trading Machines and the Looming Threat to Wall Street (Random House, 2013). See review by David Leinweber, Are Algorithmic Monsters Threatening The Global Financial System? (Forbes, 11 July 2012)

Richard Veryard, Building Organizational Intelligence (LeanPub, 2012)

Related Posts

Black Swans and Complex Systems Failure (April 2011)
The Shelf-Life of Algorithms (October 2016)
Robust Against Manipulation (July 2019)

Wednesday, October 26, 2016

The Shelf-Life of Algorithms

@mrkwpalmer (TIBCO) invites us to take what he calls a Hyper-Darwinian approach to analytics. He observes that "many algorithms, once discovered, have a remarkably short shelf-life" and argues that one must be as good at "killing off weak or vanquished algorithms" as creating new ones.

As I've pointed out elsewhere (Arguments from Nature, December 2010), the non-survival of the unfit (as implied by his phrase) is not logically equivalent to the survival of the fittest, and Darwinian analogies always need to be taken with a pinch of salt. However, Mark raises an important point about the limitations of algorithms, and the need for constant review and adaptation, to maintain what he calls algorithmic efficacy.

His examples fall into three types. Firstly there are algorithms designed to anticipate and outwit human and social processes, from financial trading to fraud. Clearly these need to be constantly modified, otherwise the humans will learn to outwit the algorithms. And secondly there are algorithms designed to compete with other algorithms. In both cases, these algorithms need to keep ahead of the competition and to avoid themselves becoming predictable. Following an evolutionary analogy, the mutual adaptation of fraud and anti-fraud tactics resembles the co-evolution of predator and prey.

Mark also mentions a third type of algorithm, where the element of competition and the need for constant change is less obvious. His main example of this type is in the area of predictive maintenance, where the algorithm is trying to predict the behaviour of devices and networks that may fail in surprising and often inconvenient ways. It is a common human tendency to imagine that these devices are inhabited by demons -- as if a printer or photocopier deliberately jams or runs out of toner because it somehow knows when one is in a real hurry -- but most of us don't take this idea too seriously.

Where does surprise come from? Bateson suggests that it comes from an interaction between two contrary variables: probability and stability --
"There would be no surprises in a universe governed either by probability alone or by stability alone."
--  and points out that because adaptations in Nature are always based on a finite range of circumstances (data points), Nature can always present new circumstances (data) which undermine these adaptations. He calls this the caprice of Nature.
"This is, in a sense, most unfair. ... But in another sense, or looked at in a wider perspective, this unfairness is the recurrent condition for evolutionary creativity."

The problem with adaptation being based solely on past experience also arises with machine learning, which generally uses a large but finite dataset to perform inductive reasoning, in a way that is non-transparent to the human. This probably works okay for preventative maintenance on relatively simple and isolated devices, but as devices and their interconnections get more complex, we shouldn't be too surprised if algorithms, whether based on human mathematics or machine learning, sometimes get caught out by the caprice of Nature. Or by so-called Black Swans.

This potential unreliability is particularly problematic in two cases. Firstly, when the algorithms are used to make critical decisions affecting human lives - as in justice or recruitment systems. (See for example, Zeynap Tufekci's recent TED talk.) And secondly, when preventative maintenance has safety implications - from aeroengineering to medical implants.

One way of mitigating this risk might be to maintain multiple algorithms, developed by different teams using different datasets, in order to detect additional weak signals and generate "second opinions". And get human experts to look at the cases where the algorithms strongly disagree.

This would suggest that we maybe shouldn't be too hasty to kill off algorithms with poor efficacy, but sometimes keep them in the interests of algorithmic biodiversity.  (There - now I'm using the evolutionary metaphor.)



Gregory Bateson, "The New Conceptual Frames for Behavioural Research". Proceedings of the Sixth Annual Psychiatric Institute (Princeton NJ: New Jersey Neuro-Psychiatric Institute, September 17, 1958). Reprinted in G. Bateson, A Sacred Unity: Further Steps to an Ecology of Mind (edited R.E. Donaldson, New York: Harper Collins, 1991) pp 93-110

Mark Palmer, The emerging Darwinian approach to analytics and augmented intelligence (TechCrunch, 4 September 2016)

Zeynap Tufekci, Machine intelligence makes human morals more important (TED Talks, Filmed June 2016)


Related Posts
The Transparency of Algorithms (October 2016)

Monday, March 18, 2013

Cloud and Continuity of Supply Risk

@dougnewdick points out the risk of a company becoming over-dependent on Google. His particular example is prompted by Google's announcement that Google Reader will be discontinued.

I have previously commented on the subject of Creeping Business Dependency, the fact that many companies have allowed themselves to become dependent on a particular company, product or technology. Especially Google. If Google decides your website offends against some search engine rules, it is perfectly capable of making your website disappear from searches. (BMW disappeared from Google for three days in 2006 - see my post BMW Search Requests.) A company might well go bust before it could sort the problem out.

Of course, you can't avoid some dependencies, but I think it is important that any significant dependency should be clearly visible in the business architecture. (In general, business architects usually neglect this kind of dependency until I point out specific examples to them.)

When looking at this kind of dependency, it is important to remember the principles of asymmetry - the Product is not the Technology, and the Company is not the Product. There have been a few popular products and platforms whose owners lost interest - these included Bloglines (formerly owned by Ask) and Delicious (formerly owned by Yahoo) - but were revived under new ownership. Users of a popular platform may feel that a large user base provides grounds for optimism that someone will want to keep it going, even if the original owner doesn't wish to. However, there are many products and platforms that have not survived.

More fundamental is the question of the underlying technology. A few years ago, there was considerable confidence and investment in RSS and Atom feeds, and a number of products and platforms were developed to exploit this technology. If there is a healthy ecosystem of different products and platforms, with relatively low switching costs, it doesn't matter much if one product drops out. But if Google and others are losing interest in this technology, that's a much more fundamental problem for anyone who is heavily committed to it.

If Google stops providing a free service, those who really want it may have to pay to get a decent service elsewhere. But this alters the economics of the service ecosystem, with unpredictable consequences. Clearly there is a risk that the service you want (or the service you need your customers to use) is increasingly expensive, inconvenient and ultimately unavailable.


Doug Newdick, Cloud and Continuity of Supply Risk (March 2013)

Saturday, February 18, 2012

BYOD - Bring Your Own Device

By popular demand, many companies are shifting ownership of elements of corporate infrastructure onto their employees. This is known as BYOC (bring your own computer) or BYOD (bring your own device).

There are many aspects to this trend.

1. Culture. Talented recruits may see this kind of choice as a desirable feature of a future employer. Some of them may have a strong personal commitment to a particular device; others may ask about BYOD policy as a quick way of getting a general impression of company culture and its attitude towards employees.

(Even if BYOD is a common request at interview, this doesn't mean it is a genuine requirement. In some cases, the BYOD request could be similar to the apparently crazy riders that performers may add to contracts as a way of testing the diligence and attention to detail of the organizers. The best-known example of such a contract rider is Van Halen's insistence on a bowl of MnMs with the brown ones removed. See "Brown out" at snopes.com.)

2. Interoperability. There is a need for interoperability within the enterprise (endo-interoperability) as well as interoperability with external platforms (exo-interoperability). Within the enterprise, people expect to be able to use common services (email, communications, content management, and so on) regardless of device. When I'm in the office, I want to be able to connect my device to office devices such as printers and projectors, as well as using the office network and servers. When I'm working at home, I want to be able to connect my device into the office systems, and use my device for web conferences and other events. But I also want to be able to connect my device into public platforms such as Facebook.


3. Innovation. Early adopters like to carry the latest and most fashionable device, even if this doesn't yet support all the required corporate services in a robust manner.

4. Business continuity and risk. A person's productivity can be seriously impaired if the device is lost or develops a fault. Conversely, a company's security can be seriously impaired if an employee uses an unverified emergency device such as her teenage son's phone. Does BYOD imply the rapid availability of backup devices of every conceivable brand, or does the company provide a limited range of standard devices for emergency use?

5. Support. Does the device deliver all the required corporate services correctly, efficiently and securely? Whose responsibility is it to verify and test these services on the given device, and to sort out the (inevitable) configuration problems? What knowledge and expertise is needed to provide adequate support across the full range of devices?

6. Economics. Device provision within large organizations was traditionally based on the economics of scale. We purchase thousands of identical devices, install the same software and services on each one, and issue these to our employees. We can obtain good discounts from the hardware and software suppliers, and we can train our support staff to provide efficient support across a narrow range of products. But this approach fails to deal with the complexities of the modern business organization where each employee has different needs, often calling for additional non-standard software and services, or even newer devices. So most modern organizations shift to provision of devices based on the economics of scope - giving everyone a flexible device platform to which additional software and services can be easily added. Then the move to BYOD takes us into the economics of alignment - optimizing the lifetime cost of device provision against the lifetime benefits to the organization and the individual within the context of use.

7. BYOD represents a shift in the balance between two kinds of device vendor - the ones who sell thousands of devices at a time by schmoozing the CIO and the ones who sell devices to individuals via consumer channels. (As a result, some stakeholders may be cynical and unsympathetic to any objection to BYOD from the CIO quarter.)

8. More fundamentally, BYOD represents a shift in the balance of power between two kinds of knowledge. The corporate IT folk supposedly know more about the corporate services and about quality attributes such as reliability and security. However, the individual employee knows more about the context of use. The architectural question here is aligning the device selection, configuration and use with the emerging requirements of the individual in the job. This is ultimately a question of governance, which needs to be guided by appropriate BYOD policies.


A lot of architectural issues then.


Fiona Graham, BYOC: Should employees buy their own computers? (BBC News 14 January 2011)

Fiona Graham, BYOD: Bring your own device could spell end for work PC (BBC News 14 February 2012)

Eric Vanderburg, Four Keys to Successful BYOD (CIO 14 February 2012)


Related posts:

Bring your own expectations (May 2014)

Tuesday, November 29, 2011

Risk and Responsibility in Self-Service

A cabbie asked @jkuramot to enter his destination into the GPS. @dahowlett suggests this is because he didn't speak good English. @jkuramot confirms that the driver didn't speak English very well but adds that "this was his go-to move".

The reason we are talking about this fragment of service design is that it is unusual in this context: we normally expect the driver to enter the destination into his navigation device. But the normal procedure is prone to error; the passenger may not speak clearly, the driver may not understand correctly, there may be a lot of background noise: the passenger arrives at the wrong destination and it's the driver's fault.

However, if the passenger enters the destination directly into the navigation device, then any error is the passenger's fault. Many service providers in other areas now follow this pattern; shifting responsibility onto the customer may help to reduce administration costs, but more importantly reduces the service provider's liability. But if the customer is not able to perform these tasks easily and accurately, this kind of shift adds more to the cost and risk for the customer than it reduces for the supplier, and therefore diminishes total value. See my review of The Support Economy.

Asking the customer to do the work makes an assumption about the customer's capability. I don't know Jake personally, but he looks from his photo and his Twitter profile like someone who would know how to operate this kind of device. The driver may have had the same impression; it is conceivable that he would have treated Jake's grandmother differently. Whereas if the device (belonging to the driver) is unusual and difficult to use, we would always insist that the driver should operate it. Self-service only works if the interface design offers a reasonable level of usability.

The other difference between the passenger and the driver is the question of which is more familiar with the destination. When I get a cab home from the airport, obviously I know my address better than the driver does. But when I arrive in a strange city, I expect the cab drivers to be more familiar with the hotels than I am: if I get the name of the hotel slightly wrong, the driver should ask if I really meant something else, rather than drive for an hour to a hotel in the next city whose name exactly matches what I said.

By the way, Google has been correcting our searches for a long time now, but has now chosen to issue a series of advertisements in which this correction (and the collection of vast amounts of data to make this correction possible) is highlighted as a service enhancement feature. See my note Towards a VPEC-T analysis of Google. This kind of service enhancement is unavailable if the driver takes himself out of the loop, and regards his job as merely enacting a specification agreed between the customer and an electronic device.

Wednesday, April 06, 2011

Black Swans and Complex System Failure

Black Swan theory (Wikipedia) tells us among other things that people tend to underestimate the probability of extremely rare events.

A corollary of this theory that is of particular interest to architects and complex system engineers concerns the design of fail-safe mechanisms. Nuclear power and oil extraction are examples of environmentally critical operations; they are therefore subject to detailed risk assessment, and designed with multiple fail-safe mechanisms. And yet both the oil spillage last year in the Gulf of Mexico and the partial melt-down in Japanese nuclear reactors following the recent tsunami involved the simultaneous failure of multiple fail-safe mechanisms. Obviously that's not supposed to happen.

Simultaneous failure of supposedly independent mechanisms is a Black Swan event.


Update (August 2011)

A recent study by Oxford University and McKinsey has blamed rare but high-impact problems, dubbed "black swans", for the increasingly common phenomenon of large IT project whose cost spirals out of control. The study finds this phenomenon to be three times as common in IT than in other domains [BBC News, 26 August 2011]. See my post on Black Swan Blindness.



Update (October 2011)

Reviewing a couple of recent books about BP and the oil spill in the Gulf of Mexico, Mattathias Schwartz makes a number of relevant points.

When crucial pieces of our infrastructure fail, they do so gracelessly, without much warning and in ways that are difficult to anticipate. ... The failure to grasp the possibility of system-wide failure might be one in an accelerating series, bookended by the 2008 financial crisis and the Fukushima nuclear meltdown last spring.
One reason for the oil and gas industry’s quick comeback in the US was the successful packaging of the blowout as a ‘black swan’, an event of such low probability that it couldn’t have been anticipated. This certainly helped excuse the fact that no one – not BP, Chevron, Exxon or Shell – had a working plan for plugging a blowout as deep as Macondo .
BP ... claimed, in its own report on the blowout, that the event had eight causes, of which BP was partly responsible for one. The president’s commission concluded that the disaster had nine causes, and that BP was responsible for six or seven. And yet BP stands by what it said at the start. 

The size of the system and the complexity of the data make it possible to argue for a maddeningly wide range of positions, especially when it comes to vague legal notions like ‘negligence’ or ‘responsibility’. Both concepts hinge on proving that one linear narrative is the right one. 

Mattathias Schwartz, LRB 6 October 2011 
reviewing
  • Spills and Spin: The Inside Story of BP by Tom Bergin 
  • A Hole at the Bottom of the Sea: The Race to Kill the BP Oil Gusher by Joel Achenbach

Tuesday, March 08, 2011

Creeping Business Dependency

People are slowly waking up to the fact that we have created yet another single point of failure into our business ecosystem. It seems that businesses have gradually made themselves dependent on Global Positioning Systems (GPS) and satellite navigation (satnav). So we are now starting to hear doom-and-gloom stories about the dire economic consequences of any interruption to the service, which could apparently be caused by anything from cyberterrorism (Daily Mail 8 March 2011) to solar flares (Daily Mail 21 Sept 2010).

Those with long memories may recall the millennium bug scare, which postulated that widespread computer error might result in total economic collapse when the date went from 99 to 00. Many companies took the opportunity to carry out a long overdue inventory of their software programs, and decommissioned a fair amount of obsolete code, as well as reviewing their disaster recovery procedures; even though the scare was probably exaggerated, some useful work was done. (I myself picked up some contract work in this area, so I can't complain.)

The Royal Academy of Engineering has just issued a report on Global Navigation Space Systems, which takes a more balanced view of the subject than the Daily Mail, but still warns of the danger of over-reliance on satellite navigation [Report (pdf), Press Release].

Chairman of the RAoE working group, Dr Martyn Thomas, told the BBC:
"We're not saying that the sky is about to fall in; we're not saying there's a calamity around the corner. What we're saying is that there is a growing interdependence between systems that people think are backing each other up. And it might well be that if a number these systems fail simultaneously, it will cause commercial damage or just conceivably loss of life. This is wholly avoidable." [BBC News 8 March 2011]

Maybe this does sound pretty speculative (as @martinjmurray complains). Nonetheless it may be a good idea for any business that has gradually become dependent on this or any other technology to check out the possible risks.

From an architectural point of view, what I find most interesting about this situation is the tendency for critical business dependencies (and the associated risks) to emerge, as a particular technology migrates unobtrusively from marginal use to core business use.

Another example of a creeping business dependency is the extent to which Google has now inserted itself into the relationship between any business and its customers. If a business offends Google in some way, and consequently disappears from Google search, this will have serious business consequences. (BMW disappeared from Google for three days in 2006 - see my post BMW Search Requests). And yet it's still rare to see Google shown as a business-critical service partner in business architecture or business process diagrams.

If we think of an architecture in terms of a set of dependencies, we can distinguish between a centrally planned architecture, in which the dependencies and their implications are understood from the outset, and an emergent defacto architecture, in which unanticipated dependencies and risks can be created by a quantity of uncontrolled activity. In a planned world, all innovation must be controlled to prevent emergent risk; in an evolving world, innovation (such as the use of Google or GPS) can be encouraged provided that there is a robust mechanism to detect and manage emerging risks.


Related posts: BMW Search Requests (Feb 2006), Cloud and Continuity of Supply Risk (March 2013)

Sunday, May 16, 2010

SOA and Risk Management

#soa #risk In this post, I identify some contrasting views on the relationship between SOA and risk.

SOA involves innovation, and innovation always introduces new risk


SOA helps reduce risk


    SOA providing visibility and control of aggregate risk and unexpected behaviour


    SOA complicating visibility and control of aggregate risk


    Therefore ... risk management as one area likely to see spending increasing


    Oh yeah? Any evidence of this?

    Monday, September 15, 2008

    SOA Example - Real-Time Regulation

    In a couple of recent posts on Turbulent Markets, I asked whether real-time profit and loss could tame turbulent markets (answer: No), and asked whether some other form of real-time event-driven system could perform a regulatory function (answer: Possibly).

    Bloggers from some of the CEP vendors have been making similar suggestions for a while.

    Back in March 2008, there was a flurry of interest in the strange fate of Eliot Spitzer, who was apparently exposed by the very regulatory technologies he himself had advocated.
    Jesper Joergensen of BEA (BEA now part of Oracle, Jesper has now joined SalesForce, Jesper's BEA blog has disappeared) took the opportunity to put in a plug for BEA's event processing products. "If anyone working in a bank's anti-money laundering, compliance or fraud detection unit is reading this", he writes, "go check out [my company's products]. This is the technology you need to automate these compliance requirements."

    There is obviously a need for systems to trap people like Eliot Spitzer. But I'm not convinced that simple compliance systems need to be real-time service-oriented event-driven systems. See my post on Real-Time Fraud Detection.

    But a much stronger case can be made for real-time risk management. Chris Martins (Progress Apama) put the case for CEP and real-time risk in March 2008. More recently, Jeff Wotton (Aleri) has put the case for real-time risk consolidation, drawing on an interview with Nick Leeson. Meanwhile, Progress Actional has a product page on Real-Time Risk Profiling for Banks.

    The point about real-time risk aggregation is that you need to produce a rapid and reliable picture of total risk, drawing on data from many heterogeneous sources. In a typical business environment, new types of risk and sources of data are constantly being added, and you want to be able to plug these into your risk consolidation straightaway. In this kind of scenario, it should be very easy to justify SOA.

    Saturday, September 06, 2008

    Tangential Service

    Sometimes a service can be provided as a cheap or unimportant side-effect of something else.

    I was struck by this thought when I saw a news story indicating shortages of Molybdenum-99, a radioactive isotope used for medical scans. There are five nuclear reactors in the UK that normally produce this isotope as a by-product, but two of the five are currently down for scheduled maintenance and a third has encountered an unexpected problem. [Isotope problem may delay scans, BBC News, 5 Sept 2008] Nuclear medicine is dependent on rare isotopes produced in nuclear reactors; however, the nuclear reactors' primary purpose is the generation of electricity and not the supply of radioactive materials to hospitals.

    Here's another example. Sally is planning to drive her daughter to a swimming gala, so she offers to drive her daughter's school friend as well. In the morning of the gala, Sally's daughter has a fever and cannot swim. It is now too late for the friend's parents to make alternative arrangements. Is Sally morally obliged to stand by her promise to take her daughter's friend, although she would prefer to stay at home with her sick daughter?

    Here are the common factors in these examples
    • A is providing a service to B
    • This service is not (and never will be) the primary mission or purpose of A.
    • B is dependent on this service
    • Under normal circumstances, the cost to A of providing the service is negligible.
    • If A's circumstances change, it may become inconvenient, expensive or impossible for A to provide the service to B.
    I'm going to call services that fit this pattern tangential services, and I think there are important considerations for both providers and consumers.
    • A may be unwilling to invest any resources in monitoring or improving the service. B might be willing to contribute some resources to improving the service (especially improving its reliability), but this may be difficult to manage.
    • B needs to monitor the service, but may not have access to information that would provide advance warning of service problems (because this is private to A).
    • Service level agreements may be weakly specified or completely absent. A may have no formal obligations to B.
    • Under normal circumstances, everyone is happy. There may be a sudden jump from happy (mutally convenient, value-adding) to unhappy (inconvenient to A and/or B, unexpected cost).
    There is of course nothing wrong with providing or consuming tangential services. For the provider, it may be a bit of extra revenue or an opportunity to provide some social value. For the consumer, it may represent a significant cost-saving, or provide access to something that might otherwise have been impossible. However, it is important to understand the implications of a service's being tangential, and to think through questions of service levels, liabilities and contingency plans.

    Tuesday, August 12, 2008

    Responding to Uncertainty

    How does a system respond intelligently to uncertain events?
    "A person may take his umbrella, or leave it at home, without any ideas whatsoever concerning the weather, acting instead on general principles such as maximin or maximax reasoning, i.e. acting as if the worst or the best is certain to happen. He may also take or leave the umbrella because of some specific belief concerning the weather. … Someone may be totally ignorant and non-believing as regards the weather, and yet take his umbrella (acting as if he believes that it will rain) and also lower the sunshade (acting as if he believes that the sun will shine during his absence). There is no inconsistency in taking precautions against two mutually exclusive events, even if one cannot consistently believe that they will both occur." [Jon Elster, Logic and Society (Chichester, John Wiley, 1978) p 84]

    Austrian physicist Erwin Schrödinger proposed a thought experiment known as Schrödinger's cat to explore the consequences of uncertainty in quantum physics. If the cat is alive, then Schrödinger needs to buy catfood. If the cat is dead, he needs to buy a spade. According to Elster's logic, he might decide to buy both.

    At Schrödinger's local store, he is known as an infrequent purchaser of catfood. The storekeeper naturally infers that Schrödinger is a cat-owner, and this inference forms part of the storekeeper's model of the world. What the storekeeper doesn't know is that the cat is in mortal peril. Or perhaps Schrödinger is not buying the catfood for a real cat at all, but to procure a prop for one of his lectures.

    Businesses often construct imaginary pictures of their customers, inferring their personal circumstances and preferences from their buying habits. Sometimes these pictures are useful in predicting future behaviour, and for designing products and services that the customers might like. But I think there is a problem when businesses treat these pictures as if they were faithful representations of some reality.

    This is an ethical problem as well as an epistemological one. You've probably heard the story of a supermarket, which inferred that some of its female customers were pregnant and sent them a mailshot that presumed they were interested in babies. But this mailshot was experienced as intrusive and a breach of privacy, especially as some of the husbands and boyfriends hadn't even been told yet. (A popular version of the story involves the angry father of a teenaged girl.)

    Instead of trying to get the most accurate picture of which customers are pregnant and which customers aren't, wouldn't it be better to construct mailshots that would be equally acceptable to both pregnant and non-pregnant customers? Instead of trying to accurately sort the citizens of an occupied country into "Friendly" and "Terrorist", wouldn't it be better to act in a way that reinforces the "Friendly" category?

    Situation models are replete with indeterminate labels like these ("pregnant", "terrorist"), but I think it is a mistake to regard these labels as representing some underlying reality. Putting a probability factor onto these labels just makes things more complicated, without solving the underlying problem. These labels are part of our way of making sense of the world, they need to be coherent, but they don't necessarily need to correspond to anything.


    Minor update 10 Feb 2019

    Saturday, June 21, 2008

    Does Multi-Tenancy Matter?

    Multi-tenancy basically means that the service provider is supporting several customers with the same resources. (There is some ambiguity about exactly which resources we are talking about - hence the embarrassingly public disagreement between Oracle and one of its reference SaaS users described by Phil Wainewright in Many degrees of multi-tenancy.)

    Gianpaolo Carraro previously made the point that if I'm a service consumer, I shouldn't care about multi-tenancy - The multi-tenant emperor has not clothes (August 2006), and now adds I can't believe we're still talking about this (June 2008).

    With services like laundry, I really shouldn't care if my dirty clothes are put into the same load as everyone else's, as long as the service provider can reliably sort them all out and return them correctly. Some providers may think that the cost savings from multi-tenancy of the washing machine doesn't justify the hassle of labelling and sorting the clothes, but that's surely their problem not mine.

    But if I have doubts about their competence and reliability, and if I have to double-check everything (= increased transaction cost) because I feel there is an increased risk of error on his part, then it becomes my problem as well.

    Gianpaolo uses the example of a restaurant kitchen. If someone on the next table orders the same dish at the same time, surely I don't care if the chef puts two slices of meat together into the same pan. Well I do care if it means that the chef is tempted to compromise, or pays insufficient attention to my special requirements. As a service consumer, I may have some theory about the likely behaviour and incentives of the service provider (Gianpaolo talks about people wanting to show off their architectural capacities). But this only matters to the extent that it affects what I end up with, or when.

    SaaS inherits from SOA the principle of encapsulation - the idea of separating the specification (WHAT) from the implementation (HOW). But a lack of trust between service consumer and service provider, as well as possible incentive incompatibility, leads to a breach in encapsulation. For SOA and SaaS to work properly, you need a good line on managing quality and risk. But that's a story for another post.

    Saturday, January 19, 2008

    Technological Perfecta

    There are several technologies that might work well together, indeed they certainly should work well together. At various times in this blog, I've talked about the potential synergies between (i) SOA and Business Intelligence, (ii) SOA and Business Process Management, and (iii) SOA/EDA and Complex Event Processing. The third of these synergies is currently getting some attention, following some enthusiastic remarks by Jerry Cuomo, WebSphere CTO (see Rich Seeley and Joe McKendrick).

    All four together would be amazing, but a lot of organizations aren't ready for this. Moreover each technology has its own set of tools and platforms, and its own set of disciplines and disciples.

    In Betting on the SOA Horse, Tim Bass describes this potential synergy using the language of gambling - exacta and trifecta. I'm not very familiar with this language, but what I think this means is that you only win the bet if the horses pass the post in the correct sequence. Tim writes:
    "Betting on horses is a risky business. Exactas and trifecta have enormous payouts, but the odds are remote."

    In On Trifecta and Event Processing, Opher Etzion disagrees with this metaphor. He argues that these technologies are mutually independent (he calls them "orthogonal"). If he is correct, this would have three consequences: (i) flexibility of deployment - you can implement and exploit them in any sequence; (ii) flexibility of benefit - you can get business benefits from any of them in isolation, and then additional benefits if and when they are all deployed together; and therefore (iii) considerably lower risk.

    My position on this is closer to Opher. I think there are some mutual dependencies between these technologies, but they are what I call soft dependencies. P has a hard dependency on Q if Q is necessary for P. Whereas P has a soft dependency on Q if Q is desirable for P.

    In planning a technology change programme, it is very useful to recognize soft dependencies, because it permits some deconfliction between different elements. Deconfliction here means forced decoupling, understanding that the results may be sub-optimal (at least initially), but accepting this in the interests of getting things done.

    In a perfect world, we might want to deploy all four technologies together, or in a precisely defined sequence. But pragmatism suggests we don't bet on the impossible or highly improbable. The challenge for the technology architect is to organize a technology portfolio to get the best balance of risk and reward. This is not primarily about comparing the features of different products, but about understanding the fundamental structural principles that allow these technologies to be deployed in a flexible and efficient manner.

    Discussion continues: Technological Perfecta 2

    Wednesday, May 16, 2007

    Service Escrow - Iron Mountain

    Last month I riffed with Gianpaolo Carraro about Service Escrow. A few days later, a company called Iron Mountain launched an SaaS escrow service, covering the source code, system documentation and data.
    Gianpaolo (who is one of Microsoft's SaaS experts) refers to companies providing this service as SaaS undertakers. But it might be better to call them SaaS support. By acting as a safety net for SaaS providers and their customers, they may reduce the risk associated with SaaS and contribute indirectly to SaaS market growth.

    I phoned Iron Mountain to find out more about their SaaS Escrow service, and the extent to which it differed from the traditional software escrow service. Here are some of the main points of our conversation.

    Storage Model

    Iron Mountain stores both the software (source code, object code and documentation - all of which belongs to the SaaS provider) and the data (which belongs to the SaaS user). Whereas traditional software escrow can often operate on a fairly slow cycle, with new software versions deposited in a fairly leisurely manner, SaaS escrow generally calls for live data backup.

    If the escrow conditions are triggered, Iron Mountain will release the software and data to the SaaS user. To avoid any perceived conflict of interest, Iron Mountain does not operate the software - even on an emergency basis. But this means that the SaaS developers (or some appointed third party) must install and test the software on a new server, load and test the data, and then restore the service.

    This is probably not something that can or should be done overnight - so it is not going to provide much protection against a sudden and unexpected failure of a business critical service. A more likely scenario is a gradual worsening of the relationship between the SaaS provider and the SaaS user, and a growing dissatisfaction with service levels and support arrangements, approaching the point where the SaaS provider is in breach of contract. SaaS escrow means that the SaaS user has a reasonable exit from an unsatisfactory relationship, and cannot be held to ransom because the SaaS provider controls a key business asset.

    The economic benefits of escrow therefore fall under the economics of governance - making sure the SaaS user has proper control of the relationship.

    Charging Model

    Although the SaaS provider may benefit from the existence of SaaS escrow, the prime beneficiary is the SaaS user. Historically, it has always been the user who has negotiated and paid for escrow. However, Iron Mountain is increasingly seeing more complex arrangements whereby the software provider pays for the escrow and passes these charges onto the sofware user.

    In the case of SaaS escrow, a typical arrangement would be that the SaaS provider pays for the deposit of the software, while the SaaS user pays for the deposit of the data (depending on the data volumes).

    It is in the interests of the SaaS user that Iron Mountain always holds the latest version of the software. Iron Mountain therefore encourages the SaaS provider to deposit software upgrades as frequently as necessary - preferably online - and does not charge on the basis of frequency. (Remember that SaaS software may undergo a faster improvement cycle than traditional software packages.)

    Management Process

    SaaS escrow only works if the SaaS provider has adequate software configuration management and data management, so that it becomes a matter of routine to send controlled copies to Iron Mountain. There is a certain amount of ongoing verification and audit that needs to be carried out, involving all the parties to the escrow arrangement, and Iron Mountain sees this as an important aspect of its own role.

    Remember that the SaaS user does not see the software unless and until the escrow conditions are triggered - so it isn't possible to test the escrow arrangements in a trial run exercise. Instead, it is necessary to test the process at a higher level of abstraction, to provide some reassurance to the SaaS user that the escrow would work adequately.

    Future

    This is early days for the SaaS escrow market, and I shall be interested to see how the market develops ...

    Monday, April 16, 2007

    Service Escrow

    A small SaaS company bites the dust, reports Gianpaolo Carraro, Director of SaaS Architecture at Microsoft.


    Obviously this is not just an SaaS problem. Companies fold all the time. If you are dependent on some external capability, you'd better have a plan for business continuity. And if you have entrusted your supplier with important assets (e.g. data) you'd better be able to get your assets back quick.

    A pessimist might regard this risk as a reason to avoid SaaS altogether. But I don't think Microsoft employs many pessimists. For his part Gianpaolo sees this risk as an opportunity for some enterprising SaaS undertaker - selling services to those whose beloved supplier has gone to the great SaaS graveyard.

    I prefer to see this as an architectural challenge:
    • How can we design a robust network of services, one not affected by a single point of failure? (This is of course a classic problem of distributed systems.)
    • Are there patterns of collaborative networks that will stand up to the loss of any single organization in the network?
    • What are the appropriate SLAs to support these patterns?
    • And what are the interoperability requirements to make this work?
    Structural problems call for structural solutions. That's what architects are for.

    Update

    Gianpaolo's initial response here: SaaS Undertaker (April 2007). Shortly after my exchange with Gianpaolo, a company called Iron Mountain launched an escrow service. See my post Service Escrow - Iron Mountain (May 2007).

    Tuesday, May 30, 2006

    Business Case for SOA

    With pre-SOA technologies, the business case is typically weakened by uncertainty. We have to factor in the possibility that the benefits will be less than expected, and the costs and timescales greater.

    With SOA, the business case may be strengthened by uncertainty. SOA helps to protect the business and its systems against risk, and keeps more options open. The greater the uncertainty, the greater the value of options.

    Business Case for SOA

    1. The classic way of constructing a business case is to estimate the costs and the benefits, estimate the timescale in which these costs and benefits will occur, and use this to calculate a return on investment (ROI). Timescale is important in investment decisions - costs and benefits are usually spread over an extended period. Earlier benefits are worth more than later benefits. A standard accounting technique for calculating ROI by aggregating costs and benefits over different times is known as discounted cash flow (DCF). Future cash flows are discounted even if they are 100% certain.
    2. But in most situations there are various dimensions of uncertainty - about whether, how much, and when benefits will be achieved; and whether additional costs and delays may be incurred. So the business case needs to factor in an adjustment for risk, which can reduce the economic feasibility of a proposition. Higher uncertainty is usually reflected in a higher discount rate, and this reduces the value of future benefits, especially longer-term ones.
    3. If we now add SOA into the equation, we may be able to identify differences to cost, timescale and benefit that are produced by SOA. SOA is then justified if the net effect is positive (in other words, if there is any additional cost from SOA, it is more than offset by faster or increased benefits, or reduced cost elsewhere). In some cases, SOA may convert a non-feasible proposition into a feasible one.
    4. Finally, we may adjust the SOA increment for its effect on risk. In many cases, one of the most important elements of a business case for SOA is its positive contribution to risk management, and therefore feasibility. A well-design loosely-coupled architecture will be protected against events that might otherwise increase total lifetime costs or reduce benefits.
    Technorati Tags:

    Thursday, February 09, 2006

    BMW Search Requests

    Another controversy hits Google. German car-maker BMW offends against Google's webspam policies and gets delisted. After three days of "death penalty", the listing is resurrected. (Smaller companies might be lucky to be reinstated months after a lesser offence.)

    Sources: Matt Cutts (Google), Search Engine Watch

    Google is here playing the same defacto regulatory role as a stock exchange - if you want us to list you, please play according to our rules. (Google can therefore add extra rules to those rules imposed by national or international law. Whether Google can subtract rules is of course an entirely different and perhaps even more controversial question.)

    But there is a crucial difference. With a stock market, there is a documented agreement between the stock exchange and the listed company, and any disputes can be taken to legal process. However, it is difficult to see how any company, however large, might establish in law that it had been unfairly treated by Google in such a matter. Even if there was clear evidence of financial loss as a result of a material change in search ranking, it would probably be impossible to pin liability for this loss onto Google.

    Look at BMW's position from a service-oriented perspective. BMW's promotional and marketing capability is critically dependent on Internet search services provided to BMW's customers by Google and its competitors. This is of course true of many companies, large and small. But for most or all of these companies (and I presume this is the case for BMW), there is no formal contract with Google that governs this particular dependency, and no defined service level agreement.

    Google's published statements appear to offer some comfort, implying that it will only take this kind of action against a company in response to some detected non-compliance on the company's part with Google's rules. (We can interpret these statements as part of a contractual pseudo-contractual specification of the service provided by Google.) But there are two ways in which this official specification can fail. Firstly, Google might make a mistake - may incorrectly punish a company for something it hasn't done. And secondly, Google might be tricked by a third party (competitor or vandal) into punitive action. (For example, by a clever combination of cybersquatting and hacking.)

    Bottom Line: If you are dependent on a service you don't control, then there is a risk that needs to be managed. In a service economy, there will be many services that you want to use, because the potential rewards of these services are well high enough to justify the risks. But these risks still exist, and service management needs to include proper attention to the relevant risk factors.

     

    Update

    A related example (March 2006), via Seth Godin. Kinderstart sues Google over lower page ranking (Reuters).

    See also:  Creeping Business Dependency (March 2011), Unruly Google and VPEC-T (January 2012)

    Friday, July 07, 2000

    Business Services and Risk Management

    A common reaction to the concept of web service based architectures is to be deeply concerned about operational integrity. It is primarily for this reason that perhaps the majority of media commentators are advising that web services will not become pervasive for three to five years. However this perspective is based on an extrapolation of current practice. In the future, outsourced services are more likely to be business services or software services, but not application services. Software services will rapidly become pervasive in the same timeframe as the current PC model of computing transitions to the web based model, and we rent our usage of whatever personal productivity software we choose to use. Business services will be strongly favored over application services because this places the commercial costs and risks together with the operational responsibility and overcomes the operational issues.

    Suppose I represent an insurance company, and I use a component-based service from another company to help me perform the underwriting. Don't I need to know the algorithm that the other company is using? Suppose that the algorithm is based on factors that I don't believe in, such as astrology? Suppose that the algorithm neglects factors that I believe to be important, such as genetics or genomics? Am I not accepting a huge risk by allowing another company to define an algorithm that is central to my business?

    There are three main attitudes to this risk. One attitude, commonly found among civil servants, lawyers and software engineers, is to break encapsulation, crawl all over the algorithm in advance, and spend months testing the algorithm across a large database of test cases. If and when the algorithm is finally accepted and installed, such people will insist on proper authorization (with extensive retesting) before the smallest detail of the algorithm can be changed. The second attitude is denial: impatient businessmen and politicians simply ignore the warnings and delays of the first group.

    There is a third approach: which is to use the forces of competition as a quality control mechanism. Instead of insisting that we find and maintain a single perfect algorithm, or kidding ourselves that we've already achieved this, we deliberately build a system that sets up several algorithms for competitive field-testing, a system that is sufficiently robust to withstand failure of any one algorithm.

    The most direct mechanism is a straight commercial one. If the company operating the underwriting algorithm also bears all or some of the underwriting risk, then its commercial success should be directly linked to the "correctness" of the algorithm.

    Where this kind of direct mechanism is not available, then we're looking for feedback mechanisms that simulate this, as closely as possible. Just as the survival of the company using these underwriting services may depend on having access to several competing services, so the survival of the underwriting services themselves may depend on being used by several different insurance companies, with different customer profiles and success criteria. (This reduces the risk that all the customers for your service disappear at the same time, and gives you a better chance to fix problems.)

    The bottom line is that I'm buying an underwriting service rather than an underwriting calculation service, a business service rather than an application service. ASPs will have to rebrand themselves yet again, and present themselves as genuine business process outsourcers.

    This relates to my own contention (to be explored in my forthcoming book on the Component-Based Business) that we should focus on business components that deliver business services, rather than merely information or application services.

    extract from article Are You Being Served? (CBDI Forum Journal, July 2000)