Showing posts with label algebra. Show all posts
Showing posts with label algebra. Show all posts

Monday, May 23, 2022

Risk Algebra

In this post, I want to explore some important synergies between architectural thinking and risk management. 


The first point is that if we want to have an enterprise-wide understanding of risk, then it helps to have an enterprise-wide view of how the business is configured to deliver against its strategy. Enterprise architecture should provide a unified set of answers to the following questions.

  • What capabilities delivering what business outcomes?
  • Delivering what services to what customers?
  • What information, process and resources to support these?
  • What organizations, systems, technologies and external partnerships to support these?
  • Who is accountable for what?
  • And how is all of this monitored, controlled and governed?

Enterprise architecture should also provide an understanding of the dependencies between these, and which ones are business-critical or time-critical. For example, there may be some components of the business that are important in the long-term, but could easily be unavailable for a few weeks before anyone really noticed. But there are other components (people, systems, processes) where any failure would have an immediate impact on the business and its customers, so major issues have to be fixed urgently to maintain business continuity. For some critical elements of the business, appropriate contingency plans and backup arrangements will need to be in place.

Risk assessment can then look systematically across this landscape, reviewing the risks associated with assets of various kinds, activities of various kinds (processes, projects, etc), as well as other intangibles (motivation, brand image, reputation). Risk assessment can also review how risks are shared between the organization and its business partners, both officially (as embedded in contractual agreements) and in actual practice.

 

Architects have an important concept, which should also be of great interest for enterprise risk management - the idea of a single point of failure (SPOF). When this exists, it is often the result of a poor design, or an over-zealous attempt to standardize and strip out complexity. But sometimes this is the result of what I call Creeping Business Dependency - in other words, not noticing that we have become increasingly reliant on something outside our control.


There are also important questions of scale and aggregation. Some years ago, I did some risk management consultancy for a large supermarket chain. One of the topics we were looking at was fridge failure. Obviously the supermarket had thousands and thousands of fridges, some in the customer-facing parts of the stores, some at the back, and some in the warehouses.

Fridges fail all the time, and there is a constant processes of inspecting, maintaining and replacing fridges. So a single fridge failing is not regarded as a business risk. But if several thousand fridges were all to fail at the same time, presumably for the same reason, that would cause a significant disruption to the business.

So this raised some interesting questions. Could we define a cut-off point? How many fridges would have to fail before we found ourselves outside business-as-usual territory? What kind of management signals or dashboard could be put in place to get early warning of such problems, or to trigger a switch to a "safety" mode of operation.

Obviously these questions aren't only relevant to fridges, but can apply to any category of resource, including people. During the pandemic, some organizations had similar issues in relation to staff absences.


Aggregation is also relevant when we look beyond a single firm to the whole ecosystem. Suppose we have a market or ecosystem with n players, and the risk carried by each player is R(n). Then what is the aggregate risk of the market or ecosystem as a whole? 

If we assume complete independence between the risks of each player, then we may assume that there is a significant probability of a few players failing, but a very small probability of a large number of players failing - the so-called Black Swan event. Unfortunately, the assumption of independence can be flawed, as we have seen in financial markets, where there may be a tangled knot of interdependence between players. Regulators often think they can regulate markets by imposing rules on individual players. And while this might sometimes work, it is easy to see why it doesn’t always work. In some cases, the regulator draws a line in the sand (for example defining a minimum capital ratio) and then checks that nobody crosses the line. But then if everyone trades as close as possible to this line, how much capacity does the market as a whole have for absorbing unexpected shocks?

 

Both here and in the fridge example, there is a question of standardization versus diversity. On the one hand, it's a lot simpler for the supermarket if all the fridges are the same type, with a common set of spare parts. But on the other hand, having more than one type of fridge helps to mitigate the risk of them all failing at the same time. It also gives some space for experimentation, thus addressing the longer term risk of getting stuck with an out-of-date fridge estate. The fridge example also highlights the importance of redundancy - in other words, having spare fridges. 

So there are some important trade-offs here between pure economic optimization and a more balanced approach to enterprise risk.

Wednesday, March 04, 2020

Economic Value of Data

How far can general principles of asset management be applied to data? In this post, I'm going to look at some of the challenges of putting monetary or non-monetary value on your data assets.

Why might we want to do this? There are several reasons why people might be interested in the value of data.
  • Establish internal or external benchmarks
  • Set measurable targets and track progress
  • Identify underutilized assets
  • Prioritization and resource allocation
  • Threat modelling and risk assessment (especially in relation to confidentiality, privacy, security)
Non-monetary benchmarks may be good enough if all we want to do is compare values - for example, this parcel of data is worth a lot more than that parcel, this process/practice is more efficient/effective than that one, this initiative/transformation has added significant value, and so on.

But for some purposes, it is better to express the value in financial terms. Especially for the following:
  • Cost-benefit analysis – e.g. calculate return on investment
  • Asset valuation – estimate the (intangible) value of the data inventory – e.g. relevant for flotation or acquisition
  • Exchange value – calculate pricing and profitability for traded data items

There are (at least) five entirely different ways to put a monetary value on any asset.
  • Historical Cost The total cost of the labour and other resources required to produce and maintain an item. 
  • Replacement Cost The total cost of the labour and other resources that would be required to replace an item. 
  • Liability Cost The potential damages or penalties if the item is lost or misused. (This may include regulatory action, reputational damage, or commercial advantage to your competitors, and may bear no relation to any other measure of value.) 
  • Utility Value The economic benefits that may be received by an actor from using or consuming the item. 
  • Market Value The exchange price of an item at a given point in time. The amount that must be paid to purchase the item, or the amount that could be obtained by selling the item. 

But there are some real difficulties in doing any of this for data. None of these difficulties are unique to data, but I can't think of any other asset class that has all of these difficulties multiplied together to the same extent.

  • Data is an intangible asset. There are established ways of valuing intangible assets, but these are always somewhat more complicated than valuing tangible assets.
  • Data is often produced as a side-effect of some other activity. So the cost of its production may already be accounted for elsewhere, or is a very small fraction of a much larger cost.
  • Data is a reusable asset. You may be able to get repeated (although possibly diminishing) benefit from the same data.
  • Data is an infinitely reproducible asset. You can sell or share the same data many times, while continuing to use it yourself. 
  • Some data loses its value very quickly. If I’m walking past a restaurant, this information has value to the restaurant. Ten minutes later I'm five blocks away, and the information is useless. And even before this point, suppose there are three restaurants and they all have access to the information that I am hungry and nearby. As soon as one of these restaurants manages to convert this information, its value to the remaining restaurants becomes zero or even negative. 
  • Data combines in a non-linear fashion. Value (X+Y) is not always equal to Value (X) + Value (Y). Even within more tangible asset classes, we can find the concepts of Assemblage and Plottage. For data, one version of this non-linearity is the phenomenon of information energy described by Michael Saylor of MicroStrategy. And for statisticians, there is also Simpson’s Paradox.


The production costs of data can be estimated in various ways. One approach is to divide up the total ICT expenditure, estimating roughly what proportion of the whole to allocate to this or that parcel of data. This generally only works for fairly large parcels - for example, this percent to customer transactions, this percentage to transport and logistics, etc.  Another approach is to work out the marginal or incremental cost: this is commonly preferred when considering new data systems, or decommissioning old ones. We can compare the effort consumed in different data domains, or count the number of transformation steps from raw data to actionable intelligence.

As for the value of the data, there are again many different approaches. Ideally, we should look at the use-value or performance value of the data - what contribution does it make to a specific decision or process, or what aggregate contribution does it make to a given set of decisions and processes. 
  • This can be based on subjective assessments of relevance and usefulness, perhaps weighted by the importance of the decisions or processs where the data are used. See Bill Schmarzo's blogpost for a worked example.
  • Or it may be based on objective comparisons of results with and without the data in question - making a measurable difference to some key performance indicator (KPI). In some cases, the KPI may be directly translated into a financial value. 
However, comparing performance fairly and objectively may only be possible for organizations that are already at a reasonable level of data management maturity.

In the absence of this kind of metric, we can look instead at the intrinsic value of the data, independently of its potential or actual use. This could be based on a weighted formula involving such quality characteristics as accuracy, alignment, completeness, enrichment, reliability, shelf-life, timeliness, uniqueness, usability. (Gartner has published a formula that uses a subset of these factors.)

Arguably there should be a depreciation element to this calculation. Last year's data is not worth as much as this year's data, and the accuracy of last year's data may not be so critical, but the data is still worth something.

An intrinsic measure of this kind could be used to evaluate parcels of data at different points in the data-to-information process. For example, showing the increase of enrichment and usability from 1. to 2. and from 2. to 3., and therefore giving a measure of the added-value produced by the data engineering team that does this for us.
    1. Source systems
    2. Data Lake – cleansed, consolidated, enriched and accessible to people with SQL skills
    3. Data Visualization Tool – accessible to people without SQL skills

If any of my readers know of any useful formulas or methods for valuing data that I haven't mentioned here, please drop a link in the comments.



Heather Pemberton Levy, Why and How to Value Your Information as an Asset (Gartner, 3 September 2015)

Bill Schmarzo, Determining the Economic Value of Data (Dell, 14 June 2016)

Wikipedia: Simpson's Paradox, Value of Information

Related posts: Information Algebra (March 2008), Does Big Data Release Information Energy? (April 2014), Assemblage and Plottage (January 2020)

Wednesday, January 01, 2020

Assemblage and Plottage

John Reilly of @RealTown explains the terms Assemblage and Plottage.
Assemblage is the process of joining several parcels to form a larger parcel; the resulting increase in value is called plottage.

If we apply these definitions to real estate, which appears to be Mr Reilly's primary domain of expertise, the term parcel refers to parcels of land or other property. He explains why combining parcels increases the total value.

However, Mr Reilly posted these definitions in a blog entitled The Data Advocate, in which he and his colleagues promote the use of data in the real estate business. So we might reasonably use the same terms in the data domain as well. Joining several parcels of data to form a larger parcel (assemblage) is widely recognized as a way of increasing the total value of the data.

While calculation of plottage in the real estate business can be grounded in observations of exchange value or use value, calculation of plottage in the data domain may be rather more difficult. Among other things, we may note that there is much greater diversity in the range of potential uses for a large parcel of data than for a large parcel of land, and that a large parcel of data can often be used for multiple purposes simultaneously.

Nevertheless, even in the absence of accurate monetary estimates of data plottage, the concept of data plottage could be useful for data strategy and management. We should at least be able to argue that some course of action generates greater levels of plottage than some other course of action.



By the way, although the idea that the whole is greater than the sum of its parts is commonly attributed to Aristotle, @sentantiq argues that this attribution is incorrect.



John Reilly, Assemblage vs Plottage (The Data Advocate, 10 July 2014)

Sententiae Antiquae, No, Aristotle Didn’t Write A Whole is Greater Than the Sum of Its Parts (6 July 2018)

Related post: Economic Value of Data (March 2020)

Thursday, November 07, 2019

On Magic Numbers - Privacy and Security

People and organizations often adopt a metrical approach to sensemaking, decision and policy. They attach numbers to things, perhaps using a weighted scorecard or other calculation method, and then make judgements about status or priority or action based on these numbers. Sometimes called triage.

In the simplest version, a single number is produced. More complex versions may involve producing several numbers (sometimes called a vector). For example, if an item can be represented by a pair of numbers, these can be used to position the item on a 2x2 quadrant. See my post Into The Matrix.

In this post, I shall look at how this approach works for managing risk, security and privacy.



A typical example of security scoring is the Common Vulnerability Scoring System (CVSS), which assigns numbers to security vulnerabilities. These numbers may determine or influence the allocation of resources within the security field.

Scoring systems are sometimes used within the privacy field as part of Privacy by Design (PbD) or Data Protection Impact Assessment (DPIA). The resultant numbers are used to decide whether something is acceptable, unacceptable or borderline. And in 2013, two researchers at ENISA published a scoring system for assessing the severity of data breaches. Scores less than 2 indicated low severity, scores higher than 4 indicated very high severity.

The advantage of these systems is that they are (relatively) quick and repeatable, especially across large diverse organizations with variable levels of subject matter expertise. The results are typically regarded as objective, and may therefore be taken more seriously by senior management and other stakeholders.

However, these systems are merely indicative, and the scores may not always provide a reliable or accurate view. For example, I doubt whether any Data Protection Officer would be justified in disregarding a potential data breach simply on the basis of a low score from an uncalibrated calculation.

Part of the problem is that these scoring systems operate a highly simplistic algebra, assuming you can break a complex situation into an number of separate factors (e.g. vulnerabilities), and then add them back together with some appropriate weightings. The weightings can be pretty arbitrary, and may not be valid for your organization. More importantly, as Marc Rogers argues (as reported by Shaun Nichols), the more sophisticated attacks rely on combinations of vulnerabilities, so assessing each vulnerability separately completely misses the point.

Thus although two minor bugs may have low CVSS ratings, interaction between them could allow a high severity attack. It is complex, but there is nothing in the assessment process to deal with that, Rogers said. It has lulled us into a false sense of security where we look at the score, and so long as it is low we don't allocate the resources.

One organization that has moved away from the scorecard approach is the Electronic Frontier Foundation. In 2014, they released a Secure Messaging Scorecard for evaluating messaging apps. However, they later decided that the scorecard format dangerously oversimplified the complex question of how various messengers stack up from a security perspective, so they archived the original scorecard and warned people against relying on it.




Nate Cardozo, Gennie Gebhart and Erica Portnoy, Secure Messaging? More Like A Secure Mess (Electronic Frontier Foundation, 26 March 2018)

Clara Galan Manso and Sławomir Górniak, Recommendations for a methodology of the assessment of severity of personal data breaches (ENISA 2013)

Shaun Nichols, We're almost into the third decade of the 21st century and we're still grading security bugs out of 10 like kids. Why? (The Register, 7 Nov 2019)

Wikipedia: Common Vulnerability Scoring System (CVSS)

Related posts: Into The Matrix (October 2015), False Sense of Security (June 2019)

Saturday, January 26, 2013

The Calculus of Cost 2

@remembermytweet and @tetradian explore the Types of Cost (Jan 2013).

Alex identifies a number of different types of cost, which as Tom points out are largely monetary costs. But what is a cost anyway?

An enterprise incurs a great deal of cost. and these can be broken down and classified in various ways. Accountants like to express all costs in monetary terms - so for example,  human effort is translated into labour cost.

But that only works if the enterprise is directly paying for the labour, in the form of wages or contractor bills. Wasting the customers' time doesn't count as a direct cost to the enterprise, although it may well have an indirect cost in terms of customer complaints and lost revenue.

We might also think of anxiety as a cost. As Seth Godin comments in relation to the airline industry, "By assuming that their customer base prefers to save money, not anxiety, they create an anxiety-filled system." Eleven things organizations can learn from airports (Jan 2013). See also my post on Anxiety as a Cost (Jan 2013).

Even for direct labour, the human cost may not be fully reflected by the wages and monetary overheads associated with employment. Many employees do unpaid overtime and incur other personal costs, and this is only visible to the enterprise accountants when it results in high levels of sickness and staff turnover.


So we need to remember the difference between a real cost and its monetary measure.


Wednesday, November 07, 2012

On Business Architecture and Management Accounting

Someone asked on Linked-In Should EA be knowledgeable about accounting and its pitfalls? Here is a rewritten version of my reply.


Management accounting provides a view of the distribution of costs, benefits and risks. Costs include labour costs, materials and other purchases, and so-called overheads. So if we want to know the total production cost of a particular product, or the total cost of a particular activity, we need to work out what share of the total expenditure should be allocated to this product or this activity. These calculations are important for many reasons, including deciding whether to invest in improved systems and technology, deciding whether to keep some capability inhouse or outsource, determining how profitable a given product or service is at a given price and volume.

In One Strategy One PL (January 2013), John R Moran argues that a company can only have a unified strategy if it has a single unified accounting view, and suggests that "the entire logic of profit centers rests on the assumption that maximizing the pieces will maximize the whole". Clearly the relationship between the performance of the pieces and the performance of the whole is an important topic for architects.

Accountants use various simple methods for cost allocation, including so-called Activity-Based Costing. These methods all make some structural assumptions about the dependencies between activities, capabilities, resources and other things. They also involve rules about handling expenditure spanning more than one accounting period. These assumptions also affect project evaluation (e.g. return on investment).

If these structural assumptions are simplistic or incorrect, the management accounting view may lead management to make poor decisions - for example, about investment or cost-cutting or outsourcing. (See my post on Architecture as Jenga.) So business architects need to appreciate what structural assumptions are implicit in the management accounts, and be prepared to challenge these assumptions when necessary.

This is not just because business architects should understand the structure of the business, but also because architecturally-led initiatives may depend on producing a business case that relies on a correct allocation of costs and benefits. For example, architects often wish to advocate long-term investment in shared services and platforms, but such investment may sometimes appear unattractive or unfundable when viewed from a conventional accounting viewpoint, and may be hard to get through the conventional budgeting process. If architects don't understand the potential distortion of the conventional accounting viewpoint, and allow management to take the accounting viewpoint at face value, then they are effectively ceding control of the business structure to the accountants.

Business architects need to pay attention to the structure of cost. Accountants allocate costs according to implicit (and often simplistic) architectural assumptions. For example, accountants use activity-based costing, based on a very simple activity architecture. If left unchallenged, these cost allocation rules can create difficulties for architects in establlshing the business case for shared services and shared infrastructure platforms, and other architecturally-led initiatives. This is one reason why architects need to take on the accountants rather than accept the accountancy view at face value.

See also my post on the Calculus of Cost. This is one of a series of posts on The Purpose of Business Architecture.

See also Tom Graves, Financial-architecture and enterprise-architecture (30 September 2013)



Updated 2 September 2013

Friday, August 07, 2009

The Calculus of Cost

@fbjk2000 (Florian Krueger) poses a set of questions about costs, under the heading Review your calculus of cost. The questions are fair, but the heading is wrong: Florian's questions are all about reviewing cost, and nothing about the calculus of cost.

Cost review looks either at the total cost (is this too much for customers to bear) or at individual cost items (are we paying too much for this item)?

Cost calculus looks at how the total cost is composed from individual cost items, understanding how costs are added and subtracted and multiplied, getting an appropriate balance between fixed costs and variable costs, clustering expenditure to achieve cost synergies, and so on.

Here are some examples
  • In a traditional high-volume non-volatile business, cost reduction can be based on economies of scale, with high fixed costs and low variable costs.
  • In a volatile business, it may make more sense to reduce the fixed costs and accept higher variable costs.
  • Purchasing several different items from the same supplier reduces transaction costs and allows bulk discounts to be negotiated, but may not achieve the lowest price for every item.
  • Allocating costs to a given project can yield widely different results, depending on the formula that is used.
For any reasonably large business or IT operation, questions like these require non-trivial mathematics. And you shouldn't use words like "calculus" unless you are prepared to do that.

Is this just quibbling about semantics? Not at all. Business people may worry about costs, but business analysts and business architects need to pay attention to the structure of costs rather than the tactical details.

Monday, March 17, 2008

Information Algebra

I get more information from two newspapers than from one - but not twice as much information. So how much more, exactly? That depends how much difference there is between the two newspapers. 

Even if two newspapers report the same general facts, they typically report different details, and they may have different sources. To the extent that there are differences in style and detail between the two newspapers, this typically reinforces my confidence in the overall story because it indicates that the journalists are not merely reusing a common source (such as a company press release). 

In the real world, we are accustomed to the fact that information and intelligence needs double-checking and corroboration. And yet in the computer world, there is a widespread belief that it is always a good thing to have a single source of information - that repeated messages are not only unnecessary but wasteful. Data cleansing wipes out difference in the name of consistency and standardization, leaving the resulting information flat and attenuated. A single source of information ("single source of truth") sometimes means a single source of failure - never a good idea in an open distributed system.

Writing about this in an SOA context - when three heads are better than one - Steve Jones describes this as redundancy, and points out the potential value of redundancy to increase reliability. He quotes Lewis Carroll (as Andrew Clarke points out, it was actually the Bellman): "What I tell you three times is true."

The same quote can be found at the head of Chapter 3 of Gregory Bateson's Mind and Nature, available online as Multiple Versions of the World. This expands on Bateson's earlier slogan "Two descriptions are better than one".

Bateson himself used the word "redundancy", but it is not a simple redundancy that can be plucked out without a second thought. Thinking about the consequences of adding and subtracting redundancy is a hard problem - Paulo Rocchi calls it calculus, but I prefer to call it algebra.

Thursday, December 27, 2007

Doubting Events

In my previous post on Cancelling Events, I pointed out that new information sometimes causes us to revise our opinion as to whether a given event had occurred. Marco Seiriö wants to handle my example by adding probability into an event model. According to Marco, you would require that each part of the system can deal with these probabilities.

I can see that probabilistic events could be useful sometimes. However, I am not convinced we need to propagate these probabilities (and the accompanying complexity) throughout the system. I'd prefer to find a way of containing the complexity, so that some parts of the system are presented with a simple binary event statement (either it happened or it didn't) while other parts of the system may be presented with a more complex probabilistic event statement (it might have happened, with probability X%). This is a form of attenuation. - it can be regarded as an application of the need-to-know principle.

This attenuation could be managed architecturally by layering - for example we might separate a process coordination layer (which knows about the probabilities) from a process execution layer (which doesn't).

There is another problem with probabilities - which is that they may change continuously. If a promised action does not appear, the probability that the other person has forgotten increases with time. In the car accident example, if the driver fails to respond under certain conditions, it becomes increasingly likely (but still not certain) that the driver is dead or unconscious. The emergency response parts of the system may need to respond in real-time or near-real-time to this continously shifting probability, but we want to decouple this from other parts of the system that do not have this requirement.

More fundamentally, probability introduces some algebraic challenges that can be solved for relatively simple examples (and perhaps the car accident example is relatively simple) but don't scale for more complex examples (I guess I'll have to construct one).

Wednesday, December 12, 2007

Economic Rent

Can the business case for SOA be based on the concept of economic rent? I'd be interested to get some reaction to the following preliminary discussion ...

Economists use the term "rent" to mean something different to the everyday meaning of the term. It is not the rental income generated by renting out an asset (such as land) but the market power associated with a given market position. For example, let's say that an executive would work for $250,000 per year, but manages to negotiate a salary of $1,250,000 per year. The difference between these two figures ($1m) is a measure of his market power, and this is what economists refer to as economic rent. Economic rent is also earned by professionals with some barrier to entry (such as doctors and lawyers) and by celebrity performers (such as musicians and sportsmen).

There are two alternative definitions of economic rent, yielding somewhat different calculations. In classical economics, it is defined as the difference between the actual income and the necessary income. For example, imagine a popular musician who currently earns around £50,000 per concert. If his popularity waned, his earnings might drop to around £10,000 per concert, but he would continue to play. The difference between these two figures is pure excess - economic rent.

Followers of Pareto use a slightly different definition - the difference between the actual income and the next best opportunity. For example, if our musician wasn't playing concerts, his next best activity might be as a session musician at £4000 per week.

John Kay writes: "economic rent arises from differentiation and migrates to wherever in the value chain scarcity is found". Kay illustrates this with The Story of Champagne. Under perfect market competition, there are no clusters of market power, and no economic rent - prices are adjusted until the level of supply exactly equals the level of demand.

With software and services, of course, the potential income sometimes bears no relationship whatsoever to the costs of production. Most of the profits of Google or Microsoft can be regarded as economic rent in this sense. Some people regard these profits as excessive, while others would justify such profits as reward for past risks and initiatives taken; in any case, the continuation of these profits depends on maintaining some exclusive advantage, something that is difficult for others to replicate.

If a service provider is to invest in developing a new service, then this may be done with the expectation of receiving sufficient reward to cover the (mostly front-loaded) costs and risk. The level of value generated by a service depends on its exclusivity and differentiation: if the service would be easy to copy then there is little prospect of economic rent; but if the service gains a strong foothold, then it may be difficult to dislodge, at least in the short term.

It is for this reason that network effects are so important to the economics of SOA. A service with ten million consumers delivers greater value to each consumer than a similar service with only ten thousand consumers. This advantage enables the provider of the more popular service to extract an economic rent.

Another important consideration is the relationship between fixed development costs (sunk costs?) and ongoing operational costs (opportunity costs?). (If I understand the Wikipedia article correctly, this relationship is one of the areas where classical economics and paretian economics diverge.) SOA and other forms of virtualization (including grid) make it easier to shift costs from fixed to variable - but of course this may not deliver the best economic advantage to the service provider.

If you have some market power, how do you calculate the best possible price? There is some tricky mathematics involved here: let's assume that the more you charge, the fewer users you have, and let's assume that there is a non-zero cost of servicing each user. Under certain conditions, the maximum rent is given by a formula known as Hotelling's Rule.

In case you were wondering, Hotelling has nothing to do with hotels. Harold Hotelling was a mathematical statistician, also known for Hotelling's Law ("in many markets it is rational for producers to make their products as similar as possible") and Hotelling's Lemma (relates the supply of a good to the profit of the good's producer).

Wikipedia: Comparative Advantage Economic Rent, Hotelling's Law, Hotelling's Rule

Saturday, June 02, 2007

SOA Algebra 2

Following my earlier post on SOA Algebra, Andrew Johnston argues that SOA doesn't need a new algebra. He claims "it just requires the right application of techniques we should already understand".

But I never said SOA needed a new algebra. I said it needed an algebra. I expect (indeed hope) that some (perhaps even all) of the elements of this algebra are already known (somewhere). I agree that most (perhaps even all) of the problems are older than SOA.

But the reason I use the word algebra is because I want a systematic and coherent approach (not just a random collection of techniques) to address a set of problems (composition and decomposition) that are currently handled very badly. Most of the people in the SOA world seem to be ignoring or fudging these problems. And there is very little work being done in this area.

Andrew mentions some techniques for calculating reliability (based on probability maths and statistics), and recommends a simple Fault Tree notation (for which he sells a Visio plug-in called RelQuest). He suggests that "service design tools ought to be able to generate the Fault Tree model directly from a graphical service composition". It would also be useful to integrate the Fault Tree model with something that can calculate or simulate the probabilities.

But of course reliability is only one of the things we need to think algebraically about. There are many aspects of service engineering where we need to reason from structure (for example a specific decomposition or collaboration pattern) to quantity (for example, a specific cost/performance envelope). If the structures are simple enough, we may be able to calculate the quantities using simple arithmetic (addition and multiplication), but as soon as the structures get interesting we need to have a more sophisticated algebra.

Friday, January 05, 2007

SOA Algebra

How do we reason about services and SOA? In particular, how do we reason about composition and decomposition?

One important type of reasoning involves the ability to think numerically. For example:
  • If service A costs $x to maintain, and service B costs $y to maintain, how much does it cost to maintain A+B?
  • If service A has an availability of 95%, and service B has an availability of 98%, what is the availability of A+B?
Many people will solve these kind of problems by adding the first and multiplying the second. In very simple situations, addition and multiplication may be good enough. But to deal with more complex situations, we need something called algebra.

Algebra is the branch of mathematics that deals with structure, relation and quantity. The word Al-Jabr is the arabic for reunion - in other words composition. This is therefore the form of reasoning that formalizes and quantifies the composition and decomposition of separate elements. Algebra tells us when it is safe to use addition and multiplication, and when we need something more sophisticated.

For example, if artefact A is composed from components {A1, A2, ..., An}, then the reliability of A is some algebraic function of the reliability of the components A1 to An. Conversely, if you have a requirement for a given level of reliability, you need some algebraic function to decompose this requirement into an equivalent set of sub-requirements.

In an earlier post on Reliability and Availability, I took issue with a vendor white paper, which assumed that all you have to do is multiply the parts to get the whole. This is obviously only true if you make some fairly simplistic assumptions about the composition structure, and is not true in the general case.

Meanwhile, Jeff Schneider takes issue with a confusing formula of SOA costs proposed by Dave Linthicum, which apparently involves adding different kinds of complexity.

If people are getting into these difficulties, where is the SOA algebra to haul them back onto solid ground? I spoke to a friendly professor (Bashar Nuseibeh), who advised me that the formal theory of composition is not advanced enough to support precise quantification. He referred me to some of the work he has been doing with Michael Jackson and others on formalizing composition.

But if I can't have precise, I'll settle for approximate. Surely anything would be better than the algebraic fallacies and muddle currently in circulation?

R. Laney, L. Barroca, M. Jackson, and B. Nuseibeh , Composing Requirements Using Problem Frames , Proceedings of 12th IEEE International Requirements Engineering Conference (RE'04), Kyoto, Japan, 6-10 September 2004 (PDF).

Wikipedia: Algebra

Thursday, November 16, 2006

Reliability and Availability

One of the pleasures of being an industry analyst is that you get to read a lot of vendor material.

Yesterday I came across the following statement in a white paper by Jonathan Purdy of Tangosol on Data Grids and SOA.
"As Business Services are integrated into increasingly complex work-flows, the added dependencies decrease availability. If a Business Process depends on a number of services, the availability of the process is actually the product of the weaknesses of all the composed services. For example, if a Business Process depends on six services, each of which achieves 99% uptime, then the Business Process itself will have a maximum of 94% uptime, corresponding to more than 500 hours of unplanned downtime each year."
This might be true under certain circumstances, but it depends on how the business process is designed, and the degree of coupling within the process. A typical design objective is to compose services in a way that doesn't multiply the dependencies in this way. One of the principles of distributed systems has always been to avoid single points of failure, and this principle is surely inherited by SOA. When used intelligently, loose coupling and asynchrony should make a business more robust.

It is sometimes possible to orchestrate services in a way that that the reliability of the whole is greater than the reliability of the parts. Firstly, there may be underlying services that are not required for every transaction, so the reliability of these underlying services only partially impacts the reliability of the process. Secondly, there may be services that provide multiple or alternative provision of a given capability - alternative process paths can be defined to make the process more fault-tolerant.

That's not going to make the problem of reliability and availability go away of course, and there are undoubtedly some useful things that vendors such as Tangosol can offer in the physical implementation of SOA. And Purdy is right to warn his readers of the risks associated with complexity.

But what this raises for me is the general difficulty of reasoning about non-functional requirements. Do you add them, do you multiply them, do you average them, or do you need to perform a more complicated bit of algebra?


Technorati Tags:

Friday, June 23, 2006

Business Service Architecture - Railway Edition

Martin Geddes draws some interesting parallels between the telecom system/infrastructure and the UK railway system/infrastructure, and discusses some of the effects of economic and political incentives in both systems. He says there are no easy answers, not even easy questions.

I agree about the (lack of) easy answers, but I think there are some important questions. My first question is an architectural one - what is the geometry of the platform stack?

Railway System Business Stack

Getting the stack right isn’t easy. The UK rail system got it disastrously wrong. The rail company (formerly Railtrack, now Network Rail) bought rail maintenance services from engineering companies, and sold rail availability to train operating companies.

But the service level agreements didn’t add up. What service levels do you need from the engineering companies in order to guarantee a given level of rail availability to the train operating companies? And how do you verify that you are getting the required service levels? (These are what I call algebraic problems.)

This is an extremely complex business, for which the rail company lacked the necessary coordination capability. One of the consequences of this incompetence was a serious rail crash, caused by grossly inadequate rail maintenance. But this isn’t just a question of local incompetence. There is a fundamental architectural flaw in the design of the stack as a whole, which fails to tackle some serious and complex questions.

The architectural question here is not just the proper distribution of capability between the layers of the stack, but the bridging between platforms with radically different ontologies - the ontology of rail maintenance is quite different to the ontology of train operation.

So although I agree with Martin about the economic and political problems, I think there are some deeper structural (geometrical, architectural) problems which make it impossible to fix the system merely by tweaking the economics and politics.


Related posts: Business Geometry (September 2004), Railway Edition 2 (August 2006), SOA Algebra (Jan 2007), Services Not All Like Laundry (July 2008)


updated 18 July 2017

Friday, December 30, 2005

REASC

Jean-Jacques Dubray (now with SAP) has posted an interesting SOA pattern on his blog. REASC: a pattern for constructing Composite Applications.

image

This pattern seems to assume a fairly simple event algebra - each event refers to a state-change of a single resource. This appears to restrict the pattern to atomic events.

How can the pattern be extended to support compound events? For example, in building an SOA to support the real-time business, I may want to create BI services that generate compound events. For example, an event may be triggered when the frequency of some transaction exceeds some threshold, or when some new pattern is detected in the data. These compound events might possibly be composed from atomic events, but this may not be the best way to specify them. In any case, I do not want to be forced to define compound (aggregate) resources that correspond to these compound events.

It is possible that JJ intends this kind of event algebra to be contained within the Coordinator. But I should prefer to elaborate the event itself to allow for event composition. This would also allow for amplification and attenuation (as found in Stafford Beer).

I am also interested in exploring the use of the REASC pattern for the service-based business, where resource perhaps equates to business asset. How might we interpret the Coordinator function in service-based B2B collaborations?

Technorati Tags:

Friday, May 06, 2005

Buffering

In many activities, there is a highly variable relationship between the quantity of input and the quantity of output.

For example, Naba Barkakati discusses the nonlinearity of creative endeavors, and argues for buffering as a form of decoupling or desynchronization. Buffering is a way to bridge between an internal world with variable levels of production, and an external world with "linear expectations".

Naba's point doesn't just apply to creative endeavours. It also applies to any activities involving other human beings, such as sales and marketing. On a day-to-day basis, there is no linear relationship between the input (quantity of sales effort, time) and the output (quantity of sales). So sales people (and sales organizations) build in exactly the kind of buffers Naba is talking about, to smooth away the visible peaks and troughs. This may include booking sales in the following period.

Buffers are also used to provide some smoothing between a fixed set of spending budgets, and a variable (and unpredictable) set of spending requirements.

But this mechanism has several negative effects.
  1. It makes it much harder to identify and execute system improvements. A writer may need more active support from an editor or publisher, but the buffers work as a defence mechanism,with the result that the writer doesn't get this support.
  2. The writer (the productive agent) carries more responsibility and takes a greater risk. This may increase the stress on the writer, who may be operating above his/her bearing limit.
  3. The attempted smoothing may be counter-productive, particularly if the buffers get larger and larger, resulting in a small number of major disappointments instead of more frequent minor disappointments.
In many contexts, buffering is regarded as borderline malfeasance. Senior management, investors and regulators can take a particularly dim view of buffering that misrepresents the true state of an operation.

So what's the answer? We may wish to implement various forms of loose coupling and asynchronicity, but this needs to be done within an appropriate governance framework, with shared attention to the economic and ethical behaviour of the whole system.

Technorati Tags:

Monday, June 21, 2004

Beyond Binary Logic

Traditionally, business information systems have operated with a binary boolean logic. For example, some business processing may depend on determining one of two alternative states: e.g. "customer-on-file" or "customer-not-known". 

Open distributed processing (ODP), on which SOA is based, represented a radical challenge to this binary logic – a challenge which the IS world has still not properly embraced. With remote data, there is always a third alternative state: e.g. "customer-file-unavailable".

In distributed systems, perhaps the most obvious reason for the third alternative state was initially the unreliability of telecommunications links. Although the internet acts as a fairly robust medium to link two or more nodes, it also acts as a transmission channel for malware that can take out individual nodes for extended periods. Over the years, system designers have thought of ingenious ways to handle this so-called exception.

But ODP doesn’t just involve geographic distribution – it also distributes over heterogeneous technologies and systems, with possible ontological, semantic and pragmatic impedance.

Trust is an important example. Many systems are designed to discriminate definitively into those individuals (customers, passengers, business partners) that can be trusted, and those that cannot. Mechanisms such as single sign-on and hierarchical trust are called upon to impose binary notions of trust across multiple organizations. Binary trust attempts to suppress the possibility of doubt as to whether someone can/should be trusted.

But mechanisms for trust and security are always vulnerable to false positives (Type 2 errors) and false negatives (Type 1 errors). A system that properly manages these false positives and false negatives entails reintroducing the possibility of doubt. We should generally adopt a critical stance in relation to simplistic binary trust mechanisms (Type 3 errors).

Thus in the open distributed world of SOA, information may always be provisional rather than absolute. For certain purposes, it may be appropriate to construct (impose) a context where closure is viable. With tongue firmly in cheek, we call these Fortresses of Certainty. Outside these fortresses is the land of Maybe, where SOA plays free. 

Related post Complexity and Attenuation (June 2004)