Showing posts with label privacy. Show all posts
Showing posts with label privacy. Show all posts

Wednesday, December 04, 2019

Data Strategy - Assurance

This is one of a series of posts looking at the four key dimensions of data and information that must be addressed in a data strategy - reach, richness, agility and assurance.



In previous posts, I looked at Reach (the range of data sources and destinations), Richness (the complexity of data) and Agility (the speed and flexibility of response to new opportunities and changing requirements). Assurance is about Trust.

In 2002, Microsoft launched its Trustworthy Computing Initiative, which covered security, privacy, reliability and business integrity. If we look specifically at data, this mean two things.
  1. Trustworthy data - the data are reliable and accurate.
  2. Trustworthy data management - the processor is a reliable and responsible custodian of the data, especially in regard to privacy and security
Let's start by looking at trustworthy data. To understand why this is important (both in general and specifically to your organization), we can look at the behaviours that emerge in its absence. One very common symptom is the proliferation of local information. If decision-makers and customer-facing staff across the organization don't trust the corporate databases to be complete, up-to-date or sufficiently detailed, they will build private spreadsheets, to give them what they hope will be a closer version of the truth.

This is of course a data assurance nightmare - the data are out of control, and it may be easier for hackers to get the data out than it is for legitimate users. And good luck handling any data subject access request!

But in most organizations, you can't eliminate this behaviour simply by telling people they mustn't. If your data strategy is to address this issue properly, you need to look at the causes of the behaviour, understand what level of reliability and accessibility you have to give people, before they will be willing to rely on your version of the truth rather than theirs.

DalleMule and Davenport have distinguished two types of data strategy, which they call offensive and defensive. Offensive strategies are primarily concerned with exploiting data for competitive advantage, while defensive strategies are primarily concerned with data governance, privacy and security, and regulatory compliance.

As a rough approximation then, assurance can provide a defensive counterbalance to the offensive opportunities offered by reach, richness and agility. But it's never quite as simple as that. A defensive data quality regime might install strict data validation, to prevent incomplete or inconsistent data from reaching the database. In contrast, an offensive data quality regime might install strict labelling, with provenance data and confidence ratings, to allow incomplete records to be properly managed, enriched if possible, and appropriately used. This is the basis for the NetCentric strategy of Post Before Processing.

Because of course there isn't a single view of data quality. If you want to process a single financial transaction, you obviously need to have a complete, correct and confirmed set of bank details. But if you want aggregated information about upcoming financial transactions, you don't want any large transactions to be omitted from the total because of a few missing attributes. And if you are trying to learn something about your customers by running a survey, it's probably not a good idea to limit yourself to those customers who had the patience and loyalty to answer all the questions.

Besides data quality, your data strategy will need to have a convincing story about privacy and security. This may include certification (e.g. ISO 27001) as well as regulation (GDPR etc.) You will need to have proper processes in place for identifying risks, and ensuring that relevant data projects follow privacy-by-design and security-by-design principles. You may also need to look at the commercial and contractual relationships governing data sharing with other organizations.

All of this should add up to establishing trust in your data management - reassuring data subjects, business partners, regulators and other stakeholders that the data are in safe hands. And hopefully this means they will be happy for you to take your offensive data strategy up to the next level.

Next post: Developing Data Strategy



Leandro DalleMule and Thomas H. Davenport, What’s Your Data Strategy? (HBR, May–June 2017)

Richard Veryard, Microsoft's Trustworthy Computing (CBDI Journal, March 2003)

Wikipedia: Trustworthy Computing

Thursday, November 07, 2019

On Magic Numbers - Privacy and Security

People and organizations often adopt a metrical approach to sensemaking, decision and policy. They attach numbers to things, perhaps using a weighted scorecard or other calculation method, and then make judgements about status or priority or action based on these numbers. Sometimes called triage.

In the simplest version, a single number is produced. More complex versions may involve producing several numbers (sometimes called a vector). For example, if an item can be represented by a pair of numbers, these can be used to position the item on a 2x2 quadrant. See my post Into The Matrix.

In this post, I shall look at how this approach works for managing risk, security and privacy.



A typical example of security scoring is the Common Vulnerability Scoring System (CVSS), which assigns numbers to security vulnerabilities. These numbers may determine or influence the allocation of resources within the security field.

Scoring systems are sometimes used within the privacy field as part of Privacy by Design (PbD) or Data Protection Impact Assessment (DPIA). The resultant numbers are used to decide whether something is acceptable, unacceptable or borderline. And in 2013, two researchers at ENISA published a scoring system for assessing the severity of data breaches. Scores less than 2 indicated low severity, scores higher than 4 indicated very high severity.

The advantage of these systems is that they are (relatively) quick and repeatable, especially across large diverse organizations with variable levels of subject matter expertise. The results are typically regarded as objective, and may therefore be taken more seriously by senior management and other stakeholders.

However, these systems are merely indicative, and the scores may not always provide a reliable or accurate view. For example, I doubt whether any Data Protection Officer would be justified in disregarding a potential data breach simply on the basis of a low score from an uncalibrated calculation.

Part of the problem is that these scoring systems operate a highly simplistic algebra, assuming you can break a complex situation into an number of separate factors (e.g. vulnerabilities), and then add them back together with some appropriate weightings. The weightings can be pretty arbitrary, and may not be valid for your organization. More importantly, as Marc Rogers argues (as reported by Shaun Nichols), the more sophisticated attacks rely on combinations of vulnerabilities, so assessing each vulnerability separately completely misses the point.

Thus although two minor bugs may have low CVSS ratings, interaction between them could allow a high severity attack. It is complex, but there is nothing in the assessment process to deal with that, Rogers said. It has lulled us into a false sense of security where we look at the score, and so long as it is low we don't allocate the resources.

One organization that has moved away from the scorecard approach is the Electronic Frontier Foundation. In 2014, they released a Secure Messaging Scorecard for evaluating messaging apps. However, they later decided that the scorecard format dangerously oversimplified the complex question of how various messengers stack up from a security perspective, so they archived the original scorecard and warned people against relying on it.




Nate Cardozo, Gennie Gebhart and Erica Portnoy, Secure Messaging? More Like A Secure Mess (Electronic Frontier Foundation, 26 March 2018)

Clara Galan Manso and Sławomir Górniak, Recommendations for a methodology of the assessment of severity of personal data breaches (ENISA 2013)

Shaun Nichols, We're almost into the third decade of the 21st century and we're still grading security bugs out of 10 like kids. Why? (The Register, 7 Nov 2019)

Wikipedia: Common Vulnerability Scoring System (CVSS)

Related posts: Into The Matrix (October 2015), False Sense of Security (June 2019)

Sunday, July 14, 2019

Trial by Ordeal

Some people think that ethical principles only apply to implemented systems, and that experimental projects (trials, proofs of concept, and so on) don't need the same level of transparency and accountability.

Last year, Google employees (as well as US senators from both parties) expressed concern about Google's Dragonfly project, which appeared to collude with the Chinese government in censorship and suppression of human rights. A secondary concern was that Dragonfly was conducted in secrecy, without involving Google's privacy team.  

Google's official position (led by CEO Sundar Pinchai) was that Dragonfly was "just an experiment". Jack Poulson, who left Google last year over this issue and has now started a nonprofit organization called Tech Inquiry, has also seen this pattern in other technology projects.
"I spoke to coworkers and they said 'don’t worry, by the time the thing launches, we'll have had a thorough privacy review'. When you do R and D, there's this idea that you can cut corners and have the privacy team fix it later." (via Alex Hern)
A few years ago, Microsoft Research ran an experiment on "emotional eating", which involved four female employees wearing smart bras. "Showing an almost shocking lack of sensitivity for gender stereotyping", wrote Sebastian Anthony. While I assume that the four subjects willingly volunteered to participate in this experiment, and I hope the privacy of their emotional data was properly protected, it does seem to reflect the same pattern - that you can get away with things in the R and D stage that would be highly problematic in a live product.

Poulson's position is that the engineers working on these projects bear some responsibility for the outcomes, and that they need to see that the ethical principles are respected. He therefore demands transparency to avoid workers being misled. He also notes that if the ethical considerations are deferred to a late stage of a project, with the bulk of the development costs already incurred and many stakeholders now personally invested in the success of the project, the pressure to proceed quickly to launch may be too strong to resist.




Sebastian Anthony, Microsoft’s new smart bra stops you from emotionally overeating (Extreme Tech, 9 December 2013)

Erin Carroll et al, Food and Mood: Just-in-Time Support for Emotional Eating (Humaine Association Conference on Affective Computing and Intelligent Interaction, 2013)

Ryan Gallagher, Google’s Secret China Project “Effectively Ended” After Internal Confrontation (The Intercept, 17 December 2018)

Alex Hern, Google whistleblower launches project to keep tech ethical (Guardian, 13 July 2019)

Casey Michel, Google’s secret ‘Dragonfly’ project is a major threat to human rights (Think Progress, 11 Dec 2018)

Iain Thomson, Microsoft researchers build 'smart bra' to stop women's stress eating (The Register, 6 Dec 2013) 

 

Related posts: Have you got Big Data in your Underwear? (December 2014), Affective Computing (March 2019)

Tuesday, March 20, 2018

Making the World More Open and Connected

Last year, Facebook changed its mission statement, from "Making The World More Open And Connected" to "Bringing The World Closer Together".

As I said in September 2005, interoperability is not just a technical question but a sociotechnical question (involving people, processes and organizations). (Some of us were writing about "open and connected" before Facebook existed.) But geeks often start with the technical interface, or what is sometimes called an API.

For many years, Facebook had an API that allowed developers to snoop on friends' data: this was shut down in April 2015. As Constine reported at the time, this was not just because the API was "kind of shady" but also to "deny developers the ability to build apps ... that could compete with Facebook’s own products". Sandy Paralikas (himself a former Facebook executive) made a similar point (as reported by Paul Lewis): Facebook executives were nervous about the commercial value of data being passed to other companies, and worried that the large app developers could be building their own social graphs.

In other words, the decision was not motivated by concern for user privacy but by the preservation of Facebook's hegemony.

When Tim Berners-Lee first talked about the Giant Global Graph in 2007, it seemed such a good idea. When Facebook launched the Open Graph in 2010, this was billed as "a taste of the future where everything can be more personalized". Like!




Philip Boxer and Richard Veryard, Taking Governance to the Edge (Microsoft Architecture Journal, August 2006)

Josh Constine, Facebook Is Shutting Down Its API For Giving Your Friends’ Data To Apps (TechCrunch, 28 April 2015)

Josh Constine and Frederic Lardinois, Everything Facebook Launched At f8 And Why (TechCrunch, 2 May 2014)

John Lanchester, You Are the Product (London Review of Books, 17 August 2017)

Paul Lewis, 'Utterly horrifying': ex-Facebook insider says covert data harvesting was routine (Guardian, 20 March 2018)

Caroline McCarthy, Facebook F8: One graph to rule them all (CNet, 21 April 2010)

Sandy Parakilas, We Can’t Trust Facebook to Regulate Itself (New York Times, 19 November 2017)

Wikipedia: Giant Global GraphOpen API,


Related Posts SOA Stupidity (September 2005), Social Networking as Reuse (November 2007), Security is Downstream from Strategy (March 2018), Connectivity Hunger (June 2018)

Sunday, March 04, 2018

The Exception That Proves the Rule

My thin clean-shaven friend @futureidentity is reassured by messages that appear to be misdirected.



But when I read his latest tweet, I thought of the exception that proves the rule. Fowler defines five uses of this phrase: I'm going to use two of them.

Firstly, when an advert is exceptionally badly targeted, we notice it precisely because it is an outlier - an exception to the normal pattern or rule. Thus reinforcing our belief in the normal pattern - the idea that many if not most messages nowadays are moderately well targeted. This is what Fowler calls the loose rhetorical sense of the phrase.

Secondly, adverts aren't necessarily misdirected by accident. Conjurers and politicians use misdirection as a form of deception, to distract the audience's attention from what they are really doing. (Some commentators regard the 45th US President as a master of misdirection.)

This is how Target does it, so the pregnant customer doesn't feel she's being stalked.
Then we started mixing in all these ads for things we knew pregnant women would never buy, so the baby ads looked random. We’d put an ad for a lawn mower next to diapers. We’d put a coupon for wineglasses next to infant clothes. That way, it looked like all the products were chosen by chance. (Forbes)

So just because a marketing message appears to be a random error, that doesn't mean it is. Further investigation might reveal it to be carefully designed to foster exactly that illusion in a specific recipient. And if it turns out to be targeted after all, this would be what Fowler calls the secondary rather complicated scientific sense of the phrase.




Related posts



Sources

Charles Duhigg, How companies learn your secrets (New York Times, 16 Feb 2012)

Kashmir Hill, How Target Figured Out A Teen Girl Was Pregnant Before Her Father Did (Forbes, 16 Feb 2012)

Wikipedia: Exception that proves the ruleMisdirection (magic)

Saturday, December 02, 2017

The Smell of Data

Retailers have long used fragrances to affect the customer in-store experience. See for example Air/Aroma.

So perhaps we can use smell to alert consumers to dodgy websites? An artist and graphic designer, Leanne Wijnsma, has built what is basically an air-defreshener: a hexagonal resin block with a perfume reservoir inside, which connects over Wi-Fi to your computer. When it notices a possible data leak (like the user connecting to an unsecured Wi-Fi network, or browsing a webpage over an unsecure connection) — puff! It releases the smell of data.

James Vincent, What does a data leak smell like? This little device lets you find out (Verge, 31 Aug 2017)

That's all very well, but it only sniffs out the most obvious risks. If you want to smell the actual data leak, you'd need a device that released a data leak fragrance when (or perhaps I should say whenever) your employer or favourite online retailer is hacked. Or maybe a device that sniffed around a corporate website looking for vulnerabilities ...

I'm sure my regular readers don't need me to spell out the flaws in that idea.



Related posts

Pax Technica - On Risk and Security (November 2017)
UK Retail Data Breaches (December 2017)

Tuesday, June 27, 2017

Digital Disruption and Consumer Trust - Resolving the Challenge of GDPR

Presentation given to the "GDPR Making it Real" workshop organized by DAMA UK and BCS DMSG, 12 June 2017.

The presentation refers to two milestones. The second milestone is 25th May 2018, the date that companies will need to comply fully with the new data protection regulations. The first milestone is the agreement of a clear and costed plan to reach the second milestone. Some organizations are now getting close to the first milestone, while others still don't have much idea how much effort and resource will be required, or how this could affect their business. Good luck with that. Let me know if I can help.


Tuesday, April 25, 2017

Uber's Self-Defeat Device

Uber's version of "rational self-interest" has led to further accusations of covert activity and unfair competitive behaviour. Rival ride company Lyft is suing Uber in the Californian courts, claiming that Uber used a secret software program known as "Hell" to invade the privacy of the Lyft drivers, in violation of the California Invasion of Privacy Act and Federal Wiretap Act.

This covert activity, if proven, would go way beyond normal competitive intelligence, such as that provided by firms like Slice Intelligence, which harvests and interprets receipts from consumer email. (Slice Intelligence has confirmed to the New York Times that it sells anonymized data from ride receipts from both Uber and Lyft, but declined to say who purchased this data.)

It has also transpired that Apple caught Uber cheating on the iPhone app, including fingerprinting and continuing to identify phones after the app was deleted, in contravention to App Store privacy guidelines. Uber CEO Travis Kalanick got a personal reprimand from Apple CEO Tim Cook, but the iPhone app remains on the App Store, and Uber continues to use fingerprinting worldwide.

Uber continues to be massively loss-making, and the mathematics remain unfavourable. So the critical question for the service economy is whether firms like Uber can ever become viable without turning themselves into defacto monopolies, either by political lobbying or by covert action.




Megan Rose Dickey, Uber gets sued over alleged ‘Hell’ program to track Lyft drivers (TechCrunch, 24 April 2017)

Mike Isaac, Uber's CEO plays with fire (New York Times, 23 April 2017)

Andrew Liptak, Uber tried to fool Apple and got caught (The Verge, 23 April 2017)

Andrew Orlowski, Uber cloaked its spying and all it got from Apple was a slap on the wrist (The Register, 24 Apr 2017)

Olivia Solon and Julia Carrie Wong, Hell of a ride: even a PR powerhouse couldn't get Uber on track (Guardian, 14 April 2017)


Related Posts

Uber Mathematics (Nov 2016) Uber Mathematics 2 (Dec 2016) Uber Mathematics 3 (Dec 2016)
Uber's Defeat Device and Denial of Service (March 2017)

Tuesday, October 25, 2016

85 Million Faces

It should be pretty obvious why Microsoft wants 85 million faces. According to its privacy policy
Microsoft uses the data we collect to provide you the products we offer, which includes using data to improve and personalize your experiences. We also may use the data to communicate with you, for example, informing you about your account, security updates and product information. And we use data to help show more relevant ads, whether in our own products like MSN and Bing, or in products offered by third parties. (retrieved 25 October 2016)
Facial recognition software is big business, and high quality image data is clearly a valuable asset.

But why would 85 million people go along with this? I guess they thought they were just playing a game, and didn't think of it in terms of donating their personal data to Microsoft. The bait was to persuade people to find out how old the software thought they were.

The Daily Mail persuaded a number of female celebrities to test the software, and printed the results in today's paper.

Talking of beards ...




Kyle Chayka, Face-recognition software: Is this the end of anonymity for all of us? (Independent, 23 April 2014)

Chris Frey, Revealed: how facial recognition has invaded shops – and your privacy (Guardian, 3 March 2016)

Rebecca Ley, Would YOU  dare ask a computer how old you look? Eight brave women try out the terrifyingly simple new internet craze (Daily Mail, 25 October 2016)


Related Post: Another 20 million faces (January 2018)


TotalData™ is a trademark of Reply Ltd. All rights reserved

Sunday, August 07, 2016

Why does my bank need more personal data?

I recently went into a High Street branch of my bank and moved a bit of money between accounts. I could have done more, but I didn't have any additional forms of identification with me.

At the end, the cashier asked me for my nationality. British, as it happens. Why do you want to know? The cashier explained that this enabled a security control: if I ever bring my passport into a branch as a form of identification, the system can check that my passport matches my declared nationality.

Really? Really? If this is really a security measure, it's a pretty feeble one. Does my bank imagine I'm going to say I'm British and then produce a North Korean passport? Like a James Bond film?

After she had explained how the bank would use my nationality data, she then asked for my National Insurance number. I declined, choosing not to quiz her any further, and left the branch planning to write a stiff letter to the head of data protection at the bank's head office.

As a data expert, I am always a little suspicious of corporate motives for data collection. So the thought did occur to me that my bank might be planning to use my personal data for some purpose other than that stated.

Of course, my bank is perfectly entitled to collect data for marketing purposes, with my consent. But in this case, I was explicitly told that the data were being collected for a very narrowly defined security purpose.

So there are two possibilities. Either my bank doesn't understand security, or it doesn't understand data protection. (Of course there will be individuals who understand these things, but the bank as an organization appears to have failed to embed this understanding into its systems and working practices.) I shall be happy to provide advice and guidance on these topics.



Saturday, June 04, 2016

As How You Drive

I have been discussing Pay As You Drive (PAYD) insurance schemes on this blog for nearly ten years.

The simplest version of the concept varies your insurance premium according to the quantity of driving - Pay As How Much You Drive. But for obvious reasons, insurance companies are also interested in the quality of driving - Pay As How Well You Drive - and several companies now offer a discount for "safe" driving, based on avoiding events such as hard braking, sudden swerves, and speed violations.

Researchers at the University of Washington argue that each driver has a unique style of driving, including steering, acceleration and braking, which they call a "driver fingerprint". They claim that drivers can be quickly and reliably identified from the braking event stream alone.

Bruce Schneier posted a brief summary of this research on his blog without further comment, but a range of comments were posted by his readers. Some expressed scepticism about the reliability of the algorithm, while others pointed out that driver behaviour varies according to context - people drive differently when they have their children in the car, or when they are driving home from the pub.

"Drunk me drives really differently too. Sober me doesn't expect trees to get out of the way when I honk."

Although the algorithm produced by the researchers may not allow for this kind of complexity, there is no reason in principle why a more sophisticated algorithm couldn't allow for it. I have long argued that JOHN-SOBER and JOHN-DRUNK should be understood as two different identities, with recognizably different patterns of behaviour and risk. (See my post on Identity Differentiation.)

However, the researchers are primarily interested in the opportunities and threats created by the possibility of using the "driver fingerprint" as a reliable identification mechanism.

  • Insurance companies and car rental companies could use "driver fingerprint" data to detect unauthorized drivers.
  • When a driver denies being involved in an incident, "driver fingerprint" data could provide relevant evidence.
  • The police could remotely identify the driver of a vehicle during an incident.
  • "Driver fingerprint" data could be used to enforce safety regulations, such as the maximum number of hours driven by any driver in a given period.

While some of these use cases might be justifiable, the researchers outline various scenarios where this kind of "fingerprinting" would represent an unjustified invasion of privacy, observe how easy it is for a third party to obtain and abuse driver-related data, and call for a permission-based system for controlling data access between multiple devices and applications connected to the CAN bus within a vehicle. (CAN is a low-level protocol, and does not support any security features intrinsically.)


Sources

Miro Enev, Alex Takakuwa, Karl Koscher, and Tadayoshi Kohno, Automobile Driver Fingerprinting Proceedings on Privacy Enhancing Technologies; 2016 (1):34–51

Andy Greenberg, A Car’s Computer Can ‘Fingerprint’ You in Minutes Based on How You Drive (Wired, 25 May 2016)

Bruce Schneier, Identifying People from their Driving Patterns (30 May 2016)

See also John H.L. Hansen, Pinar Boyraz, Kazuya Takeda, Hüseyin Abut, Digital Signal Processing for In-Vehicle Systems and Safety. Springer Science and Business Media, 21 Dec 2011

Wikipedia: CAN bus, Vehicle bus


Related Posts

Identity Differentiation (May 2006)

Pay As You Drive (October 2006) (June 2008) (June 2009)

Sunday, October 25, 2009

Towards an Architecture of Privacy

@futureidentity (Robin Wilton) posted some interesting ideas about Identity versus attributes on his blog.

"For an awful lot of service access decisions, it's not actually important to know who the service requester is - it's usually just important to know some particular thing about them. Here are a couple of examples:

  • If someone wants to buy a drink in a bar, it's not important who they are, what's important is whether they are of legal age;
  • If someone needs a blood transfusion, it's more important to know their blood type than their identity."

However, there is an important difference between Robin's two examples. Blood transfusion is a transaction with longer-lasting consequences. If a batch of blood is contaminated, there seems to be is a legitimate regulatory requirement to trace forwards (who received this blood) and backwards (who donated this blood), in order to limit the consequences of this contamination event and to prevent further occurrences.

There is a strong demand for increasing traceability. In manufacturing, we want to trace every manufactured item to a specific batch, and associate each batch with specific raw materials and employees. In food production, we want to trace every portion back to the farm, so that salmonella outbreaks can be blamed on the farmer. See Information Sharing and Joined-Up Services 1, 2.

Transactions that were previously regarded as isolated ones are now increasingly joined-up. The eggs that go into the custard tart you buy in the works canteen used to be anonymous, but in future they won't be. See Labelling as Service 1, 2.

There is also a strong demand for increased auditability. So it is not enough for the barman to check the drinker's age, the barman must keep a permanent record of having diligently carried out the check. It is apparently not enough for the hotel or bank clerk to look at my passport, they must retain a photocopy of my passport in order to remove any suspicion of collusion. (The bank not only mistrusts its customers, it also mistrusts its employees.)

There is a large (and growing) class of situations where so-called joined-up-thinking seems to require the negation of privacy. I am certainly not saying that this reasoning should always trump the needs of privacy. But privacy campaigners need to understand that all transactions belong within some system of systems, and that this provides the context for the forces they are battling against, rather than pretending that transactions can be regarded as purely isolated events. The point is that authorization is not an isolated event, but is embedded in a larger system, and it is this larger system that apparently requires greater disclosure and retention.

@j4ngis asks how long chains to use for traceability. What "length" of traceability is sound and meaningful? How do we connect all these traces? And also backward and forward in the "chain". For how long should records be kept?
  • Should we also know the batch number for the food that was given to the chicken that laid the egg you included in the cake?
  • Do we have to know the identity of the blood donor after six months? 10 years? 100 years?
The trouble is that there is no rational basis for drawing the line. It is always possible that some contamination in the chicken feed might affect the eggs and thereby the custard tart. It is always possible that the hyperactivity of certain schoolchildren, or the testosterone levels of certain adults, might be traced back to some contamination in the food chain. It is always possible that some obscure data correlation might one day save lives or protect children. And given the vanishing costs of data management, even a faint possibility of future benefit appears to provide sufficient reason for collecting and storing the data.

Robin clearly supposes that attribute-based authorization is a "Good Thing". I am sympathetic to this view, but I don't know how this view can stand up against the kind of sustained attack from a certain flavour of joined-up systems thinking that can almost always postulate the possibility (however faint) of saving lives or protecting children or catching criminals, if only we can retain everything and trace everything.

For my part, I have a vague desire for anonymity and privacy, a vague sense of the harm that might come to me as a result of breaches to my privacy, and a surge of annoyance when I am required to provide all sorts of personal data for what I see as unreasonable purposes, but I cannot base an architecture on any of these feelings.

Traditional arguments for data protection may seem to be merely rearguard resistance to integrated and joined-up systems. Traditional architectures for data protection look increasingly obsolete. But what alternatives are there?


Update May 2016

Traceability requirements for Human Blood and Blood Components are specified in Directive 2005/61/EC of the European Parliament and of the Council 30 September 2005 (pdf - 63KB)

Robin's point was that blood type was more important than identity, and of course this is true. Donor and recipient identity must be retained for 30 years, but that doesn't mean sharing this information with everybody in the blood supply chain.

Friday, April 11, 2008

Services Like Laundry

Just over a year ago, fed up with the over-use of the Lego brick metaphor, my colleague David Sprott proposed the laundry model of SOA (Explaining SOA to the Business). Antony Reynolds of Oracle quoted this in his blog today (The Laundry Model of SOA), which is what reminded me.

Both David and Antony talk about the contract basis of the laundry model.
  • I delegate something (in this case cleaning) to the laundry service.
  • Officially I don't care how it's done (although the term "dry cleaning" seems to indicate a preferred implementation).
  • The laundry service accepts some (limited) liability for any failure.
These are typical characteristics of pretty much any kind of service, more or less. But there is a particular characteristic of the laundry service I want to talk about.
  • I entrust the service with something (an asset) that belongs to me.
  • I expect the laundryman to be discrete. I do not want my dirty linen washed in public - in other words, I do not expect the laundryman to draw any inferences from the state of my clothes, or to pass these inferences to inquisitive journalists.
  • If something goes wrong with the service, then my asset is damaged. The potential liability is linked to the value of the asset, not to the value of the service. (If my dress suit is ruined, I am not going to be happy if I just get the cost of the cleaning reimbursed.)
  • I do not authorize the service provider to use the asset for his own purposes. (I do not expect the laundryman to borrow my suit for his daughter's wedding.)
These characteristics are common to many services. I do not expect my CRM provider to try and sell things to my customers, nor to allow identity thieves to access their information. (Indeed, the customer information may not belong to me either; it arguably belongs to the customers themselves, who have entrusted me with their information according to the same pattern.) I do not expect my ISP to read my email, or interpret my search history. (Perhaps I'm being naive about that.)

However, there are many services where this is not clearly agreed. Perhaps I get some cut-price service, and this is only economically viable because the service provider expects to be able to broadcast some advertising to me or my customers. But if I am careless with the small print on the contract, I may not have understood this aspect of the deal.

And there are many services where this simply doesn't apply. For example, information services often simply involve transferring information from the service provider to the consumer, while Software-as-a-Service may simply involve transferring functionality. Obviously there are trust issues here as well - I trust the BBC to provide accurate and balanced news, I trust my internet security provider to protect me from viruses, I trust Microsoft Excel to calculate my profits correctly - but these are not linked in quite the same way to specific assets.

So I think the laundry metaphor is a very useful one, but like any metaphor we must be careful not to push it too far. All services are a bit like laundry, and some services are very much like laundry, but few services are totally like laundry.


See also

Services Not All Like Laundry (July 2008)
Laundry as Intelligence (Oct 2008)
Understanding Business Services (November 2012)

Thursday, March 27, 2008

Heathrow Terminal 5

Going Live

"... technical difficulties ... staff familiarisation ... teething problems ... technical defect ... brief system fault ... a few minor problems ... time to bed down ..."
A complex bundle of services relocated into a purpose-built space. Mutual recriminations between the collaborating parties (airport and airline). What lessons for SOA design, implemention and deployment here?

Identity, Privacy and Security

"BAA has had to drop controversial plans to fingerprint domestic passengers after the information commissioner expressed concerns about the move. The airport operator said fingerprinting was needed for border security. Instead it will take photographs while the proposal is discussed with the commissioner's office."
Some aspects of the design not approved by key stakeholders and regulators. Apparently the reason for this controversial requirement is that domestic and international travellers are mixed into the same space, rather than being kept separate, and this creates additional security risks. It was assumed that technology (fingerprinting or photography) would substitute for physical barriers. Perhaps BAA chiefs have been reading articles on "deperimeterization".

Meanwhile, one technical solution is quickly substituted for another, although I'm not sure I understand why taking (and presumably transmitting and storing) photographs of passengers is any less of an invasion of privacy than taking fingerprints.

Sources

BBC News, March 27th 2008

Related Posts 

Service-Oriented Security (August 2006)
Travel Hopefully (March 2008)
For Whose Benefit Are Airports Designed? (January 2013)

Friday, January 20, 2006

DoJ Search Requests

cross-posted from Asymmetric Design blog 

Various people talking about the dispute between Google and the US Department of Justice (DoJ), including Boing Boing, Search Engine Watch, Emergent Chaos, and my old colleague Graham Shevlin.

One line of discussion relates to privacy. Some commentators are praising Google for resisting the US Government’s demands for data, when its competitors have apparently complied.

But there is another important consideration, which appears relevant to Google’s position. There is an huge gap (asymmetry) between the information requirement (as stated by the DoJ) and the data on Google’s database. (See especially Danny Sullivan’s post at SearchEngineWatch.)

So the DoJ is hoping to solve a complex problem with very large amounts of data? Does it really make sense for the DoJ to copy the raw data from Google onto its own processors? In a service-oriented grid-enabled world, it would seem to make more sense (and raise fewer privacy concerns as well) for the DoJ to collaborate with Google (and its competitors) - to compose intelligent and relevant analytical enquiries that can be run by Google (as a service, albeit commandeered by the Government) to help solve the DoJ’s problem.

Of course Google is an interested party in the outcome of the DoJ’s deliberations, but does engaging it as a trusted partner in the analysis really increase its ability to bias the outcome in a self-interested way? And if it provides knowledge and metadata services rather than raw data, this might mitigate the threat to its position as a trusted custodian of personal search records?

Tuesday, December 13, 2005

Data Ownership and Trust

This month, I've been looking at some complicated questions of data provenance, data protection and copyright, prompted by a tricky but fascinating client problem - how to convert a legacy archive into a network of information services. (Among other things, this involves dividing material according to the ownership of the content.)

So I was particularly interested to see the following three separate items appear in my blogreader today.


1. Identity Theft and Brand Damage

A UK charity had its donor list stolen by a hacking gang, which then proceeded to beg funds from the same donors. Source: Silicon.com via Emergent Chaos

This is being described as a security breach


2. Software as a Service

"If information about you is stored on your own computer, it's generally not available to others unless they are able to hack your machine or serve legal process on you. In contrast, if information about you is stored on Google's computers, the law generally treats it as Google's, not yours." Cindy Cohn via Tecosystems

This is being described as a privacy issue.


3. Platforms and Stacks

Alexa (part of Amazon) is exposing its index for commercial reuse, via a series of web services. Source: Jon Battelle via Simon Bisson

This is being described as a ground-breaking innovation


There are undoubtedly new business risks that emerge whenever we make a significant change in platform. (That's not to say we shouldn't change, merely that we need to do it with our eyes open. As Stephen O'Grady puts it, "the point here is not to be alarmist, but rather to build awareness".) The new technologies of interaction carry the potential of new forms of sociotechnical intimacy, which may take a little getting used to.

Most importantly, sociotechnical shifts like these may cause us to rethink whether we really own the data (or knowledge) we thought we owned. If an email platform can use email content to target advertising, if a communication platform can analyse message traffic to identify friendship clusters, what else is fair game?

Ultimately this comes down to an important strategic choice. Do we want intimate relationships with intelligent service providers, who can interpret (and customize) both content and context to provide deeper service value? Or do we want arms-length relationships with service providers that don't know us from Adam? Where does the platform stop and the true service begin?