Showing posts with label analytics. Show all posts
Showing posts with label analytics. Show all posts

Friday, January 01, 2021

Does Big Data Drive Netflix Content?

One thing that contributes to the success of Netflix is its recommendation engine, originally based on an algorithm called CineMatch. I discussed this in my earlier post Rhyme or Reason (June 2017).

But that's not the only way Netflix uses data. According to several pundits (Bikker, Dans, Delger, FrameYourTV, Selerity), Netflix also uses big data to create content. However, it's not always clear to what extent these assertions are based on inside information rather than just intelligent speculation.

According to Enrique Dans
The latest Netflix series is not being made because a producer had a divine inspiration or a moment of lucidity, but because a data model says it will work.
Craig Delger's example looks pretty tame - analysing the intersection between existing content to position new content. 

The data collected by Netflix indicated there was a strong interest for a remake of the BBC miniseries House of Cards. These viewers also enjoyed movies by Kevin Spacey, and those directed by David Fincher. Netflix determined that the overlap of these three areas would make House of Cards a successful entry into original programming.

This is the kind of thing risk-averse producers have always done, and although data analytics might enable Netflix to do this a bit more efficiently, it doesn’t seem to represent a massive technological innovation. Thomas Davenport and Jeanne Harris discuss some more advanced use of data in the second edition of their book Competing on Analytics.

Netflix ... has used analytics to predict whether a TV show will be a hit with audiences. ... It has used attribute analysis ... to predict whether customers would like a series, and has identified as many as seventy thousand attributes of movies and TV shows, some of which it drew on for the decision whether to create it.

One of the advantages of a content delivery platform is that you can track the consumption of your content. Amazon used the Kindle to monitor how many chapters people actually read, at what times of day, where and when they get bored. Games platforms (Nintendo, PlayStation, X-Box) can track how far people get with the games, where they get stuck, and where they might need some TLC or DLC. So Netflix knows where you pause or give up, which scenes you rewind to watch again. Netflix can also experiment with alternative trailers for the same content.

In theory, this kind of information can then be used not just by Netflix to decide where to invest, but also by content producers to produce more engaging content. But it's difficult to get clear evidence how much influence this actually has on content creation.

How much other (big) data does Netflix actually collect about its consumers. Craig Delger assumes they operate much like most other data-hungry companies.

Netflix user account data provides verified personal information (sex, age, location), as well as preferences (viewing history, bookmarks, Facebook likes).

 However, in a 2019 interview (reported by @dadehayes), Ted Sarandos denied this.

We don’t collect your data. I don’t know how old you are when you join Netflix. I don’t know if you’re black or white. We know your credit card, but that’s just for payment and all that stuff is anonymized.

Sarandos, who is Chief Content Officer at Netflix, also downplayed the role that data (big or otherwise) played in driving content.

Picking content and working with the creative community is a very human function. The data doesn’t help you on anything in that process. It does help you size the investment. … Sometimes we’re wrong on both ends of that, even with this great data. I really think it’s 70, 80% art and 20, 30% science.

But perhaps that's what you'd expect him to say, given that Netflix has always tried to attract content producers with the promise of complete creative freedom. Amazon Studios has made similar claims. See report by Roberto Baldwin.

While there may be conflicting narratives about the difference data makes to content creation, there are some observations that seem relevant if inconclusive.

Firstly, the long tail argument. The orginal business model for Amazon and Netflix was based on having a vast catalogue, in which most of the entries are of practically no interest to anyone, because the cost of adding something to the catalogue was trivial. Even if the tail doesn't actually contribute as much revenue as the early proponents of the long tail theory suggested, it helps to mitigate uncertainty and risk - not knowing in advance which are going to be hits.

But this effect is countered by the trend towards vertical integration. Amazon and Netflix have gone from distribution to producing their own content, while Disney has moved into streaming. This encourages (but doesn't prove) the hypothesis that there may be some data synergies as well as commercial synergies.

And finally, an apparent preference for conventional non-disruptive content, as noted by Alex Shephard, which is pretty much what we would expect from a data-driven approach.

Netflix is content to replicate television as we know it—and the results are deliberately less than spectacular.

Update (June 2023)

I have been reading a detailed analysis in Ed Finn's book, What Algorithms Want (2017).

Finn's answer to my question about data-driven content is no, at least not directly. Although Netflix had used data to commission new content as well as recommend existing content (Finn's example was House of Cards) it had apparently left the content itself to the producers, and then used data and algorithmic data to promote it. 

After making the initial decision to invest in House of Cards, Netflix was using algorithms to micromanage distribution, not production. Finn p99

Obviously that doesn't say anything about what Netflix has been doing more recently, but Finn seems to have been looking at the same examples as the other pundits I referenced above.


Roberto Baldwin, With House of Cards, Netflix Bets on Creative Freedom (Wired, 1 February 2013)

Yannick Bikker, How Netflix Uses Big Data to Build Mountains of Money (7 July 2020)

Enrique Dans, How Analytics Has Given Netflix The Edge Over Hollywood (Forbes, 27 May 2018), Netflix: Big Data And Playing A Long Game Is Proving A Winning Strategy (Forbes, 15 January 2020)

Thomas Davenport and Jeanne Harris, Competing on Analytics (Second edition 2017) - see extract here https://www.huffpost.com/entry/how-netflix-uses-analytics-to-thrive_b_5a297879e4b053b5525db82b

Ed Finn, What Algorithms Want: Imagination in the Age of Computing (MIT Press, 2017)

FrameYourTV, How Netflix uses Big Data to Drive Success via Inside BigData (20 January 2018) 

Daniel G. Goldstein and Dominique C. Goldstein, Profiting from the Long Tail (Harvard Business Review, June 2006)

Dade Hayes, Netflix’s Ted Sarandos Weighs In On Streaming Wars, Agency Production, Big Tech Breakups, M+A Outlook (Deadline, 22 June 2019)

Alexis C. Madrigal, How Netflix Reverse-Engineered Hollywood (Atlantic, 2 January 2014)

Selerity, How Netflix used big data and analytics to generate billions (5 April 2019)

Alex Shephard, What Netflix’s Obama Deal Says About the Future of Streaming (New Republic 23 May 2018)

Related posts: Competing on Analytics (May 2010), Rhyme or Reason - the Logic of Netflix (June 2017)

Tuesday, October 15, 2019

DataOps - Organizing the Data Value Chain

At #TalendConnect today frequent mention of #DataOps, although according to a post I found on the Talend blog from earlier this year, Talend prefers the term collaborative data management.
Data Preparation ... should be envisioned as a game-changing technology for information management due to its ability to enable potentially anyone to participate. Armed with innovative technologies, enterprises can organize their data value chain in a new collaborative way. Talend
I've always insisted that the data value chain should end not with delivering insight (so-called actionable intelligence) but with delivering business outcomes (actioned intelligence), and I was pleased to hear some of today's speakers making the same point. However, there are still voices within the industry that have a narrower view of DataOps, and I note with concern that the DataOps Manifesto identifies the goal of DataOps in terms of the early and continuous delivery of valuable analytic insights.

Although there will always be a place for analytic reports and dashboards, I always expected that these would gradually make way for analytic insights being rendered as services and integrated into operational business systems and processes, to create closed-loop business intelligence. There are many good examples of this today, especially in the manufacturing world. There are also systems that deliver insights directly to customers or end-users, perhaps in the form of recommendations. But a lot of the discussion of the data-driven enterprise still seems to be based on a dashboard mindset.

And who actually does the DataOps? A presentation from Virtusa showed a three-step DataOps process - pipeline, innovation and value - which suggests a trimodal approach. So the Town Planners would do the pipeline (building generic and highly customizable data preparation frameworks), Pioneers would do the innovation (experimental proof of concept), and the Settlers would roll out the value. I shall be interested to see some practical implementations of this approach.

Meanwhile, simplistic notions of democratization (or citizen integration) often divides people into two camps - experts and citizens - and this polarization is encouraged by Gartner's promotion of Bimodal IT. But this leads people to believe that you can have either trust or speed/agility but not both. And as Jonathan Gill of Talend emphasized in his keynote today, digital leaders don't recognize this dichotomy.




Jean-Michel Franco, 3 Key Takeaways from the 2019 Gartner Market Guide for Data Preparation (Talend, 26 April 2019)

Wikipedia: DataOps

Related posts: Service-Oriented Business Intelligence (September 2005), SPARK 2 Innovation or Trust (March 2006), Analytics for Adults (January 2013), From Networked BI to Collaborative BI (April 2016), Beyond Bimodal (May 2016), Towards the Data-Driven Business (August 2019), Beyond Trimodal - Citizens and Tourists (November 2019)

Thursday, June 29, 2017

Rhyme or Reason - The Logic of Netflix

@GuyLongworth, who teaches philosophy at Warwick, is puzzled by the Netflix recommendation algorithm, linking Annie Hall with Son of Saul.


Philosopher Guy's appeal to rhyme rather than reason seems to be based on the view that the two films have nothing else in common. But this is rather contradicted by the fact that he has actually seen both. Netflix has correctly surmised that people like Guy might possibly be interested in both films.

The first thing to understand about recommendation algorithms is that they are not solely (if at all) based on the intrinsic similarity of two products, but on what we might call relational similarity. If I tell you that people who like pizza also like ice-cream, that is primarily a statement about the "people who like". You might try to explain this statement by observing that pizza and ice-cream both have a high fat content, but then so do lots of other foods.

And when someone has just eaten a pizza, it is perhaps more likely that they will go on to eat ice-cream next, rather than eating another pizza straightaway.



The second thing to understand is that recommendation algorithms work by trial and error. Netflix wants to know if Guy will accept its suggestion to re-watch Annie Hall, and this feedback will add to its knowledge of Guy as well as its knowledge of relational similarity between films.

Trial and error works better if you have a diverse range of trials. If you watch a couple of films in a particular genre, and then Netflix only ever shows you suggestions within that genre, it will never discover that you might be interested in a completely different genre as well. And you will never discover the full range of Netflix offerings, which could result in your abandoning Netflix altogether.

Diversity of suggestion adds to the richness of the experimental data that are generated. How many members of the "people like Guy" category respond positively to suggestion A, and how many to suggestion B? Todd Yellin, Netflix VP of Product, told journalists in March that "we are addicted to the methodology of A/B testing".

What is genre anyway? In the past, genres (in book publishing, music, film, video games) were defined by the industry or by experts. In 2013, Netflix employed over 40 people hand-tagging TV shows and movies. But a data-driven approach allows genres to emerge organically from the patterns of consumption. Netflix (and Amazon and the rest) will be much more interested in data-defined genres than in industry-defined genres.

In her rant against the Netflix algorithm, @mehreenkasana makes two apparently contrary complaints. On the one hand, Netflix offers her content that is nothing like anything she has ever watched. She dismisses one suggestion with the words "I’ve never watched a show in a remotely similar vein." On the other hand, she doesn't see how Netflix can offer her challenging experiences. "Intensely curated experiences, whether you’re looking to explore movies or to meet people to date, remove one of the most critical aspects of a rich experience: risk, as in going out of your comfort zone."

But as @larakiara explains, "personalization is key to ensuring users keep coming back. But there's also the problem of over-personalization, so Netflix has to introduce variants."

Thus we can see Netflix as an embodiment of at least three of @kevin2kelly's Nine Laws of God.
  • Control from the bottom up
  • Maximize the fringes
  • Honor your errors
"A trick will only work for a while, until everyone else is doing it." (Remember Blockbuster.)




Mehreen Kasana, Netflix’s recommendation algorithm sucks (The Outline, 24 March 2017)

Kevin Kelly, Nine Laws of God. Chapter 24 of Out of Control (1994)

Lara O'Reilly, Netflix lifted the lid on how the algorithm that recommends you titles to watch actually works (Business Insider, 26 February 2016)

Janko Roettgers, Netflix Replacing Star Ratings With Thumbs Ups and Thumbs Downs (Variety, 16 March 2017), How Netflix tests Netflix: The story behind the service’s new two-thumbs-up feature (Protocol, 11 April 2022)

Tom Vanderbilt, The Science Behind the Netflix Algorithms That Decide What You’ll Watch Next (Wired, 7 August 2013)

Wikipedia: A/B Testing

Related posts: Competing on Analytics (May 2010), Emergent Similarity (February 2012), The Nature of Platforms (July 2017), Towards the Data-Driven Business (August 2019), Does Big Data Drive Netflix Content? (January 2021)

Wednesday, October 26, 2016

The Shelf-Life of Algorithms

@mrkwpalmer (TIBCO) invites us to take what he calls a Hyper-Darwinian approach to analytics. He observes that "many algorithms, once discovered, have a remarkably short shelf-life" and argues that one must be as good at "killing off weak or vanquished algorithms" as creating new ones.

As I've pointed out elsewhere (Arguments from Nature, December 2010), the non-survival of the unfit (as implied by his phrase) is not logically equivalent to the survival of the fittest, and Darwinian analogies always need to be taken with a pinch of salt. However, Mark raises an important point about the limitations of algorithms, and the need for constant review and adaptation, to maintain what he calls algorithmic efficacy.

His examples fall into three types. Firstly there are algorithms designed to anticipate and outwit human and social processes, from financial trading to fraud. Clearly these need to be constantly modified, otherwise the humans will learn to outwit the algorithms. And secondly there are algorithms designed to compete with other algorithms. In both cases, these algorithms need to keep ahead of the competition and to avoid themselves becoming predictable. Following an evolutionary analogy, the mutual adaptation of fraud and anti-fraud tactics resembles the co-evolution of predator and prey.

Mark also mentions a third type of algorithm, where the element of competition and the need for constant change is less obvious. His main example of this type is in the area of predictive maintenance, where the algorithm is trying to predict the behaviour of devices and networks that may fail in surprising and often inconvenient ways. It is a common human tendency to imagine that these devices are inhabited by demons -- as if a printer or photocopier deliberately jams or runs out of toner because it somehow knows when one is in a real hurry -- but most of us don't take this idea too seriously.

Where does surprise come from? Bateson suggests that it comes from an interaction between two contrary variables: probability and stability --
"There would be no surprises in a universe governed either by probability alone or by stability alone."
--  and points out that because adaptations in Nature are always based on a finite range of circumstances (data points), Nature can always present new circumstances (data) which undermine these adaptations. He calls this the caprice of Nature.
"This is, in a sense, most unfair. ... But in another sense, or looked at in a wider perspective, this unfairness is the recurrent condition for evolutionary creativity."

The problem with adaptation being based solely on past experience also arises with machine learning, which generally uses a large but finite dataset to perform inductive reasoning, in a way that is non-transparent to the human. This probably works okay for preventative maintenance on relatively simple and isolated devices, but as devices and their interconnections get more complex, we shouldn't be too surprised if algorithms, whether based on human mathematics or machine learning, sometimes get caught out by the caprice of Nature. Or by so-called Black Swans.

This potential unreliability is particularly problematic in two cases. Firstly, when the algorithms are used to make critical decisions affecting human lives - as in justice or recruitment systems. (See for example, Zeynap Tufekci's recent TED talk.) And secondly, when preventative maintenance has safety implications - from aeroengineering to medical implants.

One way of mitigating this risk might be to maintain multiple algorithms, developed by different teams using different datasets, in order to detect additional weak signals and generate "second opinions". And get human experts to look at the cases where the algorithms strongly disagree.

This would suggest that we maybe shouldn't be too hasty to kill off algorithms with poor efficacy, but sometimes keep them in the interests of algorithmic biodiversity.  (There - now I'm using the evolutionary metaphor.)



Gregory Bateson, "The New Conceptual Frames for Behavioural Research". Proceedings of the Sixth Annual Psychiatric Institute (Princeton NJ: New Jersey Neuro-Psychiatric Institute, September 17, 1958). Reprinted in G. Bateson, A Sacred Unity: Further Steps to an Ecology of Mind (edited R.E. Donaldson, New York: Harper Collins, 1991) pp 93-110

Mark Palmer, The emerging Darwinian approach to analytics and augmented intelligence (TechCrunch, 4 September 2016)

Zeynap Tufekci, Machine intelligence makes human morals more important (TED Talks, Filmed June 2016)


Related Posts
The Transparency of Algorithms (October 2016)

Wednesday, June 30, 2010

The Analytic Organization

There are several possible ways of defining the term analytic organization.

One way is to define a set of capabilities and working practices loosely called analytics, and use the label analytic organization for an organization that has reached a certain level of collective competence or maturity in these. This is the approach taken by Peter Graham, who recommends the following set of characteristic features.

  • Formal functional role (e.g. Chief Analytics Officer)
  • Enterprise analytics applied across all functions of the organization
  • Shared analytics
  • Understanding the drivers of the business
  • Focus on applying order to information to create knowledge
  • Common platform centrally maintained by analytic functional group

It's often a good idea to have a centre of excellence to promote and coordinate some desired capability, and Peter's centralized approach may make sense for many organizations. However, I'd be reluctant to say that Peter's approach is the only possible approach for building the analytic organization, and I should like to think that some organizations (depending on culture and circumstances) will achieve greater results with a more decentralized approach.

Meanwhile, @joemckendrick talks about an analytic organization in terms of the ability to fully leverage data (Joe McKendrick, 7 Simple Steps to Becoming an Analytic Organization, Insurance Experts' Forum, June 18, 2010). He quotes Christina Colby of CapGemini as identifying three ingredients for using analytics to its full potential as a competitive asset:
  1. treasure-troves of data
  2. the right analytic applications
  3. effective management. 
Picking a software vendor at random for contrast, I found a page on SAS's website which appears to define the analytic organization as one maintaining a large and diverse portfolio of analytic models. See SAS Analytic Advantage for Teradata.

Lex Donaldson goes one step further than everyone else. His book is called The Meta-Analytic Organization, and is based on something he calls Statistico-Organizational Theory, which appears to be a combination of statistics and psychometrics.

Meanwhile, other writers such as Bob Marshall (@flowchainsensei) focus attention on what a so-called analytic organization cannot do. An organization that is dominated by a certain kind of thinking may consequently be incapable of other kinds of thinking.

Some writers use the left-brain / right-brain metaphor, and characterize analytics as exclusively left-brain in character. I'm not convinced by this metaphor, but it may well be true that analytic organizations have characteristic patterns of strength and weakness. So I was especially interested to find a "think piece" from the CIA which calls for

an alternative analysis approach that is more an ongoing organizational process aimed at promoting mindfulness—continuous wariness of analytic failure—than a set of tools that analysts are encouraged to employ when needed.

and concludes that

Intelligence Community analytic organizations need to institutionalize sustained, collaborative efforts by analysts to question their judgments and underlying assumptions, employing both critical and creative modes of thought. For this approach to be effective, significant changes in the cultures and business processes of analytic organizations will be required.

Thus opening up the possibility for analytic organizations of transcending the stereotype.


Warren Fishbein and Gregory Treverton, Rethinking “Alternative Analysis” to Address Transnational Threats, The Sherman Kent Center for Intelligence Analysis, Occasional Papers: Volume 3, Number 2, Oct. ‘04. [For further comments on Treverton's work, see my post Puzzles and Mysteries.]

Peter Graham, Building the analytic organization (Information Management, August 2007).

Bob Marshall, The Marshall Model of Organisational Evolution (October 2010)


Updated 5 May 2020