The Cambridge Analytica Scandal and the Power Hidden in Data

One Facebook quiz reached far beyond its users, turning ordinary digital signals into inferred profiles and targeted political messages.

We may earn affiliate commissions from links on this page. Learn more.

Cyber Altitude Guide

Online Safety Starter Kit

Start with the basics for safer accounts, devices and everyday browsing.

Lock down accounts

Strong passwords, a password manager, MFA and email aliases.

Secure devices and files

Block malware, avoid risky downloads and back up important files.

Browse with less exposure

VPN, secure DNS and private browsing tools on risky networks.

In 2014 a personality quiz app collected Facebook information from a few hundred thousand people who installed it, and from tens of millions of their Facebook friends who never did. The data was used to build personality scores, match them against United States voter records and sell voter-profiling and targeted-advertising services. Four years later, in March 2018, the Cambridge Analytica scandal broke into public view and became one of the defining privacy stories of the decade.

It is worth being precise about what happened, because the shorthand version is wrong. Nobody broke into Facebook. The data moved through Facebook’s own developer platform, under permissions the platform granted at the time, and the questions that followed were about how that access was used rather than how it was breached.

The deeper issue was never only who ended up holding the data. It was what the data could be turned into. Ordinary signals, the pages someone had liked and the details on their profile, became raw material for inferred characteristics. Those inferences became profiles, and those profiles became a way of deciding which people saw which political messages. The progression from collection to inference to targeting changed the privacy question from what platforms knew to what they could do with that knowledge.

How Facebook’s Platform Made the Data Possible

Facebook’s developer platform was built to make third-party apps social. When someone installed a quiz, a game or a horoscope app, the app could ask for access to their Facebook information, and the person granted it with a tap. That much is still familiar.

The part that is no longer familiar is what came with it. Under the older version of Facebook’s platform, an app installed by one person could also obtain information connected to that person’s Facebook friends, depending on the permissions requested and the settings those friends had. The friends were never asked. They had no way to know the app existed, because their connection to someone who installed it was the only involvement they had.

So a single installation was never a single transaction. Each one reached outward along the social graph, and on a platform where an ordinary account might have a few hundred connections, the arithmetic ran away quickly. An app with a few hundred thousand users could touch a dataset in the tens of millions without any of those additional people doing anything at all.

Why One Facebook Quiz Reached Far Beyond the Quiz Taker
Under Facebook’s older developer platform, an app could access information connected to a user’s Facebook friends as well as their own, depending on permissions and settings. Those friends were never asked. One installation could therefore reach far beyond one person.

Facebook had already recognised the problem. In April 2014 it announced changes that would end developers’ ability to reach friend data in that way, but applications already on the platform were allowed to keep the older access for roughly another year while they moved across. That transition window is the detail that matters, because it is when the collection happened.

Meanwhile, the underlying idea that made the data valuable was not a secret either. Researchers at the University of Cambridge’s Psychometrics Centre had published work showing that psychological characteristics could be estimated from digital behaviour such as Facebook Likes. The scandal did not begin with an intrusion or an invention. It began with capabilities that already existed inside the platform and in the open research literature.

How Cambridge Analytica Got Facebook Data

Cambridge Analytica was a political data and consulting company connected to the British firm SCL Group, run by chief executive Alexander Nix and selling analytics and targeting services to campaigns. It did not collect the Facebook data itself.

That work was done by Aleksandr Kogan, an academic at the University of Cambridge who set up a separate commercial company called Global Science Research. The naming has caused years of confusion, so it is worth stating plainly: Kogan worked at the university, but Cambridge Analytica had no institutional connection to it, and the Psychometrics Centre has said it never collaborated with Cambridge Analytica, SCL, GSR or the campaigns involved. In early 2014 Kogan’s company began working with SCL and Cambridge Analytica on a United States data project.

How One Facebook Quiz Reached Millions

The vehicle was an app called thisisyourdigitallife, referred to in the FTC’s case as the GSRApp. People who installed it answered a personality survey and granted the app access to their Facebook information, including the pages they had liked and other profile details. Facebook has said roughly 270,000 people downloaded it.

The FTC put direct United States participation at around 250,000 to 270,000 app users, and alleged that the app also collected information associated with 50 to 65 million of their Facebook friends, including at least 30 million identifiable American consumers. The ratio is the whole point. For every person who chose to take a quiz, information connected to roughly two hundred people who did not was drawn into the same dataset.

From Facebook Likes to Personality Scores

The collected data did not simply sit in a database waiting to be read. According to the FTC, it was used to train an algorithm that generated personality scores for app users and their friends, turning observed behaviour into estimated characteristics.

Those scores were then matched against United States voter records. That step is what converted a social dataset into a political instrument, because a personality estimate attached to a name in a voter file is no longer an abstraction. It is an addressable person with an inferred disposition and a known electoral district. The FTC found that Cambridge Analytica used the resulting information in voter-profiling and targeted-advertising services.

The chain is worth holding in one piece, because most summaries collapse it. Facebook data became inferred personality scores, the scores were matched to voter records, and the matched records were used to decide which audiences received which messages. Each link is documented. What each link does not establish is what happened at the far end of it.

How the Cambridge Analytica Scandal Became a Platform Crisis

The core allegations had surfaced before 2018, but they had not yet become a broad platform-accountability crisis. The Guardian reported in December 2015 that Cambridge Analytica was using psychological profiles built from Facebook data connected to millions of people in American political work. Facebook has said it learned around that time that Kogan had passed data to Cambridge Analytica, and that it demanded certifications confirming the data had been deleted.

Then very little happened for more than two years. The facts were on the public record and the story did not take hold, which is its own uncomfortable observation about how a privacy issue becomes a scandal.

What changed was a witness. In March 2018 the Observer published the account of Christopher Wylie, a former Cambridge Analytica employee, along with documents and testimony describing the data project from the inside. Facebook suspended Cambridge Analytica, SCL, Kogan and Wylie from the platform on March 16, and the Observer’s report followed the next day. Its initial figure was more than 50 million profiles, which was what was known at the time rather than the final scope.

Within weeks the story had outgrown the company at its centre. In early April, Facebook published its own estimate that information associated with up to 87 million people, most of them in the United States, may have been improperly shared. On April 10 Mark Zuckerberg testified before a joint Senate hearing on Facebook, social media privacy, and the use and abuse of data. Regulators in Britain and the United States widened their investigations.

By that point the question had shifted. It was no longer only what Cambridge Analytica had done with a dataset. It was what responsibility Facebook carried for an ecosystem it had designed, opened to developers, and policed by asking them to certify their own good behaviour.

What “87 Million” Actually Means

The number became shorthand, and the shorthand is misleading. It is Facebook’s own estimate of the maximum number of accounts whose information may have been improperly shared with Cambridge Analytica. It is a ceiling calculated by the platform, not a confirmed count produced by an investigation.

What “87 Million” Does and Does Not Mean
Facebook estimated that information associated with up to 87 million people may have been improperly shared. It does not mean 87 million people took the quiz, received targeted ads, or had their accounts hacked. The number describes estimated reach, not confirmed manipulation.

It does not mean 87 million people took the quiz, because roughly 270,000 did. It does not mean 87 million people were assigned personality scores, or received political advertising chosen for them, or were persuaded of anything. And it does not describe a hack, because no account was broken into.

What it does describe is reach. Facebook’s estimate is a measure of how far a single app’s permissions could extend across a social network, which is precisely the mechanism the scandal exposed. Keeping the number attached to that meaning, rather than to a claim about manipulation, is the difference between understanding this story and repeating it wrongly.

The Power Was in What the Data Could Infer

The Facebook data taken on its own was not especially dramatic. Pages someone had liked. A city. Some profile fields. Read literally, most of it is the kind of thing people post in public without a second thought, and if the story ended at collection it would be a policy violation rather than a turning point.

What gave the dataset its value was what could be estimated from it. A list of Likes is an observation. A personality score derived from that list is an inference, a claim about someone that they never made about themselves and may not agree with. The distinction matters because the two things behave differently: observed data can be deleted or hidden, while an inference is generated elsewhere, held elsewhere and invisible to the person it describes.

Profiling is the step that organises those inferences into something usable. People get sorted into groups according to what has been observed and what has been estimated, and once sorted they can be counted, selected and addressed. The Cambridge Analytica scandal showed that a privacy risk does not require anyone to read a private message. It only requires enough ordinary signals for a system to make a confident guess.

Microtargeting Changed Who Saw Which Message

Profiling decides what kind of person someone is taken to be. Microtargeting decides what that person sees as a result. The two are routinely discussed as one thing and they are not, and the gap between them is where the political significance sits.

Profiling vs Microtargeting
Profiling: using observed and inferred data to decide what kind of person someone is or which group they belong to.
Microtargeting: using those groups to decide which messages they receive.

In conventional campaigning a message is public. It appears on a billboard, in a broadcast slot or in a newspaper, and anyone can see it, quote it and argue with it. A microtargeted message is delivered to a defined audience and to nobody else, so the people best placed to challenge it are often the people who never see it.

This is not a settled historical concern. Current ICO guidance still treats profiling and online political microtargeting as live data-protection issues, describing how campaigns build on electoral registers by adding factual and inferred attributes in order to aim messages at narrow groups. The practice did not end with the company that made it notorious.

Targeting Is Not the Same as Persuasion

Here the evidence runs out, and it is important to say so clearly. The documented chain ends at targeting. Data was collected, scores were generated, records were matched, audiences were selected and services were sold. What none of that establishes is that anybody changed their mind.

Cambridge Analytica worked for campaigns in the United States, and its own marketing made expansive claims about the power of psychographic methods. Promotional claims are not evidence of effect. Whether that targeting was more persuasive than conventional political advertising has never been demonstrated, and the leap from a company having a capability to that capability deciding an election is a leap this article will not make.

The Brexit question deserves the same discipline, because the common version of it is wrong. The ICO investigated Cambridge Analytica’s relationship with Leave.EU and found preliminary contacts that did not progress into the operational role widely attributed to the company. It found no evidence that Cambridge Analytica or SCL carried out data-analytics work for the referendum campaigns, and said it was impossible to determine whether data techniques used during the referendum changed its result. Work for Vote Leave was carried out by a different firm, AggregateIQ, which is a separate entity and a separate story.

Targeting Does Not Prove Persuasion
Evidence that a campaign used profiling or targeted advertising does not establish how many people were persuaded or whether an election result changed. The documented evidence supports targeting, not a proven electoral effect.

Taken together, the Cambridge Analytica scandal exposed how platform data could be turned into inferred profiles and used to decide who received particular political messages, usually without the people being profiled understanding how the process had started. The significance lies in that imbalance of power and visibility. It does not need an election result attached to it to be serious.

What Changed After the Cambridge Analytica Scandal

Cambridge Analytica itself did not last long. On May 2, 2018, less than seven weeks after the Observer published Wylie’s account, the company announced it was ceasing operations and entering insolvency proceedings. Its collapse was fast enough to feel like a conclusion, and it was not one.

The regulators kept going. The UK Information Commissioner’s Office ran a wide investigation into data analytics across political campaigning, looking at platforms, parties, brokers, analytics firms and universities rather than one company. In February 2019 the House of Commons DCMS Committee published its final report, which examined Facebook’s data-sharing practices alongside Cambridge Analytica and SCL and concluded that Facebook’s policies had helped make the episode possible.

The FTC’s actions came in two stages that are often blurred together. In July 2019 it filed cases involving Cambridge Analytica, Kogan and Nix, with Kogan and Nix settling. In December 2019 the Commission issued its final opinion, finding that Cambridge Analytica had engaged in deceptive practices in connection with the collection of Facebook data. An allegation became a finding, and the distinction is worth preserving.

Facebook Became the Bigger Accountability Question

The more consequential response was aimed at the platform. Facebook tightened developer access, reviewed applications that had obtained large volumes of data and audited older apps still holding it. Those changes addressed the mechanism directly, which the company’s critics noted it could have done years earlier.

The financial reckoning was also Facebook’s. In July 2019 the company agreed to a $5 billion penalty and extensive new privacy-governance requirements, including board-level oversight, to settle FTC allegations that it had violated an earlier privacy order and allowed third-party apps to reach data through users’ friends in ways people did not understand. That penalty belonged to Facebook, not to Cambridge Analytica, and the two are frequently confused.

The UK regulator acted on a smaller scale because it had less to work with. The ICO fined Facebook £500,000, the maximum available under the pre-GDPR regime that applied to conduct from 2014 and 2015. GDPR took effect in May 2018, weeks after the scandal peaked, having been adopted well before it. The scandal did not produce that law. It arrived just in time to make the case for it in public.

What did not change is the part worth ending on. Political advertising continued. Audience targeting continued. Campaigns went on combining voter files with commercial and behavioural data, and profiling remained a normal part of how messages are aimed. The scandal removed a company and tightened a platform. It left the underlying model intact.

Why Inferred Data Still Deserves Attention

The company is gone. The capability is not. Observe behaviour, combine datasets, estimate characteristics, define audiences, deliver messages to them: nothing in that sequence depended on Cambridge Analytica existing, and every part of it is more routine now than it was in 2014.

That is why privacy settings solve only part of the problem. A person can lock down a profile, prune a friend list, delete old posts and switch an account to private, and all of it is worth doing, because it reduces what can be directly observed. None of it reaches the inferences already generated from earlier behaviour, the audience categories an account has been placed in, or the predictions built from data that was never secret in the first place. You can hide the inputs and still be described by the outputs.

The parallels in ordinary digital life are structural rather than dramatic. Advertising profiles, recommendation systems, data brokers and political audience targeting all run on the same basic sequence, as do the behavioural inference systems now built with machine learning. It would be lazy to call any of them the new Cambridge Analytica. What is true is narrower and more useful: as inference gets cheaper and better, the distinction between what a person disclosed and what a system decided about them matters more, not less.

What Users Can See Is Only Part of the Profile

Open the privacy page of any large platform and you can see your posts, your settings, your connections and usually a list of advertisers who have your contact details. It feels like a full accounting and it is not.

What you generally cannot see is the derived layer: which interests have been inferred rather than declared, which audience segments you have been sorted into, what has been predicted about you, which outside datasets your account has been matched against, and why a specific message reached you rather than someone else. The visible profile is the part built from what you entered. The consequential profile is the part built from what you did.

That gap is the Cambridge Analytica lesson in its most ordinary form. The scandal still matters because online privacy is not limited to information people deliberately disclose. Behavioural data can be combined, analysed and used to infer characteristics that determine how platforms, advertisers, campaigns and other organisations classify and target individuals, and the people being classified are rarely in a position to check the result.

What Cambridge Analytica Left Behind

Before the scandal, it was reasonable to think of social-media privacy as a question of disclosure. You decided what to post, who could see it, and how much of yourself to put online, and the risk was that someone would see something you had shared.

What this story made difficult to hold was the idea that the risk stops there. Information collected for one purpose became material for estimating things nobody had disclosed, and those estimates became decisions about which people would see which messages, made by systems the people concerned could not inspect.

Cambridge Analytica’s lasting lesson is that the most powerful information about a person is often the information they never provided at all.