A data dilemma is a genuine conflict of values. One use of data can help some people while harming others, and no option keeps everyone satisfied. That conflict is what makes it a dilemma, with no clean right answer.
Often the benefit and the harm come from the same act. Sharing health records speeds up research while exposing private lives. A camera network makes a street safer while turning it into a place of constant watching.
The syllabus groups tensions into three families: bias, reliability and integrity (is the data fair and trustworthy?); control, ownership and access (who holds it and who may use it?); and privacy, anonymity and surveillance (what does it reveal and who watches?).
Hint
Hint
A dilemma has no clean answer, only better and worse trade-offs.
Analyse one by naming the benefit, the harm, and the competing people and communities.
The concepts of values and ethics and power are your sharpest tools here.
Bias: When Data Encodes Unfairness
Data is collected by people, from some situations and not others, so it can carry the imbalances of the world it came from. When that skew is systematic it is called data bias.
Bias enters via biased collection (over-representing one group) or from the historical record itself. When a system then produces skewed outcomes, that is algorithmic bias.
Example: a large online retailer built a CV-screening tool trained on a decade of mostly male hires. It taught itself to mark down CVs mentioning women's activities and was scrapped. Nothing said to prefer men; it faithfully copied a skewed past.
The dilemma: data can look complete and neutral while working against the people it under-counts. Naming who is missing from the data is the first move in spotting the bias.
Stakeholders feel harm unequally: the organisation sees efficiency and apparent objectivity, while the under-counted group is quietly filtered out with no visible cause, making bias as much a question of power as of accuracy.
Case study
23andMe (case study)
System: 23andMe is a consumer DNA-testing company that stores customers' genetic data and their family-relatives networks (data content topic).
Specifics: in October 2023 a credential-stuffing attack exposed data on about 6.9 million users, largely through the opt-in DNA Relatives feature; the company then filed for bankruptcy in 2025, raising fears its genetic database could be sold as an asset, and regulators urged users to delete their data.
Impacts and implications: highly sensitive, unchangeable genetic data was exposed and its future ownership thrown into doubt; you cannot reset your DNA like a password, and consent given to one company may not survive its sale, so the most personal data carries the highest stakes when control is lost.
Concepts: values and ethics (consent, sensitivity of genetic data), power (who owns and can sell your DNA), systems (one shared feature exposed millions); social and economic contexts.
Reliability and Integrity
Reliability is whether data is accurate and consistent; integrity is whether it stays complete and unaltered from creation to use.
Both break in ordinary ways: human error, system failure losing or duplicating records, and malicious attack that deliberately alters data to deceive. A single corrupted value can travel far before anyone notices.
Organisations invest in validation, backups, and audit trails. The aim is to keep data correct and to be able to prove it stayed that way.
The dilemma is trust at scale: when millions of decisions automate from one dataset, an error repeats everywhere. Data you cannot trust is often worse than no data, because it is acted on with false confidence.
When a benefits system, credit score, or medical record rests on unverified data, the person affected carries the cost of an error they cannot see. That makes integrity an ethical duty as well as a technical one.
Example
Example
A single mistyped date of birth can wrongly deny someone a service for years.
A tampered record in a supply chain can hide where a product really came from.
A lost backup can erase the only proof that something happened.
Control, Ownership, and Access
You generate a constant stream of data about yourself, yet in most cases it is stored, owned, and controlled by the companies whose services you use, not by you.
This creates a deep asymmetry: a platform can collect data with unclear consent, share it with partners you never chose, and store it under other laws, while you have little say. You come away with a working app; the company gains a durable, tradeable asset.
Contested: ownership of data is genuinely disputed. The person it describes, the company that captured it, and regulators all make competing claims. Questions of access (can you see, move, or delete your own data?) are really questions of power.
Data-portability and deletion rights are attempts to shift a little power back toward the individual, but enforcement across borders is hard, so much power still sits with whoever holds the servers.
Case study
Case study: GDPR
System: the European Union's General Data Protection Regulation, a legal response to data-control dilemmas.
Specifics: in force since 2018, it gives people rights to access, correct, port, and delete personal data, and requires clearer consent from organisations that hold it.
Impacts and implications: impact, new obligations on companies worldwide; implication, a partial rebalancing of control toward individuals, though enforcement and reach remain uneven.
Concepts: power (who controls the data) and values and ethics (what a fair claim over data looks like).
Privacy, Anonymity, and Surveillance
Privacy means being able to control what is known about you and by whom, which is more than keeping things secret. Digital systems erode it by collecting far more than people realise.
Anonymity is fragile. Data is personally identifiable information if it can single you out; seemingly anonymous records can be re-identified by combining datasets (a birth date, postcode, and gender can be enough).
Example: a streaming company released supposedly anonymous viewing records for a research contest; researchers matched them to public film ratings and re-identified users. A person's pattern of activity can be as unique as a fingerprint.
The dilemma is that the detailed data that makes services useful also makes true anonymity almost impossible. Treating anonymisation as a guarantee rather than a risk is the misplaced confidence the values and ethics concept warns against.
Dilemmas converge in surveillance, the systematic monitoring of people through their data, run by states for safety and companies for a better service.
Contested: more monitoring can reduce crime and fraud, but the same capability can chill free expression, entrench those in power, and treat whole populations as suspects. Because both benefit and cost are real, reasonable people weigh them differently.
The safeguards decide where the balance lands. Monitoring that is targeted, transparent, and open to accountability keeps costs contained. Monitoring that is broad, secret, and unaccountable slides toward oppression. To judge a given case, look at how targeted, transparent, and accountable the monitoring actually is.
Theory of Knowledge
TOK
How much privacy, if any, is it reasonable to trade for a promised gain in security?
Who should decide where the line sits, and can consent to surveillance ever be truly free?
Active recall
Self review
What makes something a data dilemma rather than a simple right-or-wrong choice?
How does data bias arise, and how does it become algorithmic bias?
Why do reliability and integrity matter so much in automated systems?
Why are ownership of, and access to, personal data contested, and why can anonymised data still identify people?
Why is the privacy-versus-security balance of surveillance genuinely contested?