Here are two claims. "Our school scored highest in the district for student health." "Four out of five students in our class drink enough water daily." Both sound solid. Before you read one more line, ask yourself one question. What would you actually need to check, to know if either one is true — or just sounds true?
Both claims can be accurate and both can still mislead. Data literacy is the discipline that tells the difference. It means reading what the data shows, then asking two formal questions before you believe any claim built from numbers. By the end of this chapter, those two questions have names: operational definition, and sample/baseline.
The second half of the chapter turns the same habit outward. Once you can read a claim honestly, a new question opens up. Was the data behind it even kept safe? Was it used the way the people it came from agreed to? Security keeps data safe. Privacy keeps it used appropriately. Both matter, and they are not the same test. And by the time you finish, you'll be building the habits that keep it that way.
Data literacy is the ability to read, understand, question, and use data. It is not one skill but three, and each sits on top of the last.
Reading is the entry level — you can look at a chart or a table and say what it shows on its face. Questioning is the upgrade. Acting is the goal. A data-literate person does not stop at the number. They ask what it actually measures, how it was collected, and what it leaves out.
Get that habit right, and you can act on a claim. You decide, using the data, without claiming more than the data can carry. Get it wrong, and a perfectly true number still fools you.
The Data Literacy Process Framework runs four steps on any claim built from data. FIND locates the data and its source — where did this number actually come from? READ takes in what is shown: which numbers, which labels, which units.
QUESTION is the step that does the real work — you ask the two formal questions this chapter builds toward. USE is where you decide, but only inside what the data supports. Most claims fail on QUESTION, not on the arithmetic. The numbers add up fine. It is the question nobody asked that lets a misleading claim through.
Take the claim from the opening page: "Students in our school are healthier than the district average." Read cold, it sounds solid. The figure beside this paragraph opens it up, bracket by bracket, at every point where someone made a choice.
Bracket one asks what was actually measured. Bracket two asks who was in the sample, and what it was compared against. Neither question changes a single number in the claim. The data did not change — the claim changed, because the two questions were asked. The right-hand panel rewrites the same claim honestly: still a good result, now precise about what it can and cannot say.
ai-data-literacy.claim-brackets-walk
/50-system/standards/book-look/cell1-a4/adv-gautama.png
Gautama
sarus crane
A claim offered without its grounds asks you to obey, not to understand — so ask it plainly: what was measured, and compared with whom?
The first formal question is: what was the camera pointed at? Its name is the operational definition — the exact specification of what is measured and how.
It names four things. First, the thing you are measuring — for example, student health. Second, the exact indicator that stands in for it — daily water intake in millilitres. Third, the method used to collect it — a self-report survey. Fourth, the time period it covers. Every measurement was a choice — a choice about what to look at, not a neutral fact handed down by the number itself.
Two studies can use the same word, say "health" or "stress," and still measure two different things. Without a stated operational definition, a reader has to guess what the number actually stands for. The person making the claim can then wave away any reading that turns out wrong.
Weaker. Two papers both claim to measure "student stress," so Meera's first glance treats the word as one finding and starts to line the two numbers up side by side. One paper reports a 10-point self-report anxiety scale — what a student is willing to say about how they feel. The other reports cortisol level in saliva — a hormone marker of the body's stress response. Same word on both covers. Two completely different indicators underneath it.
Stronger. Meera writes the two studies as two columns on her sheet before comparing anything: one headed "self-report anxiety scale," the other "cortisol level." Filled in side by side, the columns will not merge into one number — the self-report column is a report of felt stress, the cortisol column a chemical marker of physiological stress. Neither measurement is wrong, and neither is "the real" student stress; they are two different operational definitions wearing the same label. Her verdict is not which one is right — it is that the two studies cannot share one column until she knows which one each number actually is.
"Numbers are objective — they don't lie. Only graphs lie."
Every number came from a measurement, and every measurement was a choice about what to point the camera at. "42 students scored high on a self-report stress survey" and "42 students had a cortisol level under 15 nmol/L" can both get called evidence of student stress, and they measure two different things entirely. A number does not lie on its own — but an unstated operational definition lets its meaning drift away from what the measurement can actually support. Trusting a number more than a graph, simply because it is a number, is exactly the mistake.
The second formal question is: what is outside the frame? Its name is sample and baseline, and the two always travel together.
The sample is who is actually inside the frame — its limits are size, how people were selected, and what it covers. The baseline is what the sample gets compared against. A claim with no named baseline, like "our school does better," is not comparable to anything yet. A claim with a mismatched baseline, a voluntary survey set against a mandatory census, misleads even when both numbers are perfectly accurate.
Get either one wrong and a true number misleads. The sample is the frame's inside edge. The baseline is its outside edge. Both have to hold for the comparison to mean what it claims to mean.
Accurate and fair are two different tests, and a claim can pass one and fail the other every single day.
Accurate is not the same as fair. Accurate means the data was recorded correctly — the numbers themselves are right. Fair means the claim stays inside what the data can actually support. The sample, the baseline, and the operational definition all have to match what the claim implies.
A self-selected sample, a mismatched baseline, or a too-narrow operational definition — any one of these fails the fairness test. Every number can still stay perfectly accurate. "The numbers are right" is where an honest audit begins. It is never where the audit ends.
"If the numbers are accurate, the claim must be fair."
Accuracy and fairness are separate tests, and passing one says nothing about the other. Accuracy asks whether the data was recorded correctly. Fairness asks whether the claim stays inside what that data supports. A self-selected sample, a baseline collected a different way, or an operational definition narrower than the claim implies — each of these fails fairness while every number stays completely accurate. Accuracy is where the audit begins. It says nothing yet about where the claim is allowed to end.
Run both questions on the claim this chapter opened with: "Students in our school are healthier than the district average." Read cold, it sounds like a clean win.
Question ① first. What was actually measured? Water intake, self-reported, over one week. "Healthy" covers far more ground than water intake alone — the operational definition here is narrower than the claim lets on.
Question ② next. Who was in the sample? Forty-two students from one class, and nobody had to take part. What is the baseline? Last year's district figure, collected as a mandatory census — a different method entirely, on a different group.
Verdict: the numbers are accurate. The claim, read as stated, is misleading on both questions. The audit did not destroy the claim — it made it mean something specific.
Weaker. As stated: "Students in our school are healthier than the district average." It sounds broad and impressive — a full measure of health, a representative sample, a matched comparison. The data behind it supports none of those three.
Stronger. Stated honestly: "In a voluntary sample of forty-two students from one class, over one week, our average self-reported water intake beat last year's district census figure — collected by a different method." Less impressive to read. Fully precise. And still a good result.
Take three data claims — a headline, an ad, a claim a classmate makes about their own habits. For each, write a two-question audit row. First, what was actually measured (question ①)? Second, who was in the sample, and what was the baseline (question ②)? Third, reach a verdict: accurate but misleading, or both accurate and fair? Then rewrite the claim honestly. Your output is one audit sheet, three completed rows.
The two questions apply to a table exactly the way they apply to a sentence. A header names a column, say "Attendance %." That header is an unstated operational definition wearing a label. Days present divided by scheduled days? Divided by something else? The header does not say.
Who is missing from the rows is a sample gap — question ②. Show forty-seven rows out of a class of two hundred and ten. The table never tells you which forty-seven, or how they got picked.
None of this is a new table-reading move. It puts a name on what you were already doing when a table felt "off." Use those exact names in an exam answer, or in a real argument with data. The table data does not change — the reading of it does.
Data security and data privacy sound like the same idea. They are not — they ask two different questions, and the exam tests the gap between them directly.
Security asks: is the data safe from unauthorised access, modification, or destruction? Its whole job is keeping data SAFE. Privacy asks a different question. Did the person the data is about agree to how it is used? And who sees it? Its job is keeping data APPROPRIATE.
A system can be secure and still violate privacy. Lock the data down perfectly. No attacker gets in. It can still be used by people who are allowed to see it. But it can be used for a purpose the person it belongs to never agreed to. Being safe and being used appropriately are not the same success.
Weaker. A hospital encrypts every patient record. No outside attacker can read them; there has never been a breach. By the security test, the hospital has done everything right — the data is safe from unauthorised access.
Stronger. That same hospital then shares those records with an insurance company, without telling the patients. The data was secure the entire time. The patients' privacy was violated anyway, because authorised staff used the data for a purpose nobody agreed to. Security stops people without access. Privacy governs what even people with access are allowed to do.
"If my data is stored securely, my privacy is protected."
Security and privacy are distinct, and this is exactly the gap the exam tests. A system can be completely secure — encrypted, no outside access, no breach — and still violate privacy, if the people who are allowed in use the data for a purpose you never agreed to. Security stops people without access. Privacy requires that even people with access use the data only for the agreed purpose. You need both. Neither one covers the other.
A data breach is data accessed, used, or disclosed without authorisation. It happens three ways. Hacking — an outside attacker finds a weak point in the system. Insider misuse — someone who is allowed in uses the data beyond what that permission covers. Accidental exposure — a lost device, a forwarded email, a database left open by mistake.
In every one of the three, the same fact holds. Once data is out, it cannot be taken back. The harm is not only that the data was seen once — it is that it stays outside anyone's control from then on.
Three kinds of harm follow a breach. Financial — fraud, identity theft. Reputational — the information used against the person it belongs to. Safety — location data used to find someone who does not want to be found.
ai-data-literacy.data-breach
Reading a breach
How breaches happen and what they cost
A breach is unauthorised access, use, or disclosure of data. It arrives by one of three routes. Hacking: an outside attacker finds the weak point. Insider misuse: someone with permission goes beyond what it covers. Accidental exposure: a lost device, an email sent to the wrong person, or a database left open. Whichever route, the property that matters is irreversibility: once data is out, no one can call it back.
A school's student records sit exposed because a database was set up without a password. No attacker broke anything — the door was simply left unlocked. The result, accidental exposure, is exactly as permanent as if someone had broken in.
Which matters more for limiting harm — stopping a breach before it happens, or planning the response for when one does: changing passwords, checking for reuse, watching accounts?
Five habits sit at the overlap of data literacy, privacy, and security, and each one cuts real risk.
Verify before acting — check who made a claim, where it came from, and when, before you share it further. Minimise what you share. Say yes on purpose — never by default. If a service does not need your address, it does not get it. Use strong, unique authentication — a different password per account, plus two-factor where it is offered.
Recognise phishing — a real service never asks for your password by message or email. Respond to breaches. If a service you use announces one, change the password. Check where else you reused it, and watch the account. Practice 1 is question ① in habit form. Checking who sent a message, where it came from, and when is the reading-capacity form of that same question, run on a message instead of a measurement.
For years I thought password safety meant one thing: make it harder. Add a number, a symbol, a capital letter, and you were safe. I was wrong about what the actual upgrade is.
Password safety at this level is a system, not a harder word. Unique means one password per account, so a breach at one service cannot open the rest. Two-factor authentication adds a second check: a code sent to your phone. So a stolen password alone is not enough to get in. A password manager holds your unique passwords for you, so you are not the one who has to remember them all.
Each account is a separate risk, not one link in a chain where breaking one brings down the rest. That is the actual upgrade — not a harder word, a different mindset.
This is the third time privacy has shown up in this series, and the rung goes up each time. At G6, privacy was a personal habit — control what you share, watch your footprint. At G8, it became a builder's choice — what the system collects, keeps, and uses.
At G9, the third rung is a reading skill. When you read a data claim, one more question belongs beside the two you already have. Whose data made this number? Did they agree to it being used this way? Every claim about data is quietly a claim about those people — because every dataset was collected from someone.
Zoom the whole framework onto one dataset you already own: your digital footprint. It was collected with certain methods, and it records some things and not others. That is an operational definition, decided by whichever platform is doing the recording.
It covers a sample of your time and your platforms, not all of you, everywhere, always. And it gets compared against something, like other people's footprints or a platform's average user, as a baseline. A claim about you, built from your footprint, is a data claim like any other. The same two questions apply to it.
The footprint tells the collector what they measured. It does not tell them who you are.