Read this as a study guide instead

Exploratory analysis

Interviewing a dataset before you trust it: look first, model later.

Point-cloud shape
Overlay the best-fit line
yxr = 0.82

Code with its results

A notebook-style walk through the idea — every output shown is the real result of the code above it.

A five-minute interview with one small dataset

A look-first pass over one tiny in-memory dataset using only the standard library: shape, summary stats, an integrity scan, an outlier, and a correlation caveat, before any modeling.

Real output Every Out block below was produced by running the code above it. You can copy these cells into your own python3 and run them top to bottom to see the same numbers.

Before trusting a dataset, you interview it. This is a short look-first pass over a tiny table held in memory as a list of dicts, using only the standard library, so no data libraries are assumed. Nothing here builds a model; it only decides whether a model would stand on solid ground.

In [1]
rows = [
    {"id": 1, "day": "Mon", "temp": 18, "cups": 30},
    {"id": 2, "day": "Tue", "temp": 20, "cups": 33},
    {"id": 3, "day": "", "temp": 21, "cups": 35},
    {"id": 4, "day": "Thu", "temp": 23, "cups": 40},
    {"id": 5, "day": "Fri", "temp": 25, "cups": 44},
    {"id": 6, "day": "Sat", "temp": 200, "cups": 50},
    {"id": 2, "day": "Tue", "temp": 20, "cups": 33},
    {"id": 7, "day": "Sun", "temp": 27, "cups": 48},
]
print("rows:", len(rows))
print("columns:", sorted(rows[0].keys()))
for row in rows[:2]:
    print(row)
Out [1]
rows: 8
columns: ['cups', 'day', 'id', 'temp']
{'id': 1, 'day': 'Mon', 'temp': 18, 'cups': 30}
{'id': 2, 'day': 'Tue', 'temp': 20, 'cups': 33}

Shape first, then summarize. Where does the temp column center, and how far does it spread? The center and spread are a quick fingerprint of a column, and an impossible minimum or maximum shows up here before you have plotted anything.

In [2]
import statistics
temps = [row["temp"] for row in rows]
print("count", len(temps))
print("mean", statistics.mean(temps))
print("median", statistics.median(temps))
print("min", min(temps), "max", max(temps))
Out [2]
count 8
mean 44.25
median 22.0
min 18 max 200

The integrity pass hunts three classic defects: a missing value, a duplicate row, and an impossible value. It does not fix anything; it just flags candidates so you know how much to trust what follows.

In [3]
seen = set()
for row in rows:
    if row["day"] == "":
        print("missing day: id", row["id"])
    if row["id"] in seen:
        print("duplicate row: id", row["id"])
    seen.add(row["id"])
    if row["temp"] > 60 or row["temp"] < -40:
        print("impossible temp: id", row["id"], "value", row["temp"])
Out [3]
missing day: id 3
impossible temp: id 6 value 200
duplicate row: id 2

The mean of 44.25 sat far above the median of 22.0, which is the signature of an outlier: the 200 reading. Because the mean adds up every value, one extreme number drags it, while the median barely moves. Drop the impossible reading and watch which one changes.

In [4]
normal = [t for t in temps if t <= 60]
print("kept", len(normal), "of", len(temps))
print("mean without the 200 reading", statistics.mean(normal))
print("median without the 200 reading", statistics.median(normal))
Out [4]
kept 7 of 8
mean without the 200 reading 22
median without the 200 reading 21

Now look at what moves together. A correlation near 1 means two columns rise and fall in step, but co-movement is not proof of cause: a hidden factor, reversed cause, or coincidence can produce it, and a single outlier can inflate the number on its own.

In [5]
cups = [row["cups"] for row in rows]
r = statistics.correlation(temps, cups)
print("temp-cups correlation", round(r, 2))
Out [5]
temp-cups correlation 0.62

In five minutes the table confessed a missing field, a duplicate, an impossible reading, an outlier that moved the mean but not the median, and a correlation worth a second look. None of that is a conclusion. It is a short list of hypotheses to test later, which is exactly what a look-first pass is supposed to hand the next step.

The same ideas, as prose

These are the exact fragments the model serves — also available as an ordered study guide.

Look first, model later

1. Shape and summary 2. Missing, dup, impossible 3. Distributions and outliers 4. Relationships 5. Who is in the sample 6. Hypotheses Model later only then
Exploratory analysis runs six look-first steps in order — shape and summary, integrity, distributions and outliers, relationships, who is in the sample, then hypotheses — before any modeling begins.

A new dataset is a stranger, and exploratory analysis is the interview you run before you trust it with anything important. Instead of feeding the numbers straight into a model and believing whatever comes out, you first look: at the shape of the table, at what is missing or impossible, at how the values are distributed, at what moves with what, and at who the rows actually represent. The whole discipline fits in four words — look first, model later.

The habit is old and deliberate. John Tukey made the case for it in his 1977 book Exploratory Data Analysis, arguing that you should examine and graph data to see what it is trying to say before you commit to a formal model or a hypothesis test. Skipping that step is like interviewing a job candidate instead of hiring on the resume alone: the questions are where you find out whether the resume is true: you can hire on the strength of the resume alone, but the interview is where you find out whether the resume is true.

An interview follows a rough script, and so does this one: check the shape and summary, scan for broken values, eyeball the distributions and their outliers, look at how columns relate while remembering that moving together is not proof, and ask who is even in the sample. You leave not with conclusions but with questions worth testing. The rest of this guide walks that script one question at a time.

Start with shape and summary stats

The first questions are the dull ones, and they save you the most grief. How many rows are there, and how many columns? What is each column supposed to be — a number, a date, a category, a label? A file that should have fifty thousand rows and has three hundred, or a column of prices that arrived as text, is a problem you want to find in the first minute, not after an afternoon of analysis built on top of it.

Once the shape is clear, summarize each numeric column: where it centers and how far it spreads. The center and spread of a pile of numbers are familiar ground — a mean or median for the middle, a range or standard deviation for the width — and here they serve as a quick fingerprint of every column at once. A minimum age of negative four or a maximum temperature of nine hundred tells you something is wrong before you have plotted a single point.

What a summary gives you in speed it takes back in detail. A single number for center and a single number for spread flatten the whole column into two facts, and two facts can hide a great deal — a second peak, a gap, a cluster of duplicates all reading the same value. The summary is where you start looking, never where you stop.

Missing, duplicate, impossible

Before you believe a single pattern, run the integrity pass: hunt for the three classic defects that quietly poison conclusions. Missing values — blank cells, empty strings, a placeholder like NA — decide how much data you really have and whether a summary was computed over the rows you think it was. Duplicate rows inflate counts and drag averages toward whatever got recorded twice. Impossible values — a negative quantity, a date in the future, a heart rate of zero on a living patient — are the fingerprints of a broken sensor or a botched import.

This is the same care that cleaning a dataset demands, pointed at a different goal. Cleaning fixes the values so later steps can run; the integrity pass just finds and counts the damage so you know how much to trust what follows. It is like inspecting a used car before buying it: you are not repairing anything yet, only finding out what would need repairing and whether it is worth driving off the lot: you are not repairing anything yet, you are finding out what would need repairing, and whether the whole thing is worth driving off the lot.

Finding a bad value is not the same as knowing its fate. A stray reading might be a typo to drop, a real extreme to keep, or the single most important event in the batch. The integrity pass flags candidates and hands you the count; only knowing where the data came from tells you what each flag actually means.

Plot it, don't just average it

a rising line one point, off alone same mean, same spread, same correlation different shape, so you must plot it
Two datasets with identical mean, spread, and correlation can have completely different shapes, so you must plot them, not only summarize them.

A summary is a compression, and compression loses things. The most famous demonstration is Anscombe's quartet, four small datasets built by the statistician Francis Anscombe in 1973. All four share almost the same mean, the same spread, and the same correlation — by the numbers they are identical — yet plotted they look nothing alike: one a tidy line, one a smooth curve, one a straight line dragged off course by a single point. The numbers agreed while the pictures disagreed completely.

That is why the look-first pass insists on the shape of a column, not only its center and spread. An outlier — a value sitting far from the rest of the pile — makes the point on its own. Because the mean adds up every value, one extreme number drags it toward itself, while the median, which cares only about position, barely moves. Watching the mean and median pull apart is often the first hint a stray is present, and it is a hint you get for free the moment you look instead of trusting the average. Reading only the mean is like wading a river reported as three feet deep on average: the average is no comfort in the one spot that is over your head.

So plot the distribution, or at least glance at the extremes, before you draw any conclusion from a summary. The goal is not a beautiful chart; it is to catch the second peak, the gap, or the lone point that a pair of summary numbers would have hidden from you entirely.

What moves together (and the catch)

With each column understood on its own, the next question is how they move together. Do taller people weigh more; do busier days bring more sales; does a higher reading in one sensor track a higher reading in another? A scatter of one column against another, or a correlation coefficient summarizing that scatter, is the fast way to see which pairs rise and fall in step and which ignore each other entirely. This is where hypotheses start to suggest themselves.

It is also where the oldest trap in analysis waits. Things that move together need not have anything to do with each other: a hidden third factor can drive both, the arrow of cause can point the other way, or the pattern can be pure coincidence in a small sample. Correlation is not causation, and exploratory analysis is exactly the phase where that rule earns its keep — a strong scatter is a reason to ask a question, never a license to answer it.

An outlier complicates this too, because a single far-flung point can invent a correlation that the bulk of the data does not support, or hide one that it does. So read a relationship the way you read a summary: as a lead to follow, worth a closer look and a proper test, not as a verdict you already reached.

Ask who's in the data

Every dataset is a photograph of something, taken from somewhere, and the frame leaves things out. Before trusting any pattern, ask the plainest question there is: who or what do these rows actually represent, and who is missing? Survey answers only speak for the people who answered. Server logs only describe the users who reached the server. Sensor readings only cover the window the sensor was running. The pattern you found is a pattern about whoever made it into the table, and nobody else.

This matters because the tidiest analysis in the world cannot describe rows that were never collected. A model of customer behavior built only on customers who stayed says nothing certain about the ones who left; a health finding drawn from volunteers may not carry to people who never volunteer. The gap between who is in the data and who you want to talk about is where confident, well-computed, wrong conclusions come from.

You do not have to solve this during exploration, but you do have to notice it. Write down who the rows are, name the group they leave out, and carry that caveat into every claim you make later. It costs one sentence now and saves a retraction later.

You leave with questions, not verdicts

Ask Look Refine examine sharpen re-ask Hypothesis to test later not a verdict
Exploratory analysis is a loop of ask, look, and refine that exits with a hypothesis to test later, not a conclusion.

The discipline that keeps exploratory analysis honest is knowing what it is allowed to produce. It produces hypotheses, not conclusions. Looking at a distribution, an outlier, or a scatter can tell you where something interesting might be; it cannot tell you that the interesting thing is real, because you went looking through the same data you would now use to confirm it. Confirming a hunch on the data that suggested it is how noise gets promoted to a finding.

So the work runs as a loop rather than a march to an answer. You ask a question, look at what the data shows, and refine the question — a surprise in one column sends you back to check another, a suspicious summary sends you to plot the shape. Each pass sharpens what you are asking. The loop ends not with a proven claim but with a short list of specific, testable hypotheses and the caveats that came with them.

Those hypotheses are the real output, and they are meant to be tested somewhere else: on fresh data, with a formal test, or with a model whose job is exactly the question you framed. Exploratory analysis has done its job when it hands the next step a good question and an honest account of the data behind it.

Interview your own data first

The look-first pass earns its place the first time you run it on data you collected yourself. On the robot-car build, the numbers are your own driving logs — timestamps, speeds, steering angles, battery voltage — and they arrive exactly as messy as the real world makes them: a dropped frame where a reading went missing, a line logged twice, a speed that spikes to an impossible value when a sensor hiccups. Interviewing that log before you trust it is the difference between tuning a real car and tuning a glitch.

The workflow is the same one it always is: check the shape and summary, scan for the missing and duplicate and impossible, plot the distributions and spot the outliers, look at what moves with what without confusing it for cause, and note who or what the log actually captured. None of it requires a model. All of it decides whether the model you build next is standing on solid ground.

Skipping the interview does not save time; it moves the cost downstream, where a confident, well-computed, wrong answer is far more expensive than a few minutes of looking would have been. Look first, model later — not because looking is virtuous, but because it is cheaper than being wrong.