Four Plots, Three Mechanisms: Reading missingno Like a Detective

msno.bar, matrix, heatmap and dendrogram, drawn on one dataset where I know the truth. What each plot can tell you about MCAR, MAR and MNAR, and the one thing none of them can ever see.

For my EDA Insem 2, sir drew three acronyms on the board (MCAR, MAR, MNAR) and four plots from a library called missingno (bar, matrix, heatmap, dendrogram). So I did the obvious thing: ran all four on a dataset and waited for one of them to tell me which mechanism I had.

None of them did. Not because the plots are bad, but because of one line of notation that, once you see it, tells you exactly what each plot can and cannot do. This post is that line, four annotated plots, and the detective routine I use now.

The one idea: missingno only ever sees R

Every dataset with holes is secretly two tables. There is Y, the complete data you wish you had, which splits into the part you observed (Y_obs) and the part hiding behind the blanks (Y_mis). And there is R, a grid of the same shape that just says 1 if a cell was observed and 0 if it was missing. In pandas, R is df.notnull().

A small table of hostel data with some weight and sleep values faded out as hidden, next to the matching 0/1 grid R
Fig. 1 Left: the table, with the values you never get to see faded in. Right: R, the only thing missingno ever draws.

Donald Rubin’s 1976 paper sorts every missing-data problem by what R depends on:

  • MCAR (missing completely at random): R depends on nothing. P(R | Y_obs, Y_mis) = P(R).
  • MAR (missing at random): R depends only on things you did observe. P(R | Y_obs, Y_mis) = P(R | Y_obs).
  • MNAR (missing not at random): R depends on the missing values themselves, and the equation refuses to simplify.

Now the key move. All four missingno plots are computed from R and nothing else. So:

  • They can show you structure inside R (which columns go missing together).
  • If you sort the rows by an observed column first, they can hint at a link between R and Y_obs. That is MAR’s signature.
  • They can never show a link between R and Y_mis, because Y_mis is the one thing nobody has. That is MNAR’s signature, and it is invisible by construction.

Keep that in your pocket. Every plot below is just this idea applied once.

The hostel weighing machine

To check the plots against the truth, I need data where I know the truth. So I simulated it. The setup is borrowed from Stef van Buuren’s book Flexible Imputation of Missing Data, which explains all three mechanisms with one weighing scale. Mine lives in a hostel.

300 students, columns age, wing, height, weight, bmi, steps, sleep, mess_rating. A few holes are the same in every version: steps drops 12% at random (the phone app fails to sync), and sleep plus mess_rating come from one evening form that 18% of students never submit. Then I made three copies where weight goes missing in three different ways:

  • MCAR: the scale’s battery dies at random moments.
  • MAR: the scale in wing B sits on a carpet and keeps failing. wing is recorded.
  • MNAR: heavier students quietly skip the scale.

bmi is computed from weight, so it vanishes with it. All three copies lose 23 to 24% of their weights, so any difference in the plots comes from the mechanism and nothing else.

Plot 1: msno.bar, the thermometer

msno.bar(df) draws one bar per column: how many rows are not missing. That’s df.notnull().sum(), nothing more.

Two bar charts of non-missing counts per column, MCAR and MNAR, with nearly identical bars
Fig. 2 MCAR on the left, MNAR on the right. One is harmless, the other biases every average you compute. The bars cannot tell them apart.

It answers “how much is missing?” beautifully and “why?” not at all. A random battery failure and heavy students hiding produce twin bars, 227 vs 228 weights. The bar is a thermometer, not a diagnosis. I still run it first, every time, because “how much” decides whether you are cleaning a few cells or rescuing a column.

Plot 2: msno.matrix, R itself

msno.matrix(df) is the most honest plot of the four: it is R, drawn as a picture. Dark means data, pale means missing, one horizontal line per row.

Anatomy of msno.matrix: dark cells for data, pale gaps for missing values, and a sparkline on the right counting filled cells per row
Fig. 3 How to read it. The sparkline on the right counts filled cells per row, with the worst and best rows labelled (here 4 and 8 out of 8).

The trick nobody tells you: the row order is your choice, and it changes everything. The rows come in file order, which is usually arbitrary. If missingness depends on some observed column, sort by that column first and the holes will line up. That’s MAR becoming visible.

sort_then_look.py
import missingno as msno
msno.matrix(df.sort_values("wing"))
Three matrix plots, MCAR, MAR and MNAR, each sorted by wing, with the MAR panel showing holes concentrated in wing B
Fig. 4 All three tables sorted by wing (dashed line = where wing A ends). MAR lights up. MCAR and MNAR both look like salt and pepper.

Look at the middle panel: wing A is almost solid, wing B is full of holes. That’s the carpet, sitting right there in the picture. Notice two other things. The sleep and mess_rating gaps line up row for row in every panel (that evening form). And the MNAR panel on the right looks exactly as innocent as MCAR.

The one that bit me: MNAR wears an MCAR costume

I kept sorting the MNAR table by every column, hoping a band would appear. Here is what happened.

Three matrix plots of the MNAR table: in file order, sorted by height, and sorted by the true weight that would be unavailable in practice
Fig. 5 Sorting by height gives a faint drift, because height is correlated with weight (r = 0.62). The clean band only appears when sorted by the true weight, a column you never have.

Sorted by height, there is a faint drift towards the bottom, because tall students tend to be heavy and so leak a little of the truth. Sorted by the true weight, the band is unmistakable. But in real life that third panel does not exist. The cause of MNAR missingness is the very value you lost. Of course you can’t sort by it.

Plot 3: msno.heatmap, which columns go missing together

msno.heatmap(df) takes each column’s missing indicator (1 if the cell is missing, 0 if not) and correlates them pairwise. For two 0/1 columns, that Pearson r is the phi coefficient. It answers: when this column is missing, is that one missing too?

I read the library’s source to get the details right, because the picture hides some decisions:

  • Columns that are never missing (or always missing) are dropped. They have zero variance, so no correlation exists.
  • Only the lower triangle is drawn.
  • Any |r| below 0.05 is shown as a blank cell, values from 0.95 up to (but not including) 1 show as <1, and everything else is rounded to one decimal.
Nullity correlation heatmap with bmi and weight at 1, sleep and mess_rating at 0.8, and the rest blank
Fig. 6 The MAR table's nullity correlations. Each number is about missingness, not about values.

The heatmap is great at finding shared causes of holes. bmi vs weight is 1 because one is computed from the other. sleep vs mess_rating is 0.8 because they come from the same form (not 1, because 17 students submitted the form but skipped one question). Everything else is blank, meaning unrelated.

What it can’t do: say anything about MAR in the sense Rubin meant it. The carpet effect is a link between weight’s holes and the values of wing. Since wing is never missing, the heatmap drops it entirely. A near-blank heatmap is not a certificate of MCAR.

Plot 4: msno.dendrogram, the same idea at scale

With 8 columns the heatmap is readable. With 80 it isn’t, so missingno also clusters the columns. Under the hood it is one SciPy call: hierarchical clustering with average linkage on the transposed 0/1 matrix, using Euclidean distance.

That makes the y-axis surprisingly concrete. For two 0/1 columns, the Euclidean distance is sqrt(number of rows where exactly one of them is missing).

Dendrogram of nullity columns: weight and bmi join at height zero, sleep and mess_rating join at 4.12, never-missing columns at zero
Fig. 7 weight and bmi join at 0 (identical holes). sleep and mess_rating join at 4.12, which is sqrt(17): they disagree on exactly 17 rows. The never-missing columns huddle together at 0.

Read it bottom-up: the lower two columns join, the more alike their holes. Same strengths and same blind spot as the heatmap, because it’s the same information, just clustered.

So how do you actually tell them apart?

Here is the move that works. Stop staring at R alone and split every observed column by whether the target column is missing. If the people with holes look different on anything you can see, the holes are not completely random.

  • Categorical column: compare missing rates across groups (chi-square test).
  • Numeric column: compare the two distributions (a t-test, or just overlapping histograms).
  • Everything at once: Little’s MCAR test (1988), which checks whether the column means differ across missingness patterns.
Top row: percent of weight missing in wing A and B for each mechanism. Bottom row: height distributions for rows with weight observed vs missing
Fig. 8 Top: % of weight missing per wing. Bottom: height of students whose weight is observed (filled) vs missing (outline). MAR fails on wing, MNAR fails on height, MCAR passes both.

MAR gets caught by wing (4% vs 40% missing, p = 2.7e-13). And look at MNAR: the wing test passes, but the students whose weight is missing are clearly taller (p = 2.2e-08). Height is the proxy leaking the truth. Little’s test on height, weight, wing and steps says the same thing:

tableLittle’s chi-squarep-valueverdict
MCAR4.1 (df 8)0.84no evidence against MCAR
MAR57.1 (df 8)1.7e-09not MCAR
MNAR38.4 (df 8)6.4e-06not MCAR

Read the last row twice. The test caught MNAR, but it can only ever say “not MCAR”. It has no way to tell you whether that’s the carpet kind (MAR) or the self-hiding kind (MNAR). And it only caught it because height happened to be correlated with weight. Drop height from Little’s test and MNAR sails through with p = 0.55.

Try it yourself. This runs real Python in your browser (the first run takes a few seconds to load pandas and SciPy). Your numbers will differ slightly from the figures, but the pattern won’t:

python⌘↵ to run

Change the 2.2 in the MNAR line to 0 and watch it turn into MCAR. Change 0.45 to 0.05 and the carpet stops mattering.

The part no plot can do

Here’s why all of this matters, drawn with the one thing you never get in real life: the full truth.

Histograms of observed weight against the true weight distribution, for MCAR and MNAR
Fig. 9 Grey is the truth, colour is what you observe. Under MCAR the observed mean stays honest (64.2 vs 63.9). Under MNAR the heavy tail drains away and the mean drops to 61.2.

Under MCAR, deleting the holes just costs precision. Under MNAR, every average, correlation and model built on the observed rows is quietly wrong, and nothing in the data tells you. van Buuren’s advice for MNAR is the honest one: go find out why values are missing (ask whoever collected the data), and run what-if analyses. For example: if the missing students were 5 kg heavier on average, does my conclusion survive?

My routine now

Flowchart: bar and matrix first, then split observed columns by missingness; no difference means consistent with MCAR, a difference means not MCAR; both branches end with asking the domain
Fig. 10 The detective routine. Notice that both branches end with a human question, not a plot.
  1. msno.bar for how much, msno.matrix for where. Sort the matrix by every column you suspect.
  2. msno.heatmap / msno.dendrogram to find columns that share a cause (same form, same sensor, derived columns).
  3. Split every observed column by missingness, then run Little’s test. A difference means not MCAR.
  4. Ask the human question every time, even when everything passes: could the value itself be the reason it’s missing?

The mental model I keep: missingno draws R. MCAR and MAR leave fingerprints on R plus the columns you have. MNAR leaves its fingerprints on a column you lost. Plots can find the first two. Only thinking finds the third.

If you’d rather watch

I looked for one video that covers all four plots and the three mechanisms properly. It doesn’t exist (part of why this post does). These four together cover it:

Sources

Every figure here comes from one simulated dataset (seed 26), so the truth is known and every number is checkable.

Next time: what to actually do once you know the mechanism. Deletion, mean imputation, and why the second one quietly shrinks your variance.

Adrishyam: Dehazing Images with CV My Py Learnings