The Sunday Edition — the statistics column

SPSS output in plain English.

2026-05-12 · by Inés Delacroix-Vega — PhD, statistics column

The software will run almost anything it is asked to run, and it will print a tidy table whether or not the question behind it made sense. Everything difficult about a statistics assignment happens after the table appears, in the paragraph that has to say what it means.

Three questions every output answers.

Whatever procedure produced it, an output table answers three questions in sequence, and reading it in that sequence keeps the interpretation honest. What was compared or related, and in whose units. Whether the pattern observed is larger than the noise this design would throw up by accident. How large the pattern is, expressed so that a reader who has never run the procedure can picture it.

Most coursework answers only the second. A result reported as significant and then abandoned has told the reader that something is unlikely to be a fluke, which is a statement about probability rather than about people, patients or firms. The third question is what the assignment was actually asking, and it is where the marks sit.

The Sig. column, read correctly.

SPSS prints significance to three decimal places and will therefore display .000 for any value below a thousandth. That is not zero, and a write-up reporting p equals .000 has copied rather than read; the convention is to report it as less than .001. It is a small thing, and graders in every discipline notice it.

Two larger misreadings sit behind that one. A significant result is not automatically a large one, because significance responds to sample size as much as to effect, so a trivial difference measured in thousands of people will clear the threshold comfortably. And a result that is not significant is not evidence that two groups are the same; it is a failure to detect a difference, which describes the sensitivity of the study rather than the state of the world. The same discipline governs reading other people's tables, which is most of what a literature review consists of.

The assumption tables are printed too, and they are graded.

The tables a procedure prints alongside the headline result are not decoration. Levene's test appears in the comparison procedures to indicate whether the variances are similar enough for the standard version of the test, and the software helpfully prints a second row for the case where they are not. Reading the wrong row is among the most common silent errors in submitted coursework, and it is invisible unless a grader looks.

The same holds for the diagnostics around regression and correlation, for normality checks, and for whatever else the procedure requires before it is entitled to speak. State what was checked, state what the check found, and state what was done about it. A rubric criterion for assumptions is usually satisfied by three plain sentences, and it is usually left empty.

Writing the sentence underneath the table.

The APA reporting convention is fixed, which is a mercy, because it removes the guesswork entirely. Name the comparison in words, give the statistic with its degrees of freedom exactly as the table prints them, give the p value, give an effect size, then close with a clause saying what the number means in the units the study cares about. Direction belongs in the words rather than in the arithmetic; a reader should know who scored higher without decoding a sign.

Then one sentence that no procedure can supply: what the design does not permit anybody to claim. A correlation is not a cause, a convenience sample is not a population, and a single-site study describes that site. Writing that sentence yourself is considerably cheaper than having a committee write it for you two years later, and where a whole statistics sequence rather than one table is the obstacle, the statistics column carries it.

Questions to the desk.

What does a p value actually tell me?

It estimates how often data at least this extreme would arise if the effect being tested were absent. It is a statement about the data under an assumption, not the probability that your hypothesis is true and not a measure of how important the finding is. That is why an effect size is reported alongside it, and why a threshold on its own settles very little.

Why report an effect size if the result is significant?

Because significance answers whether an effect is detectable and effect size answers whether it matters. With a large sample a difference too small to change any decision will still be significant, and with a small sample a substantial difference may not be. Reporting both lets a reader judge the finding rather than take the threshold on trust, and most rubrics now require it.

Which row do I read when Levene's test is significant?

The second one, the row the software prints for unequal variances, and you say in the write-up that you did so and why. A significant Levene's result indicates the variances differ enough that the standard version of the test is not appropriate. Reporting the first row anyway is a quiet error that invalidates the paragraph resting on it.

Does it matter whether the course uses SPSS, R or Excel?

For the reasoning, no. The same three questions apply and the same reporting conventions govern the write-up, whichever tool produced the numbers. What changes is the labeling and the layout, so it is worth learning where your package hides degrees of freedom, exact p values and effect sizes. Submit output in whatever form the assignment specifies.

Inés Delacroix-Vega
Written by
Inés Delacroix-Vega
PhD, statistics column · one of eight editors on the masthead.
The Sunday Edition

One like this every Sunday morning.

Subscribe — free
Elsewhere in the paper
Take my statistics classA literature review that holdsCapstone and dissertation help

Name the deadline. The desk takes it from there.

Place a request — free quote
At the desk