20 July 2026

Analysing data with basic statistical tests

By Daniel Foster
  • 5k
  • 5k
  • 5k
Analysing data with basic statistical tests

Why statistics don't have to be scary

If your heart sinks at the sight of a scatter plot, you're in good company. Most students arrive at university able to write a decent essay but with only a hazy memory of GCSE maths, and suddenly a research methods module expects them to test hypotheses in SPSS or R. Here's the reassuring truth: the vast majority of undergraduate dissertations only need three or four basic tests. Once you understand which question each test answers, the software does the hard arithmetic for you.

This guide covers the workhorses of student research: the t-test, the chi-square test, and correlation coefficients. Get to grips with these and you'll handle most quantitative projects with confidence.

Start with your research question, not the test

The most common mistake is choosing a test because a friend used it, or because it sounds impressive. Instead, ask yourself two questions about your data.

  • What type of data do I have? Continuous data (height, test scores, reaction times) behaves differently from categorical data (yes/no, pass/fail, favourite brand).
  • What am I comparing? Are you comparing groups, or looking for a relationship between two variables?

Groups plus continuous data points towards a t-test. Categories plus groups points towards chi-square. Two continuous variables points towards correlation. Write your research question on a sticky note above your desk, and let it guide every decision.

The t-test: comparing the averages of two groups

A t-test tells you whether the difference between two group means is likely to be real or just down to chance. If you're comparing the average exam mark of students who attended revision workshops against those who didn't, a t-test is your tool.

There are three versions, and choosing correctly matters:

  • Independent samples t-test – for two separate groups, such as men and women, or two seminar classes.
  • Paired samples t-test – for the same people measured twice, such as a pre-course and post-course test.
  • One-sample t-test – for comparing your sample against a known national average.

In your write-up, report the t value, the degrees of freedom, and the p value, ideally with the means and standard deviations for each group. A p value below 0.05 is conventionally treated as statistically significant, meaning there's less than a 5% probability of seeing this difference if there were truly no effect. Always report the actual means too: "the workshop group scored 12 marks higher on average" tells your reader far more than a p value alone.

Chi-square: testing relationships between categories

Chi-square (often written as χ²) is used when both of your variables are categorical. Suppose you want to know whether there's an association between a student's faculty and whether they use the library's online resources. Neither variable is a number you can average, so a t-test won't work. Chi-square asks whether the pattern you've observed is different from what you'd expect if the two variables were completely unrelated.

You'll see the results presented in a contingency table, with observed counts in each cell and expected counts calculated by the software. Larger discrepancies between observed and expected produce a larger chi-square value and a smaller p value.

Two practical warnings. First, every expected count should ideally be five or more; if many are smaller, you may need to combine categories. Second, chi-square tells you that an association exists, not how strong it is. For strength, look at a measure such as Cramér's V, which runs from 0 (no association) to 1 (perfect association).

Correlation: measuring how two variables move together

Correlation coefficients describe the strength and direction of a relationship between two continuous variables, such as hours spent revising and final mark. The coefficient, usually written as r, ranges from −1 to +1.

  • +1 – a perfect positive relationship: as one variable rises, so does the other.
  • 0 – no linear relationship at all.
  • −1 – a perfect negative relationship: as one rises, the other falls.

As rough guidance, coefficients around 0.1 to 0.3 are weak, 0.3 to 0.5 moderate, and above 0.5 strong, though this depends on your field. Use Pearson's r when your data are roughly normally distributed and the relationship looks linear; use Spearman's rho when your data are ranked or skewed, which is common with questionnaire scales.

The golden rule: correlation is not causation. Ice cream sales and drowning incidents rise together, but neither causes the other; both are driven by warm weather. Acknowledge this limitation explicitly in your discussion.

Reporting your results honestly

Statistical significance is not the same as practical importance. With a large sample, a trivial difference can produce p < 0.05. Always report effect sizes alongside p values, and describe what the finding actually means for your research question.

Check your assumptions before running anything: normality for t-tests and Pearson's correlation, independence of observations, and adequate expected counts for chi-square. Keep a simple analysis log recording every test you run, including the ones that didn't work out. Your marker cares about your reasoning far more than a perfect result, and a transparent, well-justified analysis will always earn stronger marks than a mysteriously flawless one.