1. Expanded summary statistics

1.1 Expansion from P7 to P8

Summary statistics provide essential information to help you get to know your variables, their fundamental statistical properties and numerical characteristics. In P8, if you click on Tools > Summary Stats... you will see (Fig. 1.1) that we have greatly expanded the list of options on offer, compared to those available in P7 (where it used to be accessed under Analyse > Summary Stats...). In some cases, you may wish to calculate summary statistics across the values you are getting across a sample, so this is catered to as well (just as in P7). We have also added the possiblity to output summary statistics separately for different levels of a factor (or different groups of an indicator, in the event that you are summarising samples instead of variables).

1.Summary_Stats_dialog_P7&_P8.png

Fig. 1.1. Comparison of the 'Summary Stats...' dialog options in PRIMER 7 vs PRIMER 8.

In what follows, we shall begin by providing a brief description of each of the summary statistics, then demonstrate the use of this new tool by implementing it to summarise information about variables (i.e., across samples). Note that you can also use this tool to summarise information about samples (i.e., across variables).

1.2 Definitions of statistics

Given a set of values $\{ y_1, y_2, ..., y_n \}$ for any individual variable $Y$, the following summary statistics can be calculated by clicking on Tools > Summary Stats... in PRIMER 8:

1.3 Biotic data: summary stats

To show the utility of this tool, we will calculate some summary statistics from a study examining changes in macrofaunal communities inhabiting sediments near an oil platform (Ekofisk) in the North Sea, provided by Gray et al. (1990) . These data consist of counts of the abundances of $p$ = 174 taxa (mostly identified to species level) sampled by 3 day grabs taken at each of $N$ = 39 stations, situated in an approximately 5-spoke radial design leading out from the oil platform. The stations were classified into the following four groups according to their relative distance from the centre of drilling activity at the Ekofisk oil platform: D = less than 250m, C = 250m - 1km, B = 1 - 3.5km, and A = more than 3.5km away.

Start running PRIMER 8, then click File > Open... to open the data file named 'Ekofisk_macrofauna_counts.pri' (found inside the 'Examples_P8 > Ekofisk_macrofauna' folder).

From the 'Ekofisk_macrofauna_counts' data sheet inside PRIMER, click on Tools > Summary Stats..., as shown below:

2.Summary_Stat_menu_item[i].png

Choose from the dialog all of the summary statistics you would like to calculate on the dataset. For example, we might choose to calculate the average, median, minimum, maximum, range, standard deviation, number of zeros, singletons, and the number of non-zeros (i.e., frequency of occurrence), like this:

3.Summary_Stat_dialog(all_data,biota)[new].png

Click 'OK', and you will see a 'Summary Statistics 1' file, in which you are shown the choices you made for running that routine, viz:

4a.Summary_Stat_output_notepad(all_data,biota)[i].png

as well as a resulting output file 'Data1', which provides all of these summary statistics for each of the individual variables (taxa), as shown below:

4b.Summary_Stat_results(all_data,biota)[i].png

1.4 Split summary stats results by groups

To run summary statistics on your variables separately for multiple groups of data, just choose a factor by which you would like to split the data in the dialog. For example, for the Ekofisk dataset, suppose we wished to obtain summary statistics (means and standard errors) separately for each of the 'Distance' groups. We can achieve this by choosing Tools > Summary Stats... > (Summarise > $\bullet$Variables) & (Split by > $\checkmark$Factor > Dist) & (Statistics > $\checkmark$Average & $\checkmark$Standard Error), as shown in the dialog below:

5.Dialog(split_data_biota)_[new].png

By default, PRIMER will provide the output for each summary statistic you have asked for as a separate data sheet, and the groups (corresponding to levels of the factor by which you have chosen to do the splitting), will be the 'Samples'. In the present example, you will see that the averages for each variable for the four distance groups (A, B, C, D) have been provided in 'Data2' and the standard errors are given in 'Data3'. There will be as many new data sheets generated as there are summary statistics asked for in the dialog. Note that the name of the summary statistic provided in each sheet is given in the title at the top of the sheet, e.g., 'Ekofisk oilfield macrofauna -Summary Statistics, Average'.

6.Results(split_data_biota)_[i].png

An alternative output format can be requested here, with results for different groups given on different sheets. To do this, choose Tools > Summary Stats... > (Summarise > $\bullet$Variables) & (Split by > ($\checkmark$Factor > Dist) & (Output > $\checkmark$Levels/groups as separate sheets)) in the dialog, like so:

7.Alt_dialog_take2(split_data,biota)[new].png

Doing this yields a separate sheet of summary statistics for each of the groups (e.g., 'Data4' has all of the summary statistics you requested calculated for group D only, 'Data5' has the results for group C, and so on (see below). Note that, for this alternative type of output, the name of the factor on which the split was done and the specific group (factor level) is given in the title at the top of each new data sheet of results; for example, 'Ekofisk oilfield macrofauna -Summary Statistics, Dist: A'.

8.Alt_results(split_data_biota)_[i].png

1.5 Environmental data: summary stats

For environmental data, we might choose to calculate different sorts of summary statistics than the kinds of things we would want to know about biotic data consisting of counts. For count data, quantities like the numbers of zeros, singletons, doubletons and frequencies of occurrence might well be of interest. However, these are not typically meaningful for variables that have been recorded as (effectively) continuous values, such as temperature, dissolved oxygen, etc. Instead, when we deal with environmental data (or, more generally, continuous quantitative variables), it may be helpful to know what the smallest non-zero value is, or to know the degree of skewness. These can aid in determining an appropriate transformation (e.g., to obtain approximate symmetry).

Let's look at some environmental data from a study of benthic soft-sediment assemblages in the Firth of Clyde, SW Scotland ( Pearson & Blackstock (1984) ). The abundance and biomass of 84 macrofaunal species as well as contaminant data (organic enrichment and concentrations of heavy metals in the sediment) were sampled at a series of 12 sites along a transect that passed through the sewage-sludge disposal ground at Garroch Head. Open the data file named 'Clyde_environment.pri' in PRIMER (found inside the 'Examples_P8 > Clyde_macrofauna' folder).

9.Env_Clyde[i].png

Of course, tools such as histograms and draftsman plots (found under the Plots menu) are very useful for visualising the distributions of values for the individual variables. Summary statistics complement these visual tools, yielding some important additional detailed information. Click Tools > Summary Stats..., then choose to output the following statistics (shown in the dialog below): average, minimum, maximum, standard deviation, symmetry (0.05), skewness, kurtosis, and smallest number above a threshold of 0.

10.Env_Summary_Stats_dialog[new].png

The resulting output file (called 'Data1' and shown below) indicates that many of the variables are right-skewed, with positive values for skewness, and values less than 0.5 for the symmetry statistic, while others (Co, Ni) are apparently left-skewed (negative skewness and symmetry > 0.5). Depth, however (Dep) is a fairly flat (platykurtic) variable (with negative excess kurtosis). Also evident is the fact that values of cadmium concentration (Cd) reach a minimum value of zero, and that the smallest value recorded above zero is 0.1. Thus, if we were inclined to transform the values for Cd (e.g., to make its distribution more symmetric), we might consider a log transformation such as $log(y_i+c)$, where $c$ = 0.1.

11.Summary_Stats_Clyde_results[i].png

(Note: the variable 'Cd' has been highlighted in the above image of 'Data1' by clicking on the heading for that column. Highlighting variables (or samples), can be very useful for selecting subsets of data, or for applying a transformation to a subset of variables.)