Methods & accuracy
Which formula, which variant, and how we know it is right. Written for teachers, reviewers and anyone comparing a result with SPSS, R, JASP or jamovi.
How the engine is checked
ExplainStats is a Google Sheets add-on, but its statistics engine is plain JavaScript that also runs outside Google Sheets. That lets us compare it, automatically, with reference implementations on hundreds of random datasets. Every build runs these comparisons; a release ships only when they all pass.
| Comparison script | What it checks | Reference | Verifications |
|---|---|---|---|
| Core | Distribution functions (t, F, χ²: tails and inverses), descriptive statistics, histogram, Pearson correlation matrix, t-tests (three types, with the confidence interval of the difference and Cohen's d), F-test, one-way ANOVA, linear regression with residuals, chi-square test and Yates' correction | SciPy, statsmodels | 912 |
| Advanced | Shapiro-Wilk, Levene and Brown-Forsythe, Mann-Whitney (exact and normal, with and without ties), Wilcoxon signed-rank, Kruskal-Wallis, Spearman, Fisher's exact test, studentized-range distribution, Tukey HSD, Games-Howell, two-way ANOVA with and without replication, z-test, covariance, Cronbach's alpha, rank and percentile | SciPy, statsmodels, hand-computed references | 431 |
| Welch's ANOVA and Friedman | Welch's F, degrees of freedom and p-value; Friedman's χ² | statsmodels, SciPy | 22 |
| Pro tools | Non-central t distribution, sample size and power (t-tests, two proportions, correlation), A/B test, control charts (I-MR, X-bar/R), moving average, exponential smoothing, random-number moments | SciPy, statsmodels | 50 |
| Total | Tolerance: relative 10-8 on statistics, 10-7 on p-values (tighter than any printed output) | 1,415 | |
On top of these, 248 end-to-end scenarios run the whole add-on against a simulated spreadsheet in the four languages (every tool, exam mode, the steps sheets, error messages), and an independent reviewer re-derived Welch, ANOVA, Tukey, Games-Howell, Shapiro-Wilk, Levene, Mann-Whitney and Kruskal-Wallis results by hand against SciPy 1.17 and statsmodels 0.15, with agreement better than 10-9.
Found a discrepancy? Write to [email protected] with a small example. A confirmed error is corrected within seven days and listed on this page.
Variants and conventions, test by test
Different software make different default choices. These are ours, so that you can explain why a number differs from another program.
| Test | What ExplainStats computes | Same as |
|---|---|---|
| Descriptive statistics | Sample variance and SD (n − 1). Skewness and kurtosis with the sample-adjusted formulas of Excel (SKEW, KURT, excess kurtosis). Quartiles by linear interpolation between order statistics (QUARTILE.INC, R type 7). Outliers: beyond 1.5 × IQR from the quartiles. Confidence interval of the mean with the t distribution. | Excel, Google Sheets, R (type 7) |
| t-tests | Paired, equal variances (pooled) and Welch (unequal variances, Welch–Satterthwaite degrees of freedom). Two-tailed p-value; one-tailed also shown. Confidence interval of the difference: difference ± tα/2, df × SE, with the same SE as the test. The Assistant uses Welch by default, as recommended since Delacre et al. (2017). | SciPy ttest_ind, ttest_rel; R t.test |
| Cohen's d | Equal variances: pooled SD. Welch: average of the two variances, √((s₁² + s₂²)/2). Paired: dz = mean of the differences ÷ SD of the differences. The table names the variant. Thresholds: 0.2 small, 0.5 medium, 0.8 large (Cohen, 1988). No Hedges correction. | JASP (d and dz), G*Power |
| One-way ANOVA | Classic F with η² (SSbetween/SStotal). Tukey HSD uses the studentized-range distribution computed with the Copenhaver & Holland algorithm (the one in R's ptukey), Tukey–Kramer form for unequal group sizes. | SciPy f_oneway, statsmodels pairwise_tukeyhsd, R TukeyHSD |
| Welch's ANOVA | Welch's F with its approximate degrees of freedom; Games-Howell post-hoc tests with per-pair Welch degrees of freedom. | statsmodels anova_oneway(use_var="unequal"), R oneway.test |
| Two-way ANOVA | Balanced designs only (equal rows per sample), type I = type III sums of squares in that case, as in the Analysis ToolPak. | statsmodels OLS/ANOVA, Excel |
| Normality | Shapiro-Wilk W with Royston's 1995 algorithm (AS R94), 3 ≤ n ≤ 5,000, plus a Q-Q plot. In the Assistant, a group that fails the test but has n ≥ 30 is still analysed with a parametric test, and the report says so explicitly. | SciPy shapiro, R shapiro.test |
| Equal variances | Levene (centre = mean) and Brown-Forsythe (centre = median). The Assistant uses Brown-Forsythe. | SciPy levene |
| Mann-Whitney U | Exact distribution when there are no ties and a group has 8 values or fewer; otherwise normal approximation with tie correction and continuity correction (0.5). Two-tailed. | SciPy mannwhitneyu defaults |
| Wilcoxon signed-rank | Zero differences dropped. Exact distribution when there are no ties and n ≤ 50; otherwise normal approximation with tie and continuity corrections. | SciPy wilcoxon (zero_method="wilcox") |
| Kruskal-Wallis | H with tie correction; when significant, pairwise Mann-Whitney tests with Bonferroni correction (not Dunn's test; the table says which). | SciPy kruskal |
| Friedman | χ² statistic with tie correction, k − 1 degrees of freedom. | SciPy friedmanchisquare |
| Correlation | Pearson r and Spearman ρ (Pearson on average ranks); p-value from t = r √((n − 2)/(1 − r²)) with n − 2 degrees of freedom, two-tailed. | SciPy pearsonr, spearmanr |
| Regression | Ordinary least squares by QR decomposition (stable with correlated predictors), up to 16 predictors; R², adjusted R², standard error, ANOVA table, coefficients with t, p and confidence intervals, residuals. Same table layout as the Analysis ToolPak. | statsmodels OLS, Excel |
| Chi-square | Pearson's χ² of independence without correction, Cramér's V, warning when expected counts are below 5. For 2 × 2 tables the table also gives χ² and p with Yates' continuity correction (|O − E| reduced by 0.5, never below 0), the default of R and the "Continuity Correction" row of SPSS. | SciPy chi2_contingency (correction=False / True) |
| Fisher's exact test | 2 × 2 only; two-tailed p-value as the sum of the probabilities of all tables at most as likely as the observed one (the R convention); odds ratio. | R fisher.test |
| Cronbach's alpha | Complete rows only; alpha and "alpha if item deleted"; thresholds 0.7 acceptable, 0.8 good, 0.9 excellent. | Standard formula |
| A/B test | Two-proportion z-test with pooled standard error, two-tailed; difference with its confidence interval; sample size needed for the observed lift. | statsmodels proportions_ztest |
| Sample size and power | t-tests through the non-central t distribution (exact), two proportions through Cohen's h, correlation through Fisher's z. | statsmodels power classes, G*Power |
| Control charts | I-MR (constants 2.66 and 3.267), X-bar/R with the usual A₂, D₃, D₄ constants (subgroups of 2 to 10), p chart with normal limits. | Montgomery, Statistical Quality Control |
Things to know before comparing with other software
- Quartile chart, not a full box plot. Google Sheets has no native box plot. ExplainStats draws a candlestick chart whose box runs from Q1 to Q3 and whose lines run from the smallest to the largest value inside the 1.5 × IQR fences. It does not draw the median line or the outlier points; both are in the results table.
- Rounding. Tables keep full precision (displayed with 4 decimals); the interpretation and the APA line round to 2 or 3 decimals, and p-values below 0.001 are written "p < .001".
- Exact tests. Switching between exact and approximate methods (Mann-Whitney, Wilcoxon) follows SciPy's defaults, which differ from SPSS (asymptotic unless you ask for exact) and from R (exact below 50 without ties). The "Method" row of each table says which was used.
- Welch by default. The Assistant recommends Welch's t-test even when variances look equal, because it loses almost nothing when they are and protects you when they are not. Classic Student's t-test is still available as a tool.
- Nothing is simulated. No bootstrap, no Monte Carlo: every p-value is computed from the distribution, so running a test twice gives the same number. Only the random-number and sampling tools use a random generator (seedable).
Reproducing the checks
The comparison scripts are Python files that load the add-on's engine with Node.js, generate random datasets, run SciPy and statsmodels on the same data and compare every number. If you teach statistics and would like to run them, or to add a case that matters for your course, write to [email protected].