KiwiScience logo
KiwiScience Statistics Guide

Kiwistat Statistics GuideKiwistat 统计方法指南

KiwiScience · the tests explained — written for postgraduate researchers

KiwiScience · 各项检验详解 — 为研究生研究者而写

This guide explains the statistics behind Kiwistat: what each test does, when to use it, its assumptions, and how to interpret and report the output. It assumes a strong science background but takes nothing for granted statistically. For the application itself — entering and importing data, tabs, graphs, saving — see the User Guide.

本指南讲解 Kiwistat 背后的统计学:每种检验做什么、什么时候用、有哪些前提假设,以及如何解读和报告输出。它假定读者有扎实的科学背景,但在统计方面不预设任何基础。至于应用本身——数据录入与导入、选项卡、图表、保存——请看使用指南。

1A short statistics primer统计学简明入门

Samples estimate populations用样本估计总体

Your measurements are a sample drawn from a larger population (all possible measurements at that site, of that species…). Statistics asks: what can this sample tell us about the population, given that another sample would have come out a bit different?

你的测量结果是从一个更大的总体中抽取的样本(该样点、该物种……所有可能的测量值)。统计学要问的是:既然换一批样本结果会略有不同,那么这一批样本能告诉我们关于总体的什么?

Describing variation: SD vs SE描述变异:标准差与标准误

  • Standard deviation (SD) describes how spread out the individual measurements are.
  • Standard error (SE) = SD ⁄ √n describes how precisely you know the mean. It shrinks as you take more measurements; SD does not.
  • 标准差(SD)描述的是单个测量值分散的程度。
  • 标准误(SE) = SD ⁄ √n 描述的是你对均值了解得有多精确。测量次数越多它越小;标准差则不会。

Error bars on Kiwistat's bar charts show mean ± SE by default. Say which you plot in your figure caption — reviewers check.

Kiwistat 柱状图上的误差线默认显示"均值 ± 标准误"。请在图注中写明你画的是哪一种——审稿人会查。

What a p-value actually isp 值究竟是什么

Every test starts from a null hypothesis (H₀): “there is no real difference / no real relationship — any pattern in my sample is chance.” The p-value is the probability of getting data at least as extreme as yours if H₀ were true.

每一项检验都从一个原假设(H₀)出发:"不存在真实的差异/不存在真实的关系——我样本里的任何规律都是偶然。"p 值是指:假如 H₀ 为真,得到至少和你这份数据一样极端的数据的概率。

A small p-value (conventionally < 0.05, the significance level α) means your data would be surprising under H₀, so you reject it and call the effect statistically significant.

p 值很小(惯例是 < 0.05,即显著性水平 α)意味着在 H₀ 下你这份数据会显得很意外,于是你拒绝 H₀,并称该效应具有统计学显著性。

Common misreadings to avoid:

几种常见的误读,要避免:

  • p is not the probability that H₀ is true, and 1−p is not the probability your hypothesis is right.
  • p = 0.06 vs p = 0.04 is not “no effect” vs “effect” — p-values are continuous evidence. Report the exact value.
  • Statistical significance ≠ practical importance. With large n, tiny irrelevant differences become “significant”; with small n, large real effects can be missed (low power). Always look at effect sizes and means, not just p.
  • Running many tests inflates false positives: at α = 0.05, about 1 in 20 truly-null comparisons will come out “significant” by chance. This is why post-hoc tests adjust for multiple comparisons.
  • p 不是 H₀ 为真的概率,1−p 也不是你的假设为真的概率。
  • p = 0.06 与 p = 0.04 的区别,并不是"没有效应"与"有效应"——p 值是连续的证据强度。请报告确切数值。
  • 统计学显著 ≠ 实际重要。n 很大时,微小而无关紧要的差异也会变得"显著";n 很小时,真实的大效应可能被漏掉(功效低)。要始终看效应量和均值,而不只是 p。
  • 做很多次检验会抬高假阳性:在 α = 0.05 下,约每 20 次真正无效应的比较中,就有 1 次会碰巧显得"显著"。这正是事后检验要对多重比较作校正的原因。

Assumptions matter前提假设很重要

The parametric tests in Kiwistat (ANOVA, Pearson correlation, curve fitting) assume, to varying degrees:

Kiwistat 中的参数检验(方差分析、Pearson 相关、曲线拟合)在不同程度上都有这几条假设:

  1. Independence — each observation is a separate experimental unit. (Ten readings of the same plant are not 10 independent replicates — that's pseudoreplication, and no software can fix it.)
  2. Normality — the residuals (scatter around group means) are roughly bell-shaped. Check with Normality & Transformation. ANOVA is fairly robust to moderate non-normality, especially with balanced group sizes.
  3. Equal variances — the scatter is similar in every group. Kiwistat checks this automatically with Levene's test and tells you what to do if it fails.
  1. 独立性——每个观测都是一个独立的试验单元。(对同一株植物测十次,不等于 10 个独立重复——那叫伪重复,没有任何软件能补救。)
  2. 正态性——残差(围绕各组均值的离散)大致呈钟形。用正态性与数据变换检查。方差分析对中等程度的非正态相当稳健,各组样本量均衡时尤其如此。
  3. 方差齐性——各组的离散程度相近。Kiwistat 会用 Levene 检验自动检查,若不通过还会告诉你该怎么办。

Why log-transform?为什么要取对数?

Environmental measurements are often right-skewed: many small values, a few large ones (concentrations, counts, biomass, rainfall). Log-transforming such data usually makes it more symmetrical and evens out variances — fixing two assumptions at once. Use Use log₁₀ values for statistics in the Column pane; back-transformed results are geometric means.

环境测量数据常常是右偏的:很多小值,少数几个大值(浓度、计数、生物量、降水量)。对这类数据取对数,通常既能让分布更对称,又能拉平方差——一举满足两条假设。使用 Column 面板中的 Use log₁₀ values for statistics;反变换回来的结果是几何均值。

2Choosing the right test如何选对检验方法

QuestionData neededTest
你的问题需要的数据检验方法
Does a measurement differ between the levels of one factor (3+ groups)?1 numeric response + 1 text factorOne-way ANOVA
某个测量值在一个因子的不同水平间是否有差异(3 组以上)?1 个数值型响应变量 + 1 个文本型因子One-way ANOVA
…the same, but data are skewed / not normal even after transforming?1 numeric response + 1 text factorKruskal–Wallis
……同上,但数据偏态/即使变换后仍不服从正态?1 个数值型响应变量 + 1 个文本型因子Kruskal–Wallis
Compare just two groups?1 numeric response + 1 two-level factorOne-way ANOVA (= t-test) or Mann–Whitney U if not normal
只比较两组?1 个数值型响应变量 + 1 个两水平因子One-way ANOVA (= t-test) or Mann–Whitney U if not normal
Do two factors affect a measurement, and do they interact?1 numeric response + 2 text factors, replicated in each combinationTwo-way ANOVA
两个因子是否都影响某个测量值,它们之间有交互作用吗?1 个数值型响应变量 + 2 个文本型因子,每种组合都有重复Two-way ANOVA
One treatment factor, applied once per block (field strip, day, bench)?response + treatment + blockBlocked ANOVA (RCBD)
一个处理因子,在每个区组(田间条带、天、试验台)内各施用一次?响应变量 + 处理 + 区组Blocked ANOVA (RCBD)
Did I measure the same subject several times (a time series)?one row per subject, one column per occasion (+ optional treatment factor)Repeated measures ANOVA
我是否对同一个体反复测量了多次(时间序列)?每个个体一行,每个测量时点一列(可另加处理因子)Repeated measures ANOVA
Are my data normal? Would a transformation help?1 numeric column (optionally grouped)Normality & Transformation
我的数据服从正态吗?做个变换会有帮助吗?1 个数值列(可按组划分)Normality & Transformation
Is there a straight-line relationship between two variables (with diagnostics)?2 numeric columns (X and Y)Linear regression
两个变量之间是否存在直线关系(含诊断)?2 个数值列(X 与 Y)Linear regression
What model (curve) best describes two variables?2 numeric columns (X and Y)Curve Fitting
哪个模型(曲线)最能描述这两个变量?2 个数值列(X 与 Y)Curve Fitting
Which of my many variables move together?2+ numeric columnsCorrelation Matrix (Pearson or Spearman)
在众多变量中,哪些是同向变化的?2 个及以上数值列Correlation Matrix (Pearson or Spearman)
Which samples are similar overall? What drives the variation?3+ numeric columnsPCA
哪些样品整体上相似?变异主要由什么驱动?3 个及以上数值列PCA
How many replicates do I need? (planning stage)an expected effect size or means + SDPower / sample size
我需要多少个重复?(试验设计阶段)预期效应量,或均值 + 标准差Power / sample size
I just want to see the shape of my data first.1+ numeric columnsBoxplots / scatterplot matrix
我只想先看看数据长什么样。1 个及以上数值列Boxplots / scatterplot matrix
Comparing just two groups? One-way ANOVA with two groups is mathematically equivalent to the classic t-test (F = t²). Its rank-based counterpart is the Mann–Whitney U test.只比较两组?两组的单因素方差分析在数学上等价于经典的 t 检验(F = t²)。它基于秩次的对应方法是 Mann–Whitney U 检验。
Parametric vs non-parametric. Parametric tests (ANOVA, regression, Pearson) assume roughly normal data and gain power from that assumption. Non-parametric tests (Kruskal–Wallis, Mann–Whitney, Spearman) work on ranks, make no normality assumption, and resist outliers — the right choice for skewed counts and concentrations that won't transform to normality. When data are normal, prefer the parametric test; it detects real effects with fewer samples.参数检验与非参数检验。参数检验(方差分析、回归、Pearson 相关)假定数据大致服从正态,并从这一假定中获得更高的检验功效。非参数检验(Kruskal–Wallis、Mann–Whitney、Spearman)基于秩次,不作正态性假定,且不易受离群值影响——对于怎么变换都达不到正态的偏态计数和浓度数据,它们才是正确的选择。当数据确实服从正态时,优先用参数检验;它能用更少的样本检出真实的效应。

3Normality & transformation正态性与数据变换

Runs the Shapiro–Wilk test on your variable — the most powerful general normality test for small-to-medium samples — plus a candidate transformation (log₁₀, ln or √), and recommends whether transforming helps.

对你的变量运行 Shapiro–Wilk 检验——中小样本下功效最高的通用正态性检验——同时试算一种候选变换(log₁₀、ln 或 √),并给出是否值得做变换的建议。

  • W close to 1 = consistent with normal; p < 0.05 = significant departure from normality.
  • Skewness: 0 = symmetrical; positive = long right tail (log usually helps); negative = long left tail.
  • If the data will feed an ANOVA, set Group by to your factor — the assumption is normality within groups (of the residuals), not of the pooled data. A mixture of groups with different means can look non-normal even when every group is perfectly normal.
  • W 接近 1 表示与正态一致;p < 0.05 表示显著偏离正态。
  • 偏度:0 为对称;正值为右侧长尾(取对数通常有效);负值为左侧长尾。
  • 若这批数据接下来要做方差分析,请把 Group by 设为你的因子——假设要求的是组内(残差)正态,而不是合并后的数据正态。若干均值不同的组混在一起,即使每一组都完全正态,看上去也会像非正态。

The graph tab shows a normal Q-Q plot: sample values against theoretical normal quantiles. Points along the dashed line = normal; a bow shape = skew; S-shape = heavy or light tails.

图表选项卡给出正态 Q-Q 图:样本值对理论正态分位数作图。点沿虚线排列即为正态;呈弓形表示偏态;呈 S 形表示尾部过重或过轻。

With small n (< ~10 per group) Shapiro–Wilk has little power — “p > 0.05” may just mean “too little data to tell”. Rely on the Q-Q plot, skewness, and what you know about the measurement.样本量小时(每组约不足 10 个),Shapiro–Wilk 的功效很低——"p > 0.05"可能只是意味着"数据太少,判断不了"。此时要依靠 Q-Q 图、偏度,以及你对这项测量本身的了解。

4One-way ANOVA单因素方差分析

Analysis of variance tests whether the mean of a numeric response differs among the levels of one factor. It works by comparing variation between group means with variation within groups: if groups differ more than their internal scatter can explain, the ratio F is large and p is small.

方差分析检验一个数值型响应变量的均值在某个因子的各水平之间是否存在差异。它的做法是把各组均值之间的变异与组内的变异作比较:如果组间差异大到组内离散无法解释的程度,比值 F 就大,p 就小。

Setup: Response = your measurement column; Factor = your grouping column; choose a post-hoc test (below) and α. Each group needs at least 2 observations, and the response must be numeric.

设置:Response 选你的测量值列;Factor 选你的分组列;再选择事后检验(见下)和 α。每组至少要有 2 个观测,响应变量必须是数值型。

The ANOVA table:

方差分析表:

ColumnMeaning
列含义
SS (sum of squares)Amount of variation attributed to each source.
SS(平方和)归因于各来源的变异量。
df (degrees of freedom)Between = k−1 groups; Within = N−k observations.
df(自由度)组间 = k−1(k 为组数);组内 = N−k(N 为观测总数)。
MS (mean square)SS ÷ df — a variance.
MS(均方)SS ÷ df——即一个方差。
FMSbetween ÷ MSwithin.
FMS组间 ÷ MS组内。
p-valueProbability of an F this large if all group means were equal.
p 值若各组均值全部相等,出现这么大的 F 的概率。

A significant ANOVA says “at least one group differs” — it does not say which. That is the job of the post-hoc test.

方差分析显著,说的是"至少有一组不同"——它并不指出是哪一组。那是事后检验的任务。

Reporting: “Nitrate differed significantly among sites (one-way ANOVA, F2,15 = 12.4, p < 0.001); Tukey's HSD showed site C exceeded sites A and B.”报告写法:"各样点间硝酸盐含量差异显著(单因素方差分析,F2,15 = 12.4,p < 0.001);Tukey HSD 表明样点 C 高于样点 A 和 B。"

5Post-hoc tests: which one?事后检验:该选哪个

Post-hoc tests compare every pair of groups while controlling the family-wise error rate (the chance of any false positive across all comparisons). They differ in how strictly they control it:

事后检验对每一对组进行比较,同时控制族系误差率(所有比较中出现任何一个假阳性的概率)。它们的区别在于控制得有多严:

TestCharacterUse when…
方法特点什么时候用
Tukey's HSDBalanced; the standardDefault choice for all-pairwise comparisons. Recommended.
Tukey HSD均衡;业界标准两两全比较的默认选择。推荐。
Fisher's LSDLiberal (no multiplicity adjustment)Only defensible with 3 groups and a significant ANOVA. Finds differences easily — including false ones.
Fisher LSD宽松(不作多重性校正)只有在 3 个组且方差分析已显著时才站得住脚。很容易找出差异——包括假的差异。
BonferroniConservativeFew comparisons; simple and defensible, but loses power with many groups.
Bonferroni保守比较次数少时用;简单且有据可循,但组数一多就损失功效。
SchefféVery conservativeExploring complex contrasts, not just pairs.
Scheffé非常保守用于探索复杂的对比,而不只是两两比较。
Student–Newman–KeulsStep-down; moderately liberalTraditional in agronomy; weaker error control than Tukey.
Student–Newman–Keuls逐步降幂;中等偏宽松农学中的传统做法;误差控制弱于 Tukey。
Duncan's MRTLiberalCommon in older agricultural literature; many statisticians advise against it.
Duncan 新复极差法宽松常见于较早的农业文献;许多统计学家不建议使用。
Games-HowellDoes not assume equal variancesLevene's test failed (unequal variances) — the safe pairwise choice.
Games-Howell不假定方差齐性Levene 检验未通过(方差不齐)时——两两比较的稳妥选择。
Dunnett'sTreatments vs a control onlyYou have a control group and only care about comparisons against it (fewer comparisons = more power).
Dunnett只比较各处理与对照你有一个对照组,且只关心与它的比较(比较次数少 = 功效更高)。

Letter groupings字母标记

Results and bar charts carry compact letter displays: groups sharing a letter are NOT significantly different. So “a, ab, b” means the outer groups differ, and the middle group can't be distinguished from either. Letters start at “a” for the highest mean.

结果和柱状图上都带有紧凑的字母标记:共用同一个字母的组之间没有显著差异。因此"a, ab, b"意味着两端的组彼此不同,而中间那一组与两者都无法区分。字母从均值最高的组开始,记作"a"。

6Levene's test & unequal variancesLevene 检验与方差不齐

With every ANOVA, Kiwistat automatically runs Levene's test (Brown–Forsythe, median-centred — the robust version) on the equal-variance assumption. It is itself an ANOVA on the absolute deviations from each group's median.

每做一次方差分析,Kiwistat 都会自动对方差齐性假设运行 Levene 检验(Brown–Forsythe 中位数中心化版本,即稳健版)。它本身就是对各观测与其组中位数之绝对离差所做的一次方差分析。

  • p ≥ 0.05 — variances are homogeneous; carry on.
  • p < 0.05 — variances differ. The standard F-test can then be misleading, so Kiwistat also reports Welch's ANOVA, which does not assume equal variances — quote Welch's F and its (fractional) degrees of freedom instead — and recommends the Games-Howell post-hoc. If the spread grows with the mean (very common), ticking Use log₁₀ values for statistics on the response column often fixes the problem at the source; re-run and check Levene again.
  • p ≥ 0.05——方差齐性成立,可以继续。
  • p < 0.05——方差不齐。此时标准的 F 检验可能产生误导,因此 Kiwistat 会同时报告不假定方差齐性的 Welch 方差分析——请改引 Welch 的 F 及其(带小数的)自由度——并推荐使用 Games-Howell 事后检验。如果离散程度随均值增大(这非常常见),在响应变量列上勾选 Use log₁₀ values for statistics 往往能从根源上解决问题;重新运行一次,再看 Levene 的结果。

7Two-way ANOVA双因素方差分析

Tests two factors at once and — the real payoff — their interaction:

同时检验两个因子,而真正的收获在于它们的交互作用:

  • Main effect A: averaged over B, do A's levels differ?
  • Main effect B: averaged over A, do B's levels differ?
  • A × B interaction: does the effect of one factor depend on the level of the other? (e.g. fertiliser boosts growth in species 1 but not species 2).
  • 主效应 A:在 B 上平均之后,A 的各水平之间有差异吗?
  • 主效应 B:在 A 上平均之后,B 的各水平之间有差异吗?
  • A × B 交互作用:一个因子的效应是否取决于另一个因子的水平?(例如施肥促进物种 1 的生长,却对物种 2 无效。)
If the interaction is significant, interpret the main effects with care — “the average effect of fertiliser” means little when the effect differs by species. Look at the grouped bar chart and describe the pattern of cell means.如果交互作用显著,解释主效应时就要格外小心——当效应因物种而异时,"施肥的平均效应"意义不大。请看分组柱状图,描述各单元格均值的格局。

Requirements: replication inside every factor combination, ideally balanced (equal n per cell — Kiwistat's sums-of-squares are exact for balanced designs). Levene's test here checks variances across all combinations. The post-hoc applies to Factor A's main effect.

要求:每一种因子组合内部都要有重复,最好是均衡的(每格 n 相等——Kiwistat 的平方和对均衡设计是精确的)。此处的 Levene 检验核查的是所有组合之间的方差齐性。事后检验作用于因子 A 的主效应。

8Blocked ANOVA (Randomised Complete Block Design)区组方差分析(随机完全区组设计)

Field and glasshouse trials rarely have uniform conditions: soil, light, or time-of-day varies across the experiment. The randomised complete block design groups experimental units into blocks that are internally similar (a field strip, a bench, a sampling day), and applies every treatment once within each block. The analysis then removes the block-to-block variation from the error term, making the treatment comparison much more sensitive.

田间和温室试验的条件很少是均一的:土壤、光照或一天中的时段都会在试验范围内变化。随机完全区组设计把试验单元分成内部相似的若干区组(一条田间条带、一张试验台、一个采样日),并在每个区组内把每种处理各施用一次。分析时再把区组之间的变异从误差项中剔除,使处理间的比较灵敏得多。

Setup: a numeric response, a treatment factor, and a block factor — each treatment should appear once in each block. The results table has three rows:

设置:一个数值型响应变量、一个处理因子和一个区组因子——每种处理应在每个区组中各出现一次。结果表有三行:

  • Treatment — the effect you care about; report this F and p.
  • Block — a significant block effect confirms blocking was worthwhile (it soaked up real variation).
  • Error — the residual scatter, now free of block differences.
  • Treatment(处理)——你关心的效应;报告这一行的 F 和 p。
  • Block(区组)——区组效应显著,说明区组化确实值得(它吸收掉了真实存在的变异)。
  • Error(误差)——残余的离散,现已剔除了区组间差异。

Post-hoc comparisons and letter groupings apply to the treatment means, exactly as in one-way ANOVA. The example dataset “Wheat yield trial (RCBD)” lets you compare a blocked analysis against a naïve one-way ANOVA on the same data — the blocked test has a smaller error and a sharper treatment result.

事后比较和字母标记作用于各处理的均值,与单因素方差分析完全一样。示例数据集"Wheat yield trial (RCBD)"可以让你在同一份数据上,把区组分析与不加区组的单因素方差分析作比较——区组分析的误差更小,处理效应也更清晰。

RCBD assumes no treatment × block interaction (there is only one observation per combination, so it cannot be separated from error). If you have replication within each combination, use two-way ANOVA instead.RCBD 假定不存在"处理 × 区组"交互作用(每种组合只有一个观测,因此它无法与误差分离)。如果每种组合内部有重复,请改用双因素方差分析。

9Repeated measures ANOVA重复测量方差分析

When you measure the same subject more than once — a plant every fortnight, a plot each season, a patient before and after treatment — the measurements are not independent. Two readings from the same plant are more alike than two readings from different plants, and an ordinary ANOVA, which assumes every observation is independent, gets the error term badly wrong. A repeated measures ANOVA removes each subject's own level from the error, exactly as blocking does — in fact a simple repeated measures ANOVA is arithmetically an RCBD with subjects as the blocks.

当你对同一个体反复测量时——每两周测一次同一株植物、每季测一次同一块样地、治疗前后测同一位患者——这些测量并不独立。同一株植物的两次读数,比两株不同植物的读数更相似;而普通方差分析假定每个观测都独立,会把误差项算得离谱。重复测量方差分析把每个个体自身的水平从误差中剔除,做法与区组化如出一辙——事实上,最简单的重复测量方差分析在算术上就是一个以个体为区组的 RCBD。

How to lay the data out数据该怎么排

Kiwistat expects wide format: one row per subject, and one column per occasion. This is how time-series data usually leaves a spreadsheet or a logger.

Kiwistat 要求宽表格式:每个个体一行,每个测量时点一列。时间序列数据从电子表格或数据记录仪导出时,通常就是这个样子。

PlantTreatmentDay 0Day 7Day 14Day 21
植株处理第 0 天第 7 天第 14 天第 21 天
P01Control4.25.66.98.0
P02Control3.85.16.47.4
P07Low N4.16.08.110.0

Setup: tick the measurement columns in time order; optionally name a subject/plot ID column (used only for labelling) and a between-subjects factor such as treatment. Rows missing any occasion are dropped whole — the design needs a complete set per subject — and the count of dropped rows is reported.

设置:按时间顺序勾选各测量列;可另外指定个体/样地编号列(仅用于标注)和一个被试间因子(如处理)。凡缺少任一时点的行会被整行剔除——该设计要求每个个体都有完整的一套数据——被剔除的行数会一并报告。

Reading the table如何读这张表

  • Time (occasion) — does the measurement change over the series? This is the within-subjects effect, tested against the time × subjects error.
  • Subjects — the variation between individuals that has been taken out of the error. It is not usually a hypothesis of interest, but a large value shows why pairing mattered.
  • With a between-subjects factor the table splits in two. Between subjects tests the treatment against subject-to-subject variation; within subjects tests time and the Time × treatment interaction. That interaction is usually the real question: do the groups follow different trajectories?
  • Time(时点)——测量值在整个序列中是否发生变化?这是被试内效应,用"时点 × 个体"的误差来检验。
  • Subjects(个体)——已从误差中剔除的个体间变异。它通常不是你关心的假设,但数值很大就说明配对为什么重要。
  • 加入被试间因子后,表格会一分为二。被试间部分用个体间变异来检验处理效应;被试内部分检验时点以及时点 × 处理交互作用。那个交互作用往往才是真正的问题:各组是否走了不同的轨迹?

Sphericity — the assumption that catches people out球形度——最容易栽跟头的那条假设

Repeated measures ANOVA assumes sphericity: every pair of occasions has the same variance of differences. Time series routinely break it, because measurements close together are more alike than measurements far apart. When sphericity fails, the F-test is too liberal — p-values come out smaller than they should be.

重复测量方差分析假定球形度:任意两个时点之差的方差都相同。时间序列常常违反这一点,因为时间相近的测量比时间相隔远的更相似。球形度不成立时,F 检验会过于宽松——p 值会比应有的更小。

  • Mauchly's test checks it. A significant result (p < 0.05) means sphericity is violated.
  • Greenhouse–Geisser and Huynh–Feldt fix it by multiplying the degrees of freedom by an epsilon (ε ≤ 1). Kiwistat reports both, with fractional df — that is expected, not a bug.
  • Convention: use Greenhouse–Geisser when ε < 0.75 and Huynh–Feldt when it is larger. Many authors report Greenhouse–Geisser regardless, because Mauchly's test has little power at small n.
  • With only two occasions there is a single difference score, so sphericity cannot be violated and no correction is needed (that case is equivalent to a paired t-test).
  • Mauchly 检验用来检查它。结果显著(p < 0.05)即表示球形度被违反。
  • Greenhouse–Geisser 与 Huynh–Feldt 通过给自由度乘上一个 epsilon(ε ≤ 1)来校正。Kiwistat 两者都报告,自由度会带小数——这是正常的,不是 bug。
  • 惯例:ε < 0.75 时用 Greenhouse–Geisser,更大时用 Huynh–Feldt。许多作者一律报告 Greenhouse–Geisser,因为 Mauchly 检验在小样本下功效很低。
  • 只有两个时点时,差值只有一组,球形度不可能被违反,也就不需要校正(这种情形等价于配对 t 检验)。
Reporting: “Seedling height changed over time (F1.05, 15.7 = 6178, p < 0.001, Greenhouse–Geisser corrected), and the trajectories differed between fertiliser treatments (Time × Treatment F2.09, 15.7 = 385, p < 0.001).” Always state which correction you applied.报告写法:"幼苗株高随时间发生变化(F1.05, 15.7 = 6178,p < 0.001,经 Greenhouse–Geisser 校正),且不同施肥处理的变化轨迹不同(时点 × 处理 F2.09, 15.7 = 385,p < 0.001)。"务必写明你用的是哪一种校正。
Post-hoc comparisons between occasions use the within-subject error term, so they are paired comparisons — but they are still all-pairwise. If your real question is “which times differ from the start?”, choose Dunnett's with the baseline occasion as the control: three comparisons instead of six, and correspondingly more power.时点之间的事后比较使用被试内误差项,因此属于配对比较——但它们仍然是两两全比较。如果你真正想问的是"哪些时点与起始时点不同?",那就选 Dunnett,把基线时点当作对照:比较次数由六次降为三次,功效也相应提高。

The graph is a profile plot: mean ± SE at each occasion, joined, with one line per between-subjects group. Parallel lines mean no interaction; converging or crossing lines are the interaction made visible. The example dataset “Seedling height over time” is set up for this analysis.

图是一张轮廓图:各时点的"均值 ± 标准误"依次连线,每个被试间分组一条线。线条平行表示没有交互作用;线条收敛或交叉,就是交互作用的可视化形态。示例数据集"Seedling height over time"正是为这项分析准备的。

10Non-parametric tests非参数检验

When data are skewed, ordinal, riddled with outliers, or simply won't transform to normality — common with counts and concentrations — switch to a rank-based test. These convert values to ranks and ask whether one group tends to have higher ranks than another. They make no normality assumption and resist outliers.

当数据偏态、属于等级数据、离群值很多,或者怎么变换都达不到正态时——计数和浓度数据常常如此——请改用基于秩次的检验。它们把数值转换为秩,然后追问某一组的秩是否系统性地高于另一组。这类方法不作正态性假定,也不易受离群值影响。

Kruskal–Wallis (3 or more groups)Kruskal–Wallis(3 组及以上)

The non-parametric counterpart of one-way ANOVA. It reports a tie-corrected H statistic (compared to a chi-square distribution) and a p-value. A significant result means at least one group's distribution differs. Dunn's test then compares each pair with a Bonferroni adjustment, giving the familiar letter groupings on a boxplot. Report the medians (not means) as your summary.

单因素方差分析的非参数对应方法。它给出经结点校正的 H 统计量(与卡方分布比较)和 p 值。结果显著意味着至少有一组的分布不同。随后的 Dunn 检验对每一对做比较并作 Bonferroni 校正,在箱线图上给出我们熟悉的字母标记。汇总时请报告中位数(而不是均值)。

Mann–Whitney U (exactly 2 groups)Mann–Whitney U(恰好 2 组)

The non-parametric counterpart of a two-sample t-test. It asks whether values in one group are systematically larger than in the other. Kiwistat reports U, a z-approximation, and a two-sided p-value with tie and continuity corrections (reliable for roughly n ≥ 8 per group).

两样本 t 检验的非参数对应方法。它追问一组的数值是否系统性地大于另一组。Kiwistat 报告 U 值、z 近似,以及经结点校正和连续性校正的双侧 p 值(每组约 n ≥ 8 时可靠)。

Reporting: “Mayfly abundance differed among habitats (Kruskal–Wallis H = 15.8, df = 2, p < 0.001); Dunn's tests showed riffles > runs > pools.” Quote medians and quartiles.报告写法:"蜉蝣多度在各生境间存在差异(Kruskal–Wallis H = 15.8,df = 2,p < 0.001);Dunn 检验显示急流段 > 平流段 > 深潭。"请引用中位数和四分位数。
A non-significant result may mean “no difference” or “too little data”. And these tests compare whole distributions — if two groups have very different shapes, a significant result isn't only about the median. Boxplots (shown automatically) let you see what's driving it.结果不显著,可能意味着"没有差异",也可能意味着"数据太少"。此外,这类检验比较的是整个分布——若两组的分布形状差别很大,显著的结果就不只是中位数的问题。自动给出的箱线图能让你看清究竟是什么在起作用。

11Linear regression线性回归

Where correlation measures how tightly two variables move together, regression fits the actual line y = a + bx and quantifies it: the slope b is the change in y per unit x. Kiwistat reports the intercept and slope with standard errors, t-tests and p-values, R², and a 95% confidence interval for the slope. The key test is whether the slope differs from zero (p for b).

相关分析衡量的是两个变量同向变化的紧密程度,而回归则拟合出那条具体的直线 y = a + bx 并加以量化:斜率 b 就是 x 每变化一个单位时 y 的变化量。Kiwistat 报告截距和斜率及其标准误、t 检验与 p 值、R²,以及斜率的 95% 置信区间。关键的检验是斜率是否不等于零(b 的 p 值)。

Confidence vs prediction bands置信带与预测带

On the graph tab you can show two shaded bands around the line:

在图表选项卡上,可以在直线周围显示两条带阴影的区间:

  • Confidence band — where the true regression line probably lies (narrow).
  • Prediction band — where a new individual observation will probably fall (wide, because it also includes scatter around the line).
  • 置信带——真实的回归线大概落在哪里(较窄)。
  • 预测带——一个新的单个观测大概会落在哪里(较宽,因为它还包含了围绕直线的离散)。

Diagnostic plots — always look at these诊断图——一定要看

R² tells you how well the line fits, but not whether a line is appropriate. Two diagnostic panels appear beneath the fit:

R² 告诉你这条直线拟合得多好,但不告诉你用直线是否合适。拟合结果下方会出现两张诊断图:

  • Residuals vs fitted — should be a shapeless horizontal band around zero. A U or hump shape means the relationship is curved (try curve fitting or a transform); a widening fan means variance grows with x (try log-transforming y).
  • Residuals vs leverage — points plotted in red have a Cook's distance greater than 4/n, meaning they individually pull the line towards themselves. Check them for data-entry errors or genuine influential observations before trusting the fit.
  • 残差—拟合值图——应当是围绕零的一条没有形状的水平带。呈 U 形或驼峰形,说明关系是弯曲的(试试曲线拟合或作变换);呈向外张开的喇叭形,说明方差随 x 增大(试试对 y 取对数)。
  • 残差—杠杆值图——标为红色的点,其 Cook 距离大于 4/n,意味着它们会单独把直线拉向自己。相信这个拟合之前,请先检查它们是录入错误还是真正有影响力的观测。
Regression assumes independent observations, a linear relationship, roughly constant variance, and approximately normal residuals. Never extrapolate beyond the range of your x data, and remember that a significant slope shows association, not causation.回归假定观测相互独立、关系为线性、方差大致恒定、残差近似正态。绝不要外推到 x 的数据范围之外;并且请记住,斜率显著说明的是相关,不是因果。

12Curve fitting曲线拟合

Fits a model of Y against X by least squares and reports the equation, R² (fraction of the variation in Y explained) and adjusted R² (penalised for extra parameters — use this to compare models).

用最小二乘法拟合 Y 关于 X 的模型,并报告方程、R²(Y 的变异中被解释的比例)以及校正 R²(对多余参数作了惩罚——比较模型时用这一个)。

ModelFormTypical use
模型形式典型用途
Lineary = a + bxFirst choice; b is the rate of change.
线性y = a + bx首选;b 即变化率。
Quadratic / CubicpolynomialsCurvature, optima. Beware overfitting few points.
二次/三次多项式用于曲率、极值点。点数少时当心过拟合。
Exponentialy = aebxGrowth/decay; requires y > 0.
指数y = aebx增长/衰减;要求 y > 0。
Logarithmicy = a + b·ln(x)Rapid rise then plateau; requires x > 0.
对数y = a + b·ln(x)先快速上升后趋于平缓;要求 x > 0。
Powery = axbAllometric scaling; requires x, y > 0.
幂函数y = axb异速生长标度;要求 x、y 均 > 0。

Auto picks the best adjusted R², but prefer a model with a mechanistic justification over a marginally better empirical fit — and never extrapolate beyond your data range.

Auto 会挑选校正 R² 最高的模型;但与其取一个经验拟合只好一点点的模型,不如选一个有机理依据的——而且绝不要外推到数据范围之外。

13Correlation matrix相关矩阵

Pairwise correlation coefficients: r = +1 (perfect positive), 0 (no association), −1 (perfect negative). Bold cells are significant (p < 0.05); each pair uses all rows where both values are present. Choose the method with the Method selector:

两两之间的相关系数:r = +1(完全正相关)、0(无关联)、−1(完全负相关)。加粗的单元格表示显著(p < 0.05);每一对都使用两个变量同时有值的全部行。用 Method 选择器挑选方法:

  • Pearson (default) measures linear association — a strong curved relationship can still give r ≈ 0, and one outlier can create or destroy it. Best for roughly normal data. Plot first.
  • Spearman correlates the ranks instead of the values. It captures any monotonic relationship (consistently increasing or decreasing, even if curved), makes no normality assumption, and shrugs off outliers — a safer default for skewed environmental data.
  • Pearson(默认)衡量的是线性关联——一个很强的曲线关系仍可能得出 r ≈ 0,而单个离群值就能制造或抹掉相关。适用于大致正态的数据。请先作图看看。
  • Spearman 相关的是秩次而非数值。它能捕捉任何单调关系(始终上升或始终下降,即使是弯曲的),不作正态性假定,也不受离群值干扰——对偏态的环境数据而言,这是更稳妥的默认选择。
  • Correlation is not causation — both variables may follow a third (in environmental data, often temperature, depth or season).
  • With many variables, some cells will be “significant” by chance (≈1 in 20 at α = 0.05).
  • 相关不等于因果——两个变量可能都跟着第三个变量走(在环境数据中,常常是温度、深度或季节)。
  • 变量一多,总会有一些格子碰巧"显著"(α = 0.05 时约每 20 个中有 1 个)。

14Principal Components Analysis主成分分析

PCA condenses many correlated variables into a few new axes (principal components) that capture as much of the variation as possible. Kiwistat standardises each variable first (correlation-matrix PCA), so variables with big units don't dominate.

主成分分析把许多相互关联的变量压缩为少数几个新轴(主成分),使之尽可能多地承载原有的变异。Kiwistat 会先对每个变量作标准化(即基于相关矩阵的 PCA),以免量纲大的变量占据主导。

  • Eigenvalues / % variance: how much variation each PC explains. An eigenvalue > 1 means the PC explains more than one original variable's worth. If PC1+PC2 explain, say, 70%+, the biplot is a faithful summary.
  • Loadings (eigenvectors): how strongly each original variable contributes to each PC — use them to name the axes (“PC1 = overall nutrient enrichment”).
  • Scores plot: each sample plotted on PC1–PC2; samples that plot together have similar overall profiles. Choose a grouping column to draw 95% ellipses per group.
  • Rows missing any selected variable are dropped. The sign of an axis is arbitrary — “left” vs “right” has no meaning by itself.
  • 特征值/方差百分比:每个主成分解释了多少变异。特征值 > 1 意味着该主成分解释的变异超过一个原始变量的份量。若 PC1 + PC2 解释了比如 70% 以上,那么双标图就是一份忠实的概括。
  • 载荷(特征向量):每个原始变量对各主成分的贡献强度——可据此为坐标轴命名("PC1 = 总体养分富集程度")。
  • 得分图:每个样品在 PC1–PC2 平面上的位置;落在一起的样品,其整体特征谱相似。选定一个分组列,可为每组绘制 95% 置信椭圆。
  • 缺少任一所选变量的行会被剔除。坐标轴的正负号是任意的——单看"左"或"右"本身没有意义。

15Power analysis & sample size功效分析与样本量

The best time to think about statistics is before you collect data. Power is the probability that your study will detect an effect of a given size if it is really there. An underpowered study wastes effort: a real effect goes undetected, and a non-significant result becomes uninterpretable (“no effect” or “not enough data”?). The convention is to design for 80% power.

考虑统计问题的最佳时机,是在采集数据之前。功效是指:若某个给定大小的效应确实存在,你的研究能把它检出来的概率。功效不足的研究纯属白费力气:真实的效应被漏掉,而不显著的结果也变得无法解释(究竟是"没有效应"还是"数据不够"?)。惯例是按 80% 的功效来设计。

Kiwistat handles two designs — two groups (t-test) and one-way ANOVA with k groups. Provide:

Kiwistat 支持两种设计——两组(t 检验)和 k 组的单因素方差分析。你需要提供:

  • The effect size you need to detect, either directly (Cohen's d for two groups, f for ANOVA) or as your expected group means and within-group SD — take these from a pilot study, the literature, or the smallest difference that would matter biologically.
  • α (usually 0.05) and your target power (usually 0.8).
  • Optionally a planned n per group, to read off the power you'd achieve.
  • 你需要检出的效应量,可以直接给(两组用 Cohen's d,方差分析用 f),也可以给出预期的各组均值和组内标准差——这些数值可取自预试验、文献,或者生物学上值得关心的最小差异。
  • α(通常 0.05)和你的目标功效(通常 0.8)。
  • 也可以再给一个每组计划的 n,以便读出该样本量下能达到的功效。

The output gives the required n per group, the power of your planned n, and a power curve showing how power rises with sample size. Guideline effect sizes: d 0.2 small / 0.5 medium / 0.8 large; f 0.1 / 0.25 / 0.4.

输出会给出每组所需的 n、你计划的 n 所对应的功效,以及一条显示功效如何随样本量上升的功效曲线。效应量的经验参考:d 0.2 小/0.5 中/0.8 大;f 0.1/0.25/0.4。

After running a regression or ANOVA, you can read the residual SD from the output and feed it (with the means you hope to see) straight back into a power analysis to plan the follow-up experiment.跑完回归或方差分析之后,可以从输出中读出残差标准差,连同你希望看到的均值一起直接代入功效分析,用来规划后续试验。

16Exploring data: boxplots & scatterplot matrix数据探索:箱线图与散点图矩阵

Two graph types under Test Type produce figures with no hypothesis test — for looking before you leap. Always explore your data this way first: it reveals skew, outliers, unequal spread and non-linear relationships that decide which formal test is appropriate.

Test Type 下有两种图,只出图、不做假设检验——供你在动手之前先看一看。请务必先这样探索数据:它能揭示偏态、离群值、离散不齐和非线性关系,而这些正决定了该用哪种正式检验。

Boxplots箱线图

For one numeric variable, optionally split by a factor. The box spans the interquartile range (middle 50% of the data), the heavy line is the median, the whiskers reach the most extreme values within 1.5 × IQR of the quartiles, and points beyond are drawn as outliers. Tick Show raw data points in the graph options to overlay every observation (jittered) — increasingly expected by journals. A median sitting off-centre in its box signals skew; boxes of very different heights signal unequal variances.

针对一个数值变量,可选按某个因子分组。箱体跨越四分位距(中间 50% 的数据),粗线为中位数,须线延伸到距四分位点 1.5 × IQR 以内的最极端值,超出者绘为离群点。在图表选项中勾选 Show raw data points,可把每一个观测(加抖动)叠加上去——越来越多期刊要求这样做。中位数在箱体内偏向一侧,提示存在偏态;各箱体高度相差悬殊,提示方差不齐。

Scatterplot matrix散点图矩阵

For several numeric variables, this grid plots every pair against each other at once. Scan it for straight-line trends (candidates for regression or correlation), curves (curve fitting), clusters, and outliers — the quickest way to get to know a multivariate dataset.

针对多个数值变量,这张网格图一次性把每一对变量互相对作散点图。扫一眼就能找出直线趋势(回归或相关分析的候选)、曲线(曲线拟合)、聚集和离群值——这是快速熟悉一份多变量数据集最有效的办法。

Show raw data points is also available on ANOVA bar charts, overlaying the individual observations on each bar so readers see the real spread behind the mean.方差分析的柱状图同样提供 Show raw data points,把每根柱子背后的单个观测叠加上去,让读者看到均值之下真实的离散程度。

17Reading & reporting results结果的解读与报告

  • Red p-values are significant at 0.05. Report exact values (“p = 0.003”), reserving “p < 0.001” for very small ones.
  • Always report the test, the statistic with its degrees of freedom, the p-value, and n — e.g. “F2,15 = 13.3, p < 0.001, n = 6 per site”.
  • Mention the assumption checks: “variances were homogeneous (Levene's test, p = 0.32)” or “variances were unequal, so Welch's ANOVA and Games-Howell comparisons were used”.
  • If you used log statistics, say so: “data were log₁₀-transformed for analysis; means shown are geometric”.
  • The Sig. figs box in the app's header controls display rounding everywhere; underlying values keep full precision.
  • 红色的 p 值表示在 0.05 水平上显著。请报告确切数值("p = 0.003"),"p < 0.001"只留给确实极小的情形。
  • 务必报告所用检验、统计量及其自由度、p 值和 n——例如"F2,15 = 13.3,p < 0.001,每个样点 n = 6"。
  • 要提到前提假设的检查:"方差齐性成立(Levene 检验,p = 0.32)",或者"方差不齐,故采用 Welch 方差分析和 Games-Howell 比较"。
  • 若统计时取了对数,要写明:"分析前数据经 log₁₀ 变换;所示均值为几何均值"。
  • 应用页首的 Sig. figs 框控制各处显示的有效位数;底层数值仍保留完整精度。

For a ready-made methods paragraph describing Kiwistat's statistical implementation, see Describing Kiwistat in publications in the User Guide.

若需要一段现成的、描述 Kiwistat 统计实现方式的方法段落,见使用指南中的在论文中如何描述 Kiwistat。

18Glossary术语表

术语表条目保留英文原文——这些是你在软件输出和论文中会实际遇到的写法。

α (alpha)
The significance threshold, usually 0.05: the false-positive rate you accept.
ANOVA
Analysis of variance — compares means of 2+ groups via a variance ratio (F).
Degrees of freedom (df)
The number of independent pieces of information behind a statistic; quoted with F and t.
F statistic
Ratio of between-group to within-group variance; large F → group means differ.
Factor
A categorical explanatory variable (site, species, treatment). Its values are levels.
Block
A group of experimental units treated as internally uniform (field strip, day, bench); blocking removes their shared variation from the error.
Censored value
A measurement known only to be below (or above) a limit, e.g. “<0.05”; substituted for analysis per a stated rule.
Cook's distance
How much a single point influences a regression fit; values > 4/n flag influential points.
Effect size
The magnitude of a difference or relationship, independent of sample size (e.g. Cohen's d, f, r).
Family-wise error rate
Chance of at least one false positive across a set of comparisons; post-hoc tests control it.
Geometric mean
Back-transformed mean of logs; the natural “average” for log-scale data.
Interaction
When one factor's effect depends on another factor's level.
IQR (interquartile range)
Q3 − Q1, the spread of the middle 50% of the data; the height of a box in a boxplot.
Kruskal–Wallis test
Rank-based (non-parametric) alternative to one-way ANOVA; reports H.
Leverage
How far a point's x-value is from the mean; high-leverage points can dominate a regression.
Levene's test
Tests equality of variances between groups (Kiwistat uses the robust Brown–Forsythe form).
Mann–Whitney U
Rank-based (non-parametric) alternative to a two-sample t-test.
Non-parametric test
A test that works on ranks and makes no distributional assumption (Kruskal–Wallis, Mann–Whitney, Spearman).
Null hypothesis (H₀)
The “no effect” starting assumption that a test tries to reject.
p-value
Probability of data at least this extreme if H₀ is true.
Post-hoc test
Pairwise comparisons after a significant ANOVA, adjusted for multiplicity.
Power
The probability a test detects a real effect; grows with n and effect size.
Prediction band
Range in which a new individual observation is expected to fall around a regression line (wider than the confidence band).
Q-Q plot
Sample quantiles vs theoretical normal quantiles; straight line = normal data.
R²
Fraction of Y's variation explained by a fitted model.
RCBD
Randomised complete block design — one treatment factor applied once within each block.
Residual
Observation minus its group mean (or fitted value); the “noise” a model doesn't explain.
Shapiro–Wilk test
Formal test of normality; W near 1 and p ≥ 0.05 = consistent with normal.
Slope (b)
In regression, the change in y per one-unit change in x.
Spearman correlation
Correlation of ranks; captures any monotonic relationship, resistant to outliers.
Standard error (SE)
SD ⁄ √n — the precision of a mean.
Welch's ANOVA
ANOVA variant that does not assume equal variances.

Kiwistat · KiwiScience — this guide is reachable from Help → Statistics guide in the app and from the “?” buttons next to each test. The application itself is covered in the User Guide.

Kiwistat · KiwiScience — 本指南可从应用内的 Help → Statistics guide 以及每项检验旁的 "?" 按钮进入。应用本身则在使用指南中讲解。

↑ Back to top↑ 返回顶部