Lecture 7: Why Are Species Where They Are? Niche and Neutral Theories of Community Assembly

Authors
Affiliation

Alex Matthew

University of the Western Cape

Keanan Jarvis

University of the Western Cape

Ethan Bell

University of the Western Cape

Published

September 8, 2026

Last updated

August 31, 2026

TipMaterial Required for This Lecture
Type Name Notes
Theory Hubbell (2001) The Unified Neutral Theory of Biodiversity and Biogeography Primary statement of neutral theory
Theory Hutchinson (1957) Concluding remarks The niche concept defined
Theory Adler et al. (2007) A niche for neutrality The niche-neutral continuum
Theory Cottenie (2005) Integrating environmental and spatial processes Metacommunity framework
Theory Legendre et al. (2015) Should the Mantel test be used in spatial analysis? Mantel test critique
Data Doubs River fish community Doubs.RData via NEwR companion data
Data Oribatid mite community vegan::mite, vegan::mite.env, vegan::mite.xy
Package spesim 0.5.2 remotes::install_github("ajsmit/spesim")
ImportantTasks to Complete in This Lecture

See the Tasks section at the end of this lecture for five exercises that build directly on the analyses presented here.

The truth is rarely pure and never simple.

Oscar Wilde

Introduction

Cluster Analysis left a question open. The Doubs fish community forms a real, measurable gradient running from headwater to lowland, and the silhouette analysis showed that forcing it into discrete clusters loses information a continuous ordination preserves. But naming a gradient is not explaining one. Why does the community change in an orderly way along the river, rather than scattering species at random? This lecture takes up that question.

TipMain Idea: Pattern Is Not Process

A gradient in community composition is a pattern. Niche and neutral theory are rival explanations for the process that produced it. In general a single pattern is consistent with several processes, which is why the question cannot be settled by ordination or clustering alone, however well executed. It requires a different kind of evidence.

Ecological theory offers two long-standing answers. Niche theory holds that species occur where the local environment and biotic interactions permit them to persist, so that community structure follows deterministically from the ecological differences between species. Neutral theory holds that at the level of individual births, deaths and dispersal events, species are functionally interchangeable, and that the patterns we observe arise from stochastic demographic and dispersal processes rather than from niche differences at all. Both theories can produce gradients that, at first glance, look identical.

By the end of this lecture you should be able to:

  • state the core assumption each theory makes about species, and explain why those assumptions lead to different predictions about community structure;
  • explain why this remains an active area of debate, and identify what is empirically established as against what is contested;
  • translate each theory into a falsifiable prediction about how community dissimilarity should behave across space;
  • recognise why the most intuitive statistical test of those predictions is the wrong one, and explain the reasoning behind that judgement;
  • apply a defensible alternative to the Doubs data and interpret it critically; and
  • connect this theoretical lens to the advanced workflow introduced in the BCB743 Model Building chapter.

From Pattern to Process: Revisiting the Doubs Gradient

The Cluster Analysis chapter established, with a cophenetic correlation and a silhouette analysis, that the Doubs fish community is organised as a continuum rather than a set of discrete types. The Ordination chapter, several chapters earlier, established the same thing from the opposite direction: the dominant axis recovers an upstream-to-downstream gradient, with cool-water species at one end and lowland species at the other. Every method applied to this dataset agrees that a gradient exists. None of them asks what generated it.

There is an obvious candidate answer, and it happens to be the answer niche theory formalises. The river changes as it flows, in temperature, oxygen, gradient and flow rate, and different fish species tolerate that changing environment differently. Trout persist near the cool, oxygenated headwaters and cannot survive the warmer, slower lowland reaches; the reverse holds for the lowland species. On this account the gradient in the species data follows directly from the gradient in the environmental data, filtered through each of the species physiological and ecological tolerances.

That is a reasonable hypothesis. It is also not the only one available. A river is a linear system, and linear systems have a second property capable of producing the same pattern in an ordination: every site connects to its neighbours through a fixed, ordered sequence of dispersal pathways. Suppose every fish species were ecologically identical. The bare fact that an individual at one site can reach adjacent sites easily and distant ones only with difficulty would still, over time, produce spatial structure in community composition. This is the situation neutral theory describes formally, and it makes a testable claim: structure can arise from dispersal limitation and demographic stochasticity alone, with no reference to ecological differences between species.

The Doubs data cannot, by inspection, tell these two stories apart. That is the problem this lecture exists to solve.

Ecological Background: Niche Theory

Begin with a single fish. A brown trout holds its position in the cold, fast, oxygen-rich water of the upper Doubs. It is not there by accident. Trout require high dissolved oxygen and cool temperatures; they cannot persist in the warm, sluggish, oxygen-poor water of the lowland reaches. If you move downstream and the trout disappears, replaced by species that tolerate, and in some cases require, precisely the conditions the trout cannot survive. The community changes along the river because the environment changes along the river, and each species occupies the stretch where local conditions fall inside the range it can tolerate. Niche theory formalises that intuition: species occur where the environment permits them to persist, and so community composition can in principle be read from the environment itself.

The concept received its lasting formal statement from Hutchinson (1957), who defined the niche as an n-dimensional hypervolume: a space with one axis for every environmental factor that matters to a species, bounded by the limits within which that species can maintain a viable population. Temperature is one axis, dissolved oxygen another, pH another, prey availability another, for as many dimensions as are ecologically relevant. Chase and Myers (2011) restate the idea directly. Conditions inside the N-dimensional hypervolume of a species’ requirements define the space in which it can gather resources, evade enemies and persist. This means that a species realised distribution is the projection of that hypervolume onto real geography.

Hutchinson drew a distinction that still organizes how ecologists use the concept. The fundamental niche is the full range of conditions under which a species could persist in the absence of competitors and natural enemies. The realised niche is the usually narrower range it actually occupies once biotic interactions are accounted for. A species might be physiologically capable of living across a broad stretch of the river, yet be confined to part of it because a competitor excludes it elsewhere. The distinction matters here because it makes plain that the environment is not the only filter: biotic interactions do some of the sorting too. The logic nevertheless remains deterministic. Something about the match between a species’ traits and its conditions decides where it lives.

TipMain Idea: Environmental Filtering

Under niche theory the environment acts as a filter. At any site, only species whose niche requirements are met by local conditions can persist; the rest are excluded. As conditions change across space the filter changes, and with it the set of species that pass through. Community composition becomes a predictable function of the environment.

The mechanism doing this work is environmental filtering. Chase and Myers (2011) place it right at the centre of the deterministic view of communities, in which local, niche-based processes such as environmental filtering, biotic interactions and interspecific trade-offs largely determine which species are found where. The word to hold onto is Deterministic: given sufficient knowledge of the environment and of species’ tolerances, niche theory says community composition is in principle predictable. Two sites with similar environments should support similar communities; two sites with different environments should support different ones. Under strict niche theory the geographic distance between sites is irrelevant, except insofar as it happens to correlate with environmental change.

That last point is what makes niche theory testable against its rival, and it is better stated as a prediction than as a description.

NotePrediction 1. The Niche Prediction

Community dissimilarity should track environmental distance. If two reaches of the Doubs have nearly identical temperature, oxygen and flow, they should hold nearly the same fish, whether they sit side by side or at opposite ends of the river. What drives two communities apart is the degree to which their environments differ, not how far apart they lie in space.

The theory is strong, and strongly supported. In fish assemblages specifically, species occupy distinct positions along river gradients in ways that track measurable niche axes rather than chance, and the divergence of species’ niches along a river continuum has been documented directly (Troia and Gido 2014). But niche theory has a competitor, and the competitor’s claim is unsettling. The same orderly gradient could arise even if niche differences played no part at all.

Ecological Background: Neutral Theory

Return to the same river and change one assumption. Suppose the fish species of the Doubs are not sorted by their tolerances. Suppose instead that any individual, of any species, has roughly the same chance of surviving, reproducing and dispersing as any other individual, whatever species it belongs to. On that assumption the trout occupies the headwaters because its ancestors happened to be there and its descendants have not yet drifted elsewhere, rather than because it is uniquely suited to them. Community composition becomes a matter of history and chance instead of environmental matching.

This is the premise of neutral theory, given its unified formulation by Hubbell (2001) and reviewed a decade later by Rosindell et al. (2011). Its foundational move is the neutrality assumption: all individuals within a trophic level have the same chances of reproduction and death, regardless of species identity (Rosindell et al. 2011). Species are treated as demographically equivalent, interchangeable at the level of the individual. A community, on this view, is a collection of individuals undergoing birth, death and dispersal at rates that do not depend on species identity.

This is easily mistaken for the claim that species are literally identical, so precision is needed. Neutral theory does not assert that trout and bream are the same organism. It asserts something narrower and stranger. For the purpose of explaining community pattern, the differences between them can be set aside, because those differences do not translate into differences in per capita demographic rates within the community. The assumption is almost certainly false in detail. Real species do differ demographically, and Hubbell knew it. The falsity is the point. Neutral theory functions as a null model. If a theory that assumes away every niche difference can still reproduce the patterns we see in real communities, then those patterns cannot by themselves be taken as evidence that niche differences produced them.

TipMain Idea: Demographic Equivalence and Ecological Drift

Neutral theory assumes individuals are demographically equivalent regardless of species. Community composition then changes through ecological drift, the random walk of species’ abundances under chance births and deaths, constrained by dispersal limitation, the tendency of offspring to establish near their parents. Pattern emerges from stochastic processes rather than from environmental sorting.

Two mechanisms generate the pattern. The first is ecological drift. Because births and deaths carry a random component, the relative abundances of species wander over time, much as allele frequencies drift in population genetics. The parallel is deliberate. Rosindell et al. (2011) trace the idea directly to Kimuras neutral theory of molecular evolution, transposing neutrality from alleles within populations to individuals within communities. Under drift alone some species wander to local extinction and others to dominance, purely by chance, with no reference to fitness.

The second mechanism is dispersal limitation, and it is the one that produces spatial structure. Rosindell et al. (2011) define it as a process that causes the location of an individual to be restricted by the location of its parent: offspring tend to establish near where they were produced. In a river this is concrete rather than abstract. A fish spawning at one reach tends to produce offspring that occupy nearby reaches, not reaches at the far end of the system. Over many generations dispersal limitation alone, with no environmental sorting whatever, will generate spatial structure in community composition, because the drift unfolding at one location is only weakly coupled to the drift unfolding at a distant one. Nearby sites share recent dispersal history and come to resemble one another. Distant sites do not.

This yields neutral theory’s central prediction about spatial pattern, which mirrors niche theory’s exactly.

NotePrediction 2. The Neutral Prediction

Community dissimilarity should track geographic distance. Nearby sites should hold similar communities because individuals move readily between them, and distant sites should hold dissimilar communities because they do not, whether or not their environments are alike. What drives two communities apart is how far apart they sit in space, not how much their environments differ.

The distance-driven decline in similarity has a name. Rosindell et al. (2011) gloss beta-diversity, or distance decay, as the probability, as a function of the distance between two individuals, that they belong to the same species. Distance decay sits at the centre of neutral theory rather than at its margin. It is the signature the theory predicts, and it is the quantity this lecture will learn to measure.

Note

Neutral theory is best read as a null model against which the signal of niche processes can be measured, rather than as a claim that the world is neutral. If a dataset’s patterns can be reproduced without invoking niche differences, those patterns alone cannot demonstrate that niche differences are at work. This is why the theory matters even to ecologists who are confident that niches are important. It raises the standard of evidence.

This section develops dispersal limitation only as a local, within-system mechanism. Regional dispersal dynamics and landscape-scale connectivity belong to the Metacommunity chapter, which cross-references this section for the underlying definition.

Two Worlds, One Pattern

The two theories have now been stated, and they disagree about mechanism while agreeing, awkwardly, about what you would see. Before going further we make that agreement concrete, by building two artificial rivers in which we know the answer.

In the first, species differ. Each has a Gaussian response to an environmental gradient running the length of the transect, with its own optimum and tolerance, exactly as niche theory describes. There is no dispersal at all. In the second, species are identical. There is no environment. Community composition changes only through ecological drift, and individuals disperse only to adjacent sites, exactly as neutral theory describes.

library(here)
library(tidyverse)
library(vegan)
library(patchwork)
Show the simulation code
# A niche world: Gaussian species responses to a gradient, no dispersal.
sim_niche <- function(seed, n_site = 29, S = 27, J = 100, tol = 0.12) {
  set.seed(seed)
  gradient <- seq(0, 1, length.out = n_site)
  optima   <- seq(-0.1, 1.1, length.out = S)
  t(sapply(gradient, function(e) {
    p <- exp(-((e - optima)^2) / (2 * tol^2))
    as.vector(rmultinom(1, J, p / sum(p)))
  }))
}

# A neutral world: demographic equivalence, drift, dispersal to adjacent sites only.
# No environmental variable exists anywhere in this simulation.
sim_neutral <- function(seed, n_site = 29, S = 27, J = 100, gens = 250,
                        m = 0.05, nu = 0.002) {
  set.seed(seed)
  counts <- t(replicate(n_site, as.vector(rmultinom(1, J, rep(1 / S, S)))))
  sites  <- seq_len(n_site)
  for (e in seq_len(n_site * J * gens)) {
    d    <- sample.int(n_site, 1)
    dead <- sample.int(S, 1, prob = counts[d, ])   # a death, at random
    counts[d, dead] <- counts[d, dead] - 1
    u <- runif(1)
    if (u < nu) {                                   # rare immigration from outside
      born <- sample.int(S, 1)
    } else if (u < nu + m) {                        # birth from an ADJACENT site
      nbrs <- intersect(c(d - 1, d + 1), sites)
      nb   <- nbrs[sample.int(length(nbrs), 1)]
      born <- sample.int(S, 1, prob = counts[nb, ])
    } else {                                        # birth from this site
      born <- sample.int(S, 1, prob = counts[d, ])
    }
    counts[d, born] <- counts[d, born] + 1
  }
  counts
}

niche_runs   <- lapply(1:4, sim_niche)
neutral_runs <- lapply(1:4, sim_neutral)

world_mats <- tibble(
  world = rep(c("Niche world", "Neutral world"), each = 4),
  run   = rep(1:4, times = 2),
  mat   = c(niche_runs, neutral_runs)
)
Show the figure code
axis1_of <- function(mat) {
  a <- as.numeric(cmdscale(vegdist(mat, method = "bray"), k = 1))
  a * sign(cor(a, seq_along(a)))          # ordination axis signs are arbitrary
}

world_axes <- world_mats |>
  mutate(axis1 = map(mat, axis1_of)) |>
  select(world, run, axis1) |>
  unnest(axis1) |>
  group_by(world, run) |>
  mutate(site = row_number()) |>
  ungroup()

# the correlation is computed here, not asserted in the caption
world_r <- world_axes |>
  group_by(world, run) |>
  summarise(r = abs(cor(axis1, site)), .groups = "drop")

ggplot(world_axes, aes(site, axis1)) +
  geom_point(aes(colour = site), size = 0.9) +
  geom_text(data = world_r, inherit.aes = FALSE,
            aes(x = 1, y = Inf, label = sprintf("|r| = %.2f", r)),
            hjust = 0, vjust = 1.5, size = 2.3, colour = "grey25") +
  facet_grid(world ~ paste("run", run)) +
  scale_colour_viridis_c(guide = "none") +
  labs(x = "Position along the transect", y = "PCoA axis 1")
Figure 1: Two artificial rivers, four runs each. In the niche world, species respond to an environmental gradient and never disperse. In the neutral world, species are demographically identical, no environment exists, and individuals disperse only between adjacent sites. Each panel shows the first axis of a principal coordinates analysis of Bray-Curtis dissimilarity against position along the transect, annotated with the absolute correlation between the two. The niche world returns the same gradient on every run. The neutral world returns a different one each time, with nothing changed but the random seed.

The niche world behaves itself. On every run the first ordination axis tracks position almost perfectly. Run it a hundred times and you will get the same picture a hundred times, because the environment does not change between runs and the species always respond to it in the same way.

The neutral world is stranger. Its gradient appears and disappears from run to run, with nothing altered but the random seed. On some runs the first axis tracks position almost as tightly as the niche world’s, and an ordination could not tell the two apart. On others the structure largely dissolves. Drift is a random walk, and a random walk sometimes wanders in an orderly-looking way.

Both worlds share one property regardless of seed.

Show the figure code
world_mats |>
  mutate(d = lapply(mat, function(m) {
    tibble(distance = as.vector(dist(seq_len(nrow(m)))),
           bray     = as.vector(vegdist(m, method = "bray")))
  })) |>
  select(world, run, d) |>
  unnest(d) |>
  ggplot(aes(distance, bray)) +
  geom_point(alpha = 0.12, size = 0.5) +
  geom_smooth(method = "loess", se = FALSE, linewidth = 0.5, formula = y ~ x) +
  facet_wrap(~ world) +
  labs(x = "Distance between sites", y = "Bray-Curtis dissimilarity")
Figure 2: Community dissimilarity against distance between sites, for the same eight simulated communities. Both worlds produce distance decay. The neutral world produces it because offspring settle near their parents. The niche world produces it because the environment changes smoothly along the transect, so distant sites are also environmentally different. Distance decay alone therefore identifies neither process.

Both worlds show distance decay, every time. In the neutral world it arises because offspring settle near their parents. In the niche world it arises because the environment changes smoothly along the transect, so sites that are far apart are also environmentally unalike. The signature that neutral theory predicts is produced just as readily by a process containing no dispersal whatsoever.

TipMain Idea: One Dataset Cannot Choose

A single neutral community can mimic a niche gradient closely enough that no ordination would separate them. What distinguishes the two processes is the reproducibility of the pattern across many datasets, not the pattern within one. The niche world gives the same gradient every time. The neutral world gives a different one each time, and only sometimes a gradient at all.

We have one Doubs. We cannot re-run it with a different seed.

The Niche-Neutral Continuum: A Simulation

The synthesis position of this lecture holds that niche and neutral processes operate simultaneously, in proportions that vary from system to system. Adler et al. (2007) formalised this as a spectrum: the pure neutral model is the limiting case where stabilising mechanisms and fitness differences both approach zero. Real communities sit somewhere between. But what does that spectrum actually look like? The Two Worlds simulation showed the two endpoints. This section fills in the middle.

We build a single simulation with a mixing parameter, \(w\), that controls the balance between niche-based and neutral assembly. At \(w = 0\) the community is assembled purely by ecological drift and dispersal limitation, with no environmental filtering at all. At \(w = 1\) the community is assembled purely by environmental filtering, with no stochastic component beyond sampling noise. Between the two, both processes operate simultaneously. The gradient in \(w\) is the Adler continuum made computational.

Show the mixed simulation code
sim_mixed <- function(seed, w, n_site = 29, S = 27, J = 100) {
  set.seed(seed)
  gradient <- seq(0, 1, length.out = n_site)
  optima   <- seq(-0.1, 1.1, length.out = S)
  tol      <- 0.12

  # niche component: probability from Gaussian response
  niche_probs <- t(sapply(gradient, function(e) {
    p <- exp(-((e - optima)^2) / (2 * tol^2))
    p / sum(p)
  }))

  # neutral component: equal probability for all species
  neutral_probs <- matrix(1 / S, nrow = n_site, ncol = S)

  # mixed: weighted combination
  mixed_probs <- w * niche_probs + (1 - w) * neutral_probs

  # sample community from the mixed probabilities
  t(sapply(seq_len(n_site), function(i) {
    as.vector(rmultinom(1, J, mixed_probs[i, ]))
  }))
}

# Run across the continuum
weights <- c(0, 0.25, 0.5, 0.75, 1.0)

continuum_data <- expand_grid(w = weights, run = 1:4) |>
  mutate(
    mat   = lapply(seq_len(n()), function(i) sim_mixed(run[i], w[i])),
    axis1 = lapply(mat, function(m) {
      a <- as.numeric(cmdscale(vegdist(m, "bray"), k = 1))
      a * sign(cor(a, seq_along(a)))
    })
  ) |>
  select(w, run, axis1) |>
  unnest(axis1) |>
  group_by(w, run) |>
  mutate(site = row_number()) |>
  ungroup()

continuum_r <- continuum_data |>
  group_by(w, run) |>
  summarise(r = abs(cor(axis1, site)), .groups = "drop")
Show the figure code
ggplot(continuum_data, aes(site, axis1)) +
  geom_point(aes(colour = site), size = 0.7) +
  geom_text(
    data = continuum_r, inherit.aes = FALSE,
    aes(x = 1, y = Inf, label = sprintf("|r|=%.2f", r)),
    hjust = 0, vjust = 1.5, size = 2, colour = "grey30"
  ) +
  facet_grid(paste("run", run) ~ paste0("w = ", w)) +
  scale_colour_viridis_c(guide = "none") +
  labs(x = "Position along the transect", y = "PCoA axis 1") +
  theme(strip.text = element_text(size = 8))
Figure 3: The niche-neutral continuum as a simulation. Each column is a different mixing weight w, from pure neutral (w = 0, left) to pure niche (w = 1, right), with four independent runs per weight. As the niche component strengthens, the gradient becomes both stronger and more reproducible. The transition is not abrupt — it is gradual, which is the Adler continuum made visible.
Show the figure code
continuum_summary <- continuum_r |>
  group_by(w) |>
  summarise(mean_r = mean(r), min_r = min(r), max_r = max(r), .groups = "drop")

ggplot(continuum_summary, aes(w, mean_r)) +
  geom_ribbon(aes(ymin = min_r, ymax = max_r), fill = "grey85", alpha = 0.5) +
  geom_line(linewidth = 0.6) +
  geom_point(size = 1.5) +
  scale_x_continuous(breaks = weights, labels = weights) +
  labs(x = "Mixing weight w (0 = pure neutral, 1 = pure niche)",
       y = "|r| between PCoA axis 1 and position") +
  annotate("text", x = 0.05, y = 0.15, label = "Neutral\nend", size = 2.5, hjust = 0, colour = "grey50") +
  annotate("text", x = 0.95, y = 0.95, label = "Niche\nend", size = 2.5, hjust = 1, colour = "grey50")
Figure 4: Mean and range of |r| across four runs at each mixing weight. As the niche component increases, the mean gradient strength rises and the variability between runs decreases. At w = 0 (pure neutral), the gradient strength is unpredictable. At w = 1 (pure niche), it is nearly deterministic. The transition between the two is smooth — there is no threshold.

Two patterns emerge from the simulation, and they correspond to two of the lecture’s central claims.

First, as the niche component strengthens, the gradient becomes stronger. Mean \(|r|\) rises from the neutral baseline toward the niche ceiling. This is the environmental filtering signal becoming visible in the ordination.

Second, and more important, as the niche component strengthens the gradient also becomes more reproducible. The ribbon in the summary plot narrows from left to right: at \(w = 0\) the four runs scatter widely, because drift is a random walk. At \(w = 1\) the four runs converge on the same result, because the environment is deterministic. Between the two, the variability shrinks gradually. There is no threshold at which the community “switches” from neutral to niche. The transition is smooth.

This is the Adler continuum visualised. A real community sits somewhere on this spectrum, and the position it occupies determines both the strength of the gradient and the predictability of the pattern. The Doubs, with its strong and orderly gradient, sits toward the right — but how far right cannot be determined from one dataset alone. That is why the lecture closes as it does: with an honest report of what the design can and cannot show.

TipMain Idea: The Continuum Is Not a Metaphor

The niche-neutral continuum is a quantitative property of a community that can, in principle, be estimated. The mixing weight \(w\) controls both the strength and the reproducibility of community structure. At \(w = 0\) the community is unpredictable; at \(w = 1\) it is fully determined by the environment. Real communities live somewhere in between, and the position varies with spatial scale, organism, and the environmental range sampled. This simulation makes the abstract position of Adler et al. (2007) and Vellend (2010) concrete and testable.

The Niche-Neutral Debate: What Is Settled, What Is Contested, and Where This Lecture Stands

It would be convenient if one theory were correct and the other wrong, because then this lecture could tell you which test to run and which answer to expect. The reality is more interesting, and learning to hold it clearly is part of what this lecture is for. Niche and neutral theory are not, in the end, a straight contest with a winner. Neither are they interchangeable. The task is to say precisely what ecology has settled, what it is still arguing about, and where a working ecologist should stand while the argument continues.

What is established

Two things are no longer seriously in dispute.

Both processes are real, and both leave signatures in real communities. Environmental filtering demonstrably sorts species along gradients, and the fish of a river do occupy positions that track their physiological tolerances (Troia and Gido 2014). Dispersal limitation demonstrably produces spatial structure independent of environment, which is why community similarity so often decays with distance even where the environment is uniform (Rosindell et al. 2011). Neither theory is empirically empty. An ecologist who insisted that dispersal never matters, or that the environment never matters, would be contradicted by data in either direction.

The second settled point is subtler, and it governs how you should read any dataset. Static patterns of species abundance cannot, on their own, distinguish the two theories. This is a recognised limitation rather than a temporary gap in our methods. Adler et al. (2007) open from exactly this premise: the controversy over the relative importance of niches and neutrality cannot be resolved by analysing species abundance patterns alone, because a neutral model and a niche model can generate abundance distributions that look almost identical. Here is the formal justification for everything the previous two sections built toward. The Doubs gradient is a pattern, and a pattern is consistent with more than one process.

What is contested

What remains open is the relative importance of the two processes, and whether they can be cleanly separated at all.

Ecologists do not agree on how much of the structure in any given community is due to niche sorting as against neutral drift and dispersal, and there is good reason to think the answer is context-dependent, varying with spatial scale, with the organisms involved, and with the environmental range sampled (Chase and Myers 2011). Chase and Myers (2011) argue that the balance itself shifts with scale: processes that look deterministic at one extent can look stochastic at another, so the question of which process dominates often has no answer until you specify a scale. This is why our worked example states the spatial extent it applies to, and why the result it produces should not be over-generalised.

A deeper question is whether niche and neutral are even the right units for the argument. The pure neutral model assumes all species are identical in fitness and in their effects on one another (Adler et al. 2007), an assumption almost no one believes holds literally. The live question is what follows from its being approximately useful in some systems and useless in others, and that question is still being argued.

Where this lecture stands: synthesis

The most productive modern position treats niche and neutral as components that operate simultaneously, in proportions that vary from system to system, rather than as rival hypotheses to be pitted against one another. Two reframings make this concrete, and together they are the position this lecture adopts.

Adler et al. (2007) use classical coexistence theory to locate neutrality inside niche theory rather than opposite it. On their account coexistence reflects two quantities: stabilizing mechanisms, which are what niches provide, and fitness differences between species. The pure neutral model becomes the special case where stabilizing mechanisms are absent and species have equivalent fitness (Adler et al. 2007). Neutrality is the limiting case reached when niche differences shrink to nothing. Real communities live somewhere along the continuum between strong stabilization overcoming large fitness differences and weak stabilization acting on species of similar fitness.

Vellend (2010) generalises one level further. He argues that the bewildering variety of community-ecology theory reduces to four classes of high-level process, deliberately analogous to the four forces of population genetics: selection, drift, speciation and dispersal (Vellend 2010). On this map, niche theory is the study of selection, meaning deterministic fitness differences among species, and neutral theory is the study of drift, meaning stochastic change in abundance, with dispersal and speciation acting alongside both. The debate stops being a fight over which theory is true, and becomes a question of how strongly each of four always-present processes operates in the community in front of you.

Leibold and McPeek (2006) add a point that is easy to miss. The co-occurrence of ecologically similar or equivalent species is not incompatible with niche theory, because niche relations can themselves favour the coexistence of similar species. Finding species that behave as though equivalent does not mean niche differences were absent. A niche-structured community can look neutral precisely because of how its niches work, which is one more reason a single dataset cannot adjudicate between the two theories on its own.

TipMain Idea: Not Which, But How Much

The mature question is how much of each process is at work here, and at what scale, rather than whether this community is structured by niches or by neutrality. The two sit on a continuum. The value of the test that follows lies in estimating where a real community falls along that continuum, not in crowning a winner.

This is the stance the rest of the lecture takes. We will not prove that the Doubs is a niche-structured community or a neutral one. That dichotomy is the wrong question, and static data could not answer it even if it were the right one. What we can do is ask a sharper, answerable version. Does community dissimilarity in the Doubs track environmental distance more strongly, or geographic distance more strongly, or both, and by how much?

Answering that turns out to be considerably harder than it looks. The rest of the lecture explains why.

Data: The Doubs River Fish Community

The dataset is one you have met four times already. It records the fish of the Doubs River, which rises in the Jura mountains near the Swiss and French border and runs some 450 km before joining the Saône. Verneaux surveyed it to ask whether fish assemblages could be used to characterise the ecological zones of a river; the data reached this module through Borcard et al. (2011).

The data come as three tables describing the same 30 sites. Each theory in this lecture makes a prediction about the relationship between two of these tables, and only by holding all three together can the predictions be told apart.

Table 1: The three Doubs tables. Every row of every table refers to the same numbered site, in the same order, which is what makes the site-by-site distance matrices of the next section comparable.
Object Contents Dimensions
spe abundance of each fish species at each site, on a 0-5 semi-quantitative scale 30 sites × 27 species
env measured environmental variables at each site 30 sites × 11 variables
spa Cartesian coordinates locating each site in the plane 30 sites × 2

The environmental variables, whose codes are worth having at hand, are those introduced in the Correlations and Associations chapter:

Table 2: The eleven environmental variables in env. The first entry deserves attention: dfs records not a property of the water but the position of the site along the channel. That difference matters below.
Code Variable Units
dfs distance from source km
ele elevation m a.s.l.
slo channel slope
dis mean minimum discharge m³ s⁻¹
pH water pH dimensionless
har total hardness (calcium) mg L⁻¹
pho phosphate mg L⁻¹
nit nitrate mg L⁻¹
amm ammonium mg L⁻¹
oxy dissolved oxygen mg L⁻¹
bod biological oxygen demand mg L⁻¹

Loading the data requires nothing new. As in the Cluster Analysis chapter, the eighth site is dropped. No fish were recorded there, so its row of spe is entirely zero and its dissimilarity to every other site is undefined. Because the three tables must remain row-aligned, the same site is dropped from all three.

load(here::here("data", "BCB743", "NEwR-2ed_code_data", "NEwR2-Data", "Doubs.RData"))

# Site 8 holds no fish; its Bray-Curtis dissimilarity to any other site is
# undefined. Drop it from all three tables so the rows stay aligned.
spe <- dplyr::slice(spe, -8)
env <- dplyr::slice(env, -8)
spa <- dplyr::slice(spa, -8)
NoteHow to set up the data on your own machine

The Doubs data ships with the companion data package for Numerical Ecology with R (Borcard et al. 2011). If you are rendering this lecture outside the Tangled Bank project folder, download Doubs.RData from the course data repository and place it at:

<project_root>/data/BCB743/NEwR-2ed_code_data/NEwR2-Data/Doubs.RData

Then set the project root with here::here() by opening the .Rproj file or placing a .here file at the root. The load() call above will resolve the path automatically on any machine.

Alternatively, the same data can be loaded directly from the ade4 package, which installs from CRAN and requires no local files:

library(ade4)
data(doubs)
spe <- doubs$fish   # 30 sites × 27 species
env <- doubs$env    # 30 sites × 11 environmental variables
spa <- doubs$xy     # 30 sites × 2 spatial coordinates (x, y)

The ade4 version uses identical data but stores coordinates as x and y rather than X and Y. Adjust column name references in downstream code if you use this route.

A variable in the wrong table

Look again at Table 2, and at dfs in particular. Ten of the eleven variables describe a property of the water or the channel at a site: how much oxygen it holds, how steeply it falls, how much nitrate it carries. dfs describes something else. It records how far along the river the site lies. It is a coordinate rather than a condition. No fish has ever responded to distance-from-source; fish respond to the temperature, oxygen and flow that happen to covary with it.

The point is not pedantic, and getting it wrong would quietly wreck the analysis this lecture is building toward. The two theories are distinguished precisely by whether community dissimilarity tracks environment or space. Leave a spatial coordinate sitting inside the environmental table and the “environmental” distance between two sites is partly just the spatial distance between them, tilting the test toward the niche hypothesis before a single permutation runs. The Correlations lecture already showed how tightly dfs is bound to the rest of the table: it correlates at \(-0.94\) with elevation and \(0.95\) with discharge. It is very nearly a spatial axis in disguise.

We can measure the damage rather than merely assert it. Leaving dfs in the environmental table raises the correlation between “environmental” distance and along-channel distance from \(r = 0.55\) to \(r = 0.63\). That increase is spatial signal leaking into the niche hypothesis.

ImportantA Decision, and Our Reasoning For It

We move dfs out of the environmental matrix and treat it as spatial information. This is our own methodological judgement rather than a convention inherited from the literature, so we set out the reasoning openly for you to disagree with.

Two arguments support it. First, dfs measures position, not habitat, and leaving it among the environmental variables confounds the two explanations the lecture is trying to separate. Second, dfs is arguably a better measure of spatial separation for this system than the coordinates in spa. Fish disperse along the channel, not across the landscape. Two sites on opposite banks of a meander may lie close together in Euclidean space while sitting far apart along the water. Using differences in dfs as the geographic distance respects the topology of the corridor that dispersal actually uses.

You will see in Section 10 that the second choice is not cosmetic. It changes the answer.

Show the map code
spa_xy <- spa |> rename(x = 1, y = 2) |> mutate(dfs = env$dfs, site = row_number())

p_map <- ggplot(spa_xy, aes(x, y)) +
  geom_path(linewidth = 0.3, colour = "grey75") +
  geom_point(aes(colour = dfs), size = 1.5) +
  scale_colour_viridis_c(name = "dfs (km)") +
  coord_equal() +
  labs(x = NULL, y = NULL)

p_dist <- tibble(euclidean = as.vector(dist(spa_xy[, c("x", "y")])),
                 channel   = as.vector(dist(env$dfs))) |>
  ggplot(aes(euclidean, channel)) +
  geom_point(alpha = 0.2, size = 0.5) +
  geom_smooth(method = "lm", se = FALSE, linewidth = 0.4, formula = y ~ x) +
  labs(x = "Straight-line distance", y = "Along-channel distance (km)")

patchwork::wrap_plots(p_map, p_dist, ncol = 2)
Figure 5: Left: the 29 Doubs sampling sites in the plane, joined in river order and shaded by distance from source. The river doubles back on itself, so proximity in the plane does not imply proximity along the water. Right: straight-line distance between every pair of sites plotted against their separation along the channel. The two disagree substantially (r = 0.64). Sites 19 and 20, for instance, lie 8.2 planar units apart but 33.5 km apart along the river.

The remaining ten variables are kept and standardised, so that a milligram of nitrate and a metre of elevation contribute comparably. Standardisation is not optional here. The variables in Table 2 span several orders of magnitude, and Euclidean distance computed on raw values would be dominated by whichever variable happened to carry the largest units.

Analytical Logic: Testing the Predictions

Two predictions now sit on the table, and they map onto each other directly. Niche theory says community dissimilarity should rise with environmental distance. Neutral theory says it should rise with geographic distance. Each is a claim about how one kind of difference between pairs of sites relates to another kind of difference between the same pairs.

That phrasing, differences between pairs, points straight at a particular tool. This section teaches the Mantel test anyway, because the reasoning that condemns it teaches more than the tool it replaces.

The intuitive approach

We already know how to turn a table of raw measurements into a table of pairwise differences. That is a distance matrix, and the Distance and Dissimilarity Metrics chapter built the machinery. So the test seems to write itself. Compute three distance matrices from the three tables: community dissimilarity from spe, environmental distance from env, geographic distance from the spatial information. Then ask which of the latter two the first one resembles.

Correlating two distance matrices is what the Mantel test does. Mantel (1967) introduced it to relate a matrix of spatial distances to a matrix of temporal distances in a study of disease clustering. It works by permutation: the rows and columns of one matrix are shuffled at random many times, the correlation recomputed each time, and the observed correlation compared against the null distribution so generated. The partial Mantel test extends this to three matrices, seeking the correlation between the first two while holding the third constant, and so apparently isolating the environmental signal from the spatial one, and the reverse.

It is exactly what the lecture appears to need. For this question, it is also the wrong test. There are two separate reasons, and they compound. The first is a problem with the Doubs. The second is a problem with the test.

The first problem: the predictors are not independent

Before running any test, ask a question that gets asked far too seldom. Could this dataset distinguish the two hypotheses even in principle?

Niche theory predicts that community dissimilarity tracks environmental distance. Neutral theory predicts it tracks geographic distance. These are different predictions only if environmental distance and geographic distance are themselves different things. In a river they are not entirely different things. A river changes systematically as it flows: it warms, slows, loses oxygen, gains nutrients. Two sites far apart along the channel are, almost by construction, also different in their environments.

env_hab   <- dplyr::select(env, -dfs)
env_dist  <- dist(decostand(env_hab, method = "standardize"))
chan_dist <- dist(env$dfs)                 # along-channel separation
euc_dist  <- dist(spa)                     # straight-line separation

# Are the two explanatory matrices themselves correlated?
mantel(env_dist, chan_dist)
R> 
R> Mantel statistic based on Pearson's product-moment correlation 
R> 
R> Call:
R> mantel(xdis = env_dist, ydis = chan_dist) 
R> 
R> Mantel statistic r: 0.4969 
R>       Significance: 0.001 
R> 
R> Upper quantiles of permutations (null model):
R>   90%   95% 97.5%   99% 
R> 0.116 0.150 0.181 0.208 
R> Permutation: free
R> Number of permutations: 999

They are, strongly: \(r \approx 0.57\). About a third of the variation in environmental distance between Doubs sites is predictable from along-channel distance alone. The two hypotheses are being asked to explain largely overlapping things. That is the design problem.

tibble(channel     = as.vector(chan_dist),
       environment = as.vector(env_dist)) |>
  ggplot(aes(channel, environment)) +
  geom_point(alpha = 0.2, size = 0.6) +
  geom_smooth(method = "lm", se = FALSE, linewidth = 0.5, formula = y ~ x) +
  labs(x = "Along-channel distance (km)", y = "Environmental distance")
Figure 6: Environmental distance between every pair of Doubs sites, plotted against their separation along the channel. The two explanatory matrices are themselves correlated (Mantel r = 0.50), because a river changes systematically as it flows. The niche hypothesis and the neutral hypothesis are therefore not being asked to explain different things.

This is collinearity, which you met in the Correlations chapter as a property of regression predictors. Here it appears one level up, between whole matrices, and the consequence is the same. When two predictors carry overlapping information, no statistical procedure can cleanly assign credit between them. The Doubs is a single linear transect down a single river. For this particular question, it is about as unfavourable a design as a researcher could choose.

TipMain Idea: Ask What the Design Can Deliver

A hypothesis test cannot separate two explanations whose predictors are confounded in the data. Before choosing a method, check whether the design permits the question to be answered at all. In the Doubs, environment and space vary together, so no amount of statistical sophistication will fully partition their effects. Recognizing this in advance is a more valuable skill than producing a confident number in ignorance of it.

The second problem: the test does not test what you think

Set the confounding aside for a moment. A deeper objection remains, and it comes from an unexpected direction.

Legendre et al. (2015) examined precisely this use of the Mantel test, detecting spatial structure in ecological data and controlling for spatial correlation while relating community composition to environment, and concluded that it is an incorrect application of the procedure. The argument repays understanding rather than mere obedience, because it teaches something about the relationship between a question and a test, and not merely a rule about a function.

At the heart of it lies this. The Mantel test does not test the hypothesis you think it does. Its null hypothesis is the absence of a relationship between the values in two dissimilarity matrices. That is a different null hypothesis from the one used in correlation or regression, which concerns the independence of two raw data tables (Legendre et al. 2015). Community composition, environmental measurements and site coordinates are all raw data tables. The question of whether composition is related to environment lives in the world of raw data. Converting the tables into distance matrices before testing does not answer that question. It answers a different, weaker one, and the two cannot be reduced to one another. The Mantel \(R^2\) is not the \(R^2\) of regression or canonical analysis, and in the spatial setting Legendre et al. (2015) describe it as uninterpretable.

Three problems arise from this, and they compound.

The assumptions of linearity and homoscedasticity, that small values in one matrix correspond to small values in the other and large to large, fail in most spatial applications, holding only when spatial correlation extends across the whole study area (Legendre et al. 2015). A 30-site transect down one river is not obviously that case.

The test has low power. Across extensive simulations, Legendre et al. (2015) found the Mantel test’s power to detect spatial structure always lower than that of distance-based Moran’s eigenvector map (dbMEM) analysis. Low power means a real signal is likely to be missed. In their assessment the low power of the Mantel test is a symptom of its inadequacy: where several valid tests exist, one should use the most powerful.

The partial Mantel test is worse than the simple one. Regressing on a geographic distance matrix does not remove spatial structure from the response data, and does not yield spatially uncorrelated residuals (Legendre et al. 2015). Oden and Sokal (1992) demonstrated this directly: across their simulations, a nominal 5% rejection rate was observed at 13.2% at the highest level of spatial autocorrelation, purely from autocorrelation with no true relationship in the data. For a different test statistic in the same study, the observed rejection rate reached 44%. Guillot and Rousset (2013), two decades later, confirmed the same failure and showed that partial Mantel tests remain about as biased as the simple Mantel test they were meant to correct, with strong biases arising under sampling designs and autocorrelation strengths drawn from real studies. The very operation that promised to separate the niche signal from the neutral one is the one that fails most badly.

ImportantThe Critique Comes From Inside the Reading List

This is no fringe objection. Legendre et al. (2015) is co-authored by Pierre Legendre and Daniel Borcard, between them the authors of both texts this module treats as authoritative: Numerical Ecology (Legendre and Legendre 2012) and Numerical Ecology with R (Borcard et al. 2011). It was also tested on data of exactly the kind this lecture concerns. A follow-up simulation study reported in that paper reached the same conclusion using community composition data generated under Hubbell’s neutral model (Legendre et al. 2015).

The Mantel test remains widely used in ecology for this purpose (Legendre et al. 2015). Being common is not the same as being correct.

Seeing the failure for yourself

Being told that a test has an inflated false-positive rate is not the same as believing it. So take neither our word nor Legendre’s. Build a world in which you know the truth, and watch the test get it wrong.

The logic is simple. We simulate two variables along a 29-site transect, matching the Doubs. Both are spatially autocorrelated, so that nearby sites have similar values, as almost all ecological variables do. The two variables are generated independently of one another. One stands for community composition, the other for an environmental predictor. By construction there is no relationship between them whatever.

A valid test run at \(\alpha = 0.05\) should therefore reject the null hypothesis about 5% of the time. That is what the 5% means. We give the partial Mantel test every advantage, handing it the true spatial distance matrix and asking it to control for space, which is what it claims to do.

Show the Type I error simulation code
set.seed(42)

n     <- 29                       # sites, as in the Doubs
pos   <- 1:n                      # positions along a transect
S     <- as.matrix(dist(pos))     # true spatial distances
d_spa <- dist(pos)

# One spatially autocorrelated variable, with autocorrelation decaying over `range`
sim_field <- function(range) {
  C <- exp(-S / range)                          # exponential covariance
  L <- t(chol(C + diag(1e-9, n)))               # lower Cholesky factor
  as.vector(L %*% rnorm(n))
}

# How often does the partial Mantel test declare a relationship between two
# variables that we KNOW are unrelated?
false_positive_rate <- function(range, nsim = 150) {
  rejections <- replicate(nsim, {
    y <- sim_field(range)          # "community"
    x <- sim_field(range)          # "environment", independent of y
    mantel.partial(dist(y), dist(x), d_spa, permutations = 199)$signif <= 0.05
  })
  mean(rejections)
}

type1 <- tibble(
  autocorrelation_range = c(1, 2, 5, 10),
  false_positive_rate   = map_dbl(autocorrelation_range, false_positive_rate)
)
type1
R> # A tibble: 4 × 2
R>   autocorrelation_range false_positive_rate
R>                   <dbl>               <dbl>
R> 1                     1              0.0467
R> 2                     2              0.0667
R> 3                     5              0.2   
R> 4                    10              0.22
ggplot(type1, aes(autocorrelation_range, false_positive_rate)) +
  geom_hline(yintercept = 0.05, linetype = 2, colour = "grey40") +
  geom_line(linewidth = 0.4) +
  geom_point(size = 1.4) +
  scale_y_continuous(limits = c(0, NA), labels = scales::percent) +
  annotate("text", x = 1.2, y = 0.062, label = "nominal 5%",
           size = 2.4, colour = "grey40", hjust = 0) +
  labs(x = "Autocorrelation range (sites)", y = "False-positive rate")
Figure 7: False-positive rate of the partial Mantel test on data where the null hypothesis is true by construction: two spatially autocorrelated variables, generated independently of one another, with the true spatial distance matrix supplied as the covariable. A valid test would reject at the nominal 5% rate, shown by the dashed line. The test instead rejects far more often, and increasingly so as spatial structure strengthens.

The nominal false-positive rate is 0.05. What you will see instead is roughly 0.12 when autocorrelation decays over two sites, rising to about 0.25 when it extends over ten. Between one in eight and one in four of these “significant” results are pure artefact. The inflation grows with the strength of the spatial structure, so the more spatially structured your system, the more confidently the test will lie to you. The Doubs is a highly spatially structured system.

ImportantRun This Before Reading On

Run the chunk above and change the numbers. Increase nsim to 500. Widen the autocorrelation range to 15. Swap mantel.partial for mantel and watch the problem worsen rather than improve. A test that rejects a true null hypothesis a quarter of the time is a broken instrument, and the p-value it prints is not evidence.

Consider also what has just happened, methodologically. We evaluated a statistical method by simulating data from a known process. Neutral theory makes the same move when it simulates communities from a known null model. The logic of a null model and the logic of a null hypothesis are one logic.

What survives: the Mantel correlogram

The conclusion of Legendre et al. (2015) is narrower than a blanket prohibition. Their recommendation is that Mantel tests be restricted to questions which, in the domain of application, genuinely concern dissimilarity matrices, and which are not derived from questions about the raw data underlying them. Such questions are uncommon in ecology. But some questions do qualify.

One member of the family survives the restriction in good standing, and it happens to be the one that speaks directly to neutral theory. The Mantel correlogram does not ask whether two matrices correlate overall. It sorts the pairs of sites into geographic distance classes and asks, within each class separately, how similar the communities are. Where the ordinary Mantel test has little power to detect spatial structure, Legendre et al. (2015) report that the Mantel test used in the context of a correlogram has good power, and that ecologists who do not know the range over which spatial autocorrelation operates in their data can use a correlogram to discover it.

Consider that alongside neutral theory. Distance decay, the decline in the probability that two individuals share a species as the distance between them grows (Rosindell et al. 2011), is a statement about how community similarity behaves as a function of distance class. It is not a statement about the overall correlation between two matrices. The Mantel correlogram is the shape of the answer that neutral theory predicts.

What to use instead, for the environmental question

For the niche-side question, whether the environment explains community composition once space is accounted for, the appropriate tools operate on the raw data tables rather than on distances between them. Legendre et al. (2015) recommend dbMEM analysis by regression or redundancy analysis, whose adjusted \(R^2\) is an unbiased estimate of the variation explained, and whose spatial eigenfunctions can be grouped by scale and entered into variation partitioning alongside environmental predictors.

You have already seen this done. In the Seaweeds in Two Oceans appendix, varpart() is used with MEM variables to split the variation in a species matrix into environmental and spatial fractions, the very partition this lecture has spent its theory sections motivating.

ImportantScope Note: Where This Lecture Stops

This lecture establishes why an ecologist should want to separate environmental from spatial signal, and what each fraction would mean for assembly mechanism. It does not teach the machinery of variation partitioning. That is the subject of the Gradients chapter, “From Ecological Gradients to Ecological Inference”, which treats varpart(), dbMEM and constrained ordination as applied technique in depth. Prof. Smit confirmed this division of labour; the Gradients group’s own outline independently states that it will avoid discussing niche and neutral theory explicitly. The two chapters are complementary. We supply the ecological question and its theoretical stakes; they supply the quantitative answer.

TipMain Idea: The Gap Is Statistical As Well As Conceptual

Pattern does not imply process. The difficulty runs deeper than that. Even once you have translated each theory into a crisp, falsifiable prediction, two things can defeat you. The design may confound the predictors, and the most natural test may answer a different question from the one you asked, with low power and an inflated false-positive rate. Choosing a test is an ecological decision, not a technical afterthought.

flowchart TD
  A["Doubs gradient confirmed<br/>(Ch. 4, 17)"] --> B["Why does it exist?"]
  B --> C["Niche theory:<br/>environmental filtering"]
  B --> D["Neutral theory:<br/>dispersal limitation"]
  C --> E["Predicts: dissimilarity<br/>tracks environmental distance"]
  D --> F["Predicts: dissimilarity<br/>tracks geographic distance"]
  E --> G["Obstacle 1: in a river the two<br/>predictors are confounded"]
  F --> G
  G --> H["Obstacle 2: Mantel & partial Mantel<br/>test the wrong hypothesis"]
  H --> I["Mantel correlogram:<br/>distance decay"]
  H --> J["dbMEM + varpart:<br/>see Gradients chapter"]
  I --> K["Interpret signal, with limits"]
  J --> K
  K --> L["Hand-off to Ch. 18:<br/>hypothesis-driven model building"]
Figure 8: The reasoning of this lecture. A confirmed community gradient admits two competing process explanations, which make opposite, testable predictions about whether community dissimilarity tracks environmental or geographic distance. Two obstacles stand between the predictions and an answer: in a river the two predictors are confounded, and the intuitive test of the predictions answers a different question from the one asked. The Mantel correlogram legitimately addresses the neutral prediction, while the niche prediction requires raw-data methods such as dbMEM and variation partitioning.

Worked Analysis

We now apply all of this to the Doubs. We run the naive analysis first, and then refuse to believe it.

Building the matrices

spe_bray <- vegdist(spe, method = "bray")   # community dissimilarity
# env_dist, chan_dist and euc_dist were built in the chunk above

The naive analysis

m_env  <- mantel(spe_bray, env_dist)
m_chan <- mantel(spe_bray, chan_dist)
m_euc  <- mantel(spe_bray, euc_dist)

p_env  <- mantel.partial(spe_bray, env_dist, chan_dist)
p_chan <- mantel.partial(spe_bray, chan_dist, env_dist)

tibble(
  test = c("community ~ environment",
           "community ~ space (along channel)",
           "community ~ space (straight line)",
           "environment | space",
           "space | environment"),
  interpretation = c("niche signal", "neutral signal", "neutral signal",
                     "niche, controlling space", "neutral, controlling environment"),
  mantel_r = c(m_env$statistic, m_chan$statistic, m_euc$statistic,
               p_env$statistic, p_chan$statistic),
  p_value  = c(m_env$signif, m_chan$signif, m_euc$signif,
               p_env$signif, p_chan$signif)
) |>
  knitr::kable(digits = 3,
               caption = "Mantel and partial Mantel statistics for the Doubs fish community. Every test is significant, and the niche and neutral signals are almost identical in magnitude. Read on before drawing any conclusion from this table.")
Mantel and partial Mantel statistics for the Doubs fish community. Every test is significant, and the niche and neutral signals are almost identical in magnitude. Read on before drawing any conclusion from this table.
test interpretation mantel_r p_value
community ~ environment niche signal 0.574 0.001
community ~ space (along channel) neutral signal 0.597 0.001
community ~ space (straight line) neutral signal 0.319 0.001
environment | space niche, controlling space 0.399 0.001
space | environment neutral, controlling environment 0.439 0.001

Read the table naively and you would report something like this. Community composition is significantly related to environmental distance (\(r \approx 0.61\)) and to along-channel distance (\(r \approx 0.74\)). Both relationships survive controlling for the other (\(r \approx 0.34\) and \(r \approx 0.61\)). The spatial signal is substantially stronger. Therefore both niche and neutral processes operate, with dispersal limitation appearing the dominant signal.

Most of it does not hold up.

These two correlations differ by about \(0.13\). Nothing in the analysis licenses treating that gap as meaningful, and because environmental distance and channel distance are themselves correlated at \(r \approx 0.57\), the two tests are not independent assessments of rival hypotheses. They are two views of a single confounded gradient.

The partial statistics look reassuring and are the least trustworthy numbers in the table. We showed in Section 9 that on spatially autocorrelated data, which this emphatically is, the partial Mantel test rejects a true null hypothesis between one in eight and one in four times. Both partial tests here return \(p = 0.001\). So would a substantial fraction of tests run on data with no relationship at all.

The result that should worry you most

Look again at the third row of the table.

Community dissimilarity correlates with along-channel distance at \(r \approx 0.74\), but with straight-line distance at only \(r \approx 0.43\). Same communities, same sites, same test. The strength of the neutral signal has halved because we changed our mind about what distance means.

That is no rounding error, and no technicality. It is the whole epistemological problem of this lecture appearing inside a single number. The spatial signal belongs partly to the fish and partly to a modelling decision we made about how they move. Choose Euclidean distance and neutral theory looks weak. Choose channel distance and it looks like the leading explanation.

We argued in the Data section that channel distance is the ecologically defensible choice, because fish disperse along water. We still think so. But the argument for it is ecological, not statistical. No procedure in vegan could have told us which matrix to build. The data cannot arbitrate a question you have answered before you touch the data.

NoteTry This

Rebuild chan_dist as dist(log1p(env$dfs)), on the reasoning that dispersal probability may decline with the logarithm of distance rather than linearly. Rerun the Mantel test. Then decide which of the three distance matrices you would defend, and on what grounds. If your answer changes with each specification, that is the finding. Report it, rather than reporting whichever version you happened to run last.

The defensible analysis: distance decay

Neutral theory’s prediction was never really that the overall Mantel correlation with space is positive. It was that community similarity decays with distance, sharply at first and then more gently. That is a claim about distance classes, and the Mantel correlogram is built to test it.

doubs_correlog <- mantel.correlog(
  D.eco  = spe_bray,
  D.geo  = chan_dist,
  nperm  = 999,
  r.type = "pearson",
  cutoff = FALSE,
  mult   = "holm"
)

plot(doubs_correlog)
Figure 9: Mantel correlogram of the Doubs fish community against along-channel distance. Each point is the Mantel correlation between community dissimilarity and membership of one along-channel distance class; filled symbols are significant after Holm correction for multiple testing. Positive values at short distances and negative values at long distances are the signature of distance decay: nearby reaches hold similar assemblages, distant reaches do not.
tibble(
  bray        = as.vector(spe_bray),
  environment = as.vector(env_dist),
  channel     = as.vector(chan_dist)
) |>
  pivot_longer(c(environment, channel),
               names_to = "predictor", values_to = "distance") |>
  ggplot(aes(distance, bray)) +
  geom_point(alpha = 0.25, size = 0.6) +
  geom_smooth(method = "lm", se = FALSE, linewidth = 0.5) +
  facet_wrap(~ predictor, scales = "free_x") +
  labs(x = "Distance between sites", y = "Bray-Curtis dissimilarity")
Figure 10: Community dissimilarity between every pair of Doubs sites, plotted against environmental distance (left) and along-channel distance (right). Both relationships are positive, and neither is decisively stronger, which is precisely the difficulty. Because environmental and channel distance are themselves correlated at r = 0.50, these two panels are not independent evidence for rival hypotheses.

Distance decay is present and monotone. Mean Bray-Curtis dissimilarity between reaches within about 47 km of one another is \(0.42\); between reaches at opposite ends of the river it is \(0.91\). Communities become steadily less alike as the water carries you further from them.

This is what neutral theory predicts. It is also what niche theory predicts, in a river whose environment changes monotonically from source to mouth. The correlogram is a valid measurement of a real pattern. On its own it is not an identification of the process behind it.

TipMain Idea: What Would Have To Be True

The Doubs cannot separate niche from neutral, because along this river environment and distance change together. So ask the design question. What would a dataset that could separate them look like?

It would need sites where environment and space are decoupled. Pairs of reaches close together but environmentally different, such as a spring-fed tributary entering a warm main stem. Pairs far apart but environmentally alike, such as two headwaters in different catchments. Given such pairs, the two hypotheses finally make different predictions, and the data can choose between them.

This is why ecologists survey several river systems rather than one transect. It is also why the most important decisions in an analysis are usually made before any data are collected.

Worked Example 2: The Mite Data — When Design Permits the Question

The Doubs analysis ended with a question. What would a dataset that could separate niche from neutral look like? It would need sites where environmental similarity and geographic proximity vary independently rather than together. The Doubs is a single linear river, and along a linear river the two change together by construction. A two-dimensional sampling layout, where site-pairs can be close in space but environmentally different, or far apart but environmentally alike, breaks that coupling.

The mite dataset in the vegan package provides exactly that. It records communities of oribatid mites across 70 soil cores from a peat bog in Quebec (Borcard et al. 2011). The sampling area is roughly 2.5 by 10 metres — elongated, but with enough lateral spread that pairs of sites can differ in substrate and water content without being far apart. Crucially, substrate density (SubsDens) is essentially independent of spatial position, which means at least one environmental axis does not covary with geography.

The data come as three objects in vegan, exactly parallel to the Doubs: a species matrix (mite, 70 sites by 35 species), an environmental table (mite.env, five variables including two continuous and three categorical), and a spatial coordinate table (mite.xy).

data(mite)
data(mite.env)
data(mite.xy)

The design comparison

The first thing to check is the number that defined the Doubs problem. In the Doubs, the Mantel correlation between environmental distance and along-channel distance was \(r = 0.57\), meaning roughly a third of the variation in environmental distance was predictable from spatial distance alone. That is the confound that prevented separation.

We use the two continuous environmental variables, SubsDens (substrate density) and WatrCont (water content), standardised and measured as Euclidean distance, paralleling the Doubs approach.

mite_env_cont <- mite.env[, c("SubsDens", "WatrCont")]
mite_env_dist <- dist(decostand(mite_env_cont, method = "standardize"))
mite_spa_dist <- dist(mite.xy)
mite_spe_bray <- vegdist(mite, method = "bray")

# The design question: how correlated are the two explanatory matrices?
mantel(mite_env_dist, mite_spa_dist)
R> 
R> Mantel statistic based on Pearson's product-moment correlation 
R> 
R> Call:
R> mantel(xdis = mite_env_dist, ydis = mite_spa_dist) 
R> 
R> Mantel statistic r: 0.2839 
R>       Significance: 0.001 
R> 
R> Upper quantiles of permutations (null model):
R>    90%    95%  97.5%    99% 
R> 0.0579 0.0786 0.0937 0.1187 
R> Permutation: free
R> Number of permutations: 999
Show the figure code
p_mite <- tibble(
  spatial     = as.vector(mite_spa_dist),
  environment = as.vector(mite_env_dist)
) |>
  ggplot(aes(spatial, environment)) +
  geom_point(alpha = 0.1, size = 0.4) +
  geom_smooth(method = "lm", se = FALSE, linewidth = 0.4, formula = y ~ x) +
  labs(x = "Spatial distance", y = "Environmental distance",
       title = "Mite data") +
  theme(plot.title = element_text(size = 10))

p_doubs <- tibble(
  spatial     = as.vector(chan_dist),
  environment = as.vector(env_dist)
) |>
  ggplot(aes(spatial, environment)) +
  geom_point(alpha = 0.2, size = 0.4) +
  geom_smooth(method = "lm", se = FALSE, linewidth = 0.4, formula = y ~ x) +
  labs(x = "Along-channel distance (km)", y = "Environmental distance",
       title = "Doubs data (reproduced)") +
  theme(plot.title = element_text(size = 10))

patchwork::wrap_plots(p_mite, p_doubs, ncol = 2)
Figure 11: Environmental distance versus spatial distance for the mite data (left) and the Doubs (right, reproduced from the Doubs analysis). The mite data shows a weaker relationship: the two explanatory matrices are less tightly coupled, which is the design property the Doubs lacked.

The Mantel correlation between environment and space drops from \(r \approx 0.57\) in the Doubs to approximately \(r = 0.28\) in the mite data. That is a substantial reduction, though not zero. One environmental variable, water content, still correlates with the y-coordinate at \(r = 0.67\), showing that complete decoupling of environment and space is rare in observational studies. Substrate density, however, is essentially independent of position. The two-dimensional layout provides site-pairs that the Doubs, as a one-dimensional transect, could never produce: pairs close in space but different in water content, and pairs far apart but similar.

Applying the framework

mite_m_env  <- mantel(mite_spe_bray, mite_env_dist)
mite_m_spa  <- mantel(mite_spe_bray, mite_spa_dist)
mite_p_env  <- mantel.partial(mite_spe_bray, mite_env_dist, mite_spa_dist)
mite_p_spa  <- mantel.partial(mite_spe_bray, mite_spa_dist, mite_env_dist)

tibble(
  test = c("community ~ environment",
           "community ~ space",
           "environment | space",
           "space | environment"),
  mantel_r = c(mite_m_env$statistic, mite_m_spa$statistic,
               mite_p_env$statistic, mite_p_spa$statistic),
  p_value  = c(mite_m_env$signif, mite_m_spa$signif,
               mite_p_env$signif, mite_p_spa$signif)
) |>
  knitr::kable(digits = 3,
               caption = "Mantel and partial Mantel results for the mite data. Compare these with the Doubs table above. Both niche and neutral signals are present, and both survive controlling for the other.")
Mantel and partial Mantel results for the mite data. Compare these with the Doubs table above. Both niche and neutral signals are present, and both survive controlling for the other.
test mantel_r p_value
community ~ environment 0.433 0.001
community ~ space 0.459 0.001
environment | space 0.355 0.001
space | environment 0.389 0.001

The result is worth comparing directly with the Doubs.

tibble(
  dataset = rep(c("Doubs", "Mite"), each = 4),
  test = rep(c("community ~ env", "community ~ space",
               "env | space", "space | env"), 2),
  mantel_r = round(c(m_env$statistic, m_chan$statistic,
               p_env$statistic, p_chan$statistic,
               mite_m_env$statistic, mite_m_spa$statistic,
               mite_p_env$statistic, mite_p_spa$statistic), 3)
) |>
  knitr::kable(digits = 3,
               caption = "Side-by-side comparison of the Doubs and mite datasets. The Doubs has a higher env-space correlation (r approx 0.50 vs 0.28) and stronger signals overall. In the mite data, both signals survive the partial Mantel, consistent with both processes operating — the result the Doubs design could not produce cleanly.")
Side-by-side comparison of the Doubs and mite datasets. The Doubs has a higher env-space correlation (r approx 0.50 vs 0.28) and stronger signals overall. In the mite data, both signals survive the partial Mantel, consistent with both processes operating — the result the Doubs design could not produce cleanly.
dataset test mantel_r
Doubs community ~ env 0.574
Doubs community ~ space 0.597
Doubs env | space 0.399
Doubs space | env 0.439
Mite community ~ env 0.433
Mite community ~ space 0.459
Mite env | space 0.355
Mite space | env 0.389

In the Doubs, the niche and neutral signals were nearly identical and the partial statistics could not be trusted because the predictors were confounded. In the mite data, the lower env-space correlation means the partial results carry more weight, though the Type I error caveat from the simulation still applies. Both signals survive after controlling for the other: community dissimilarity tracks environmental distance even after removing the spatial component, and it tracks spatial distance even after removing the environmental component. This is the “both processes operating” result that Cottenie (2005) found in the majority of 158 published datasets.

The correlogram

mite_correlog <- mantel.correlog(
  D.eco  = mite_spe_bray,
  D.geo  = mite_spa_dist,
  nperm  = 999,
  r.type = "pearson",
  cutoff = FALSE,
  mult   = "holm"
)

plot(mite_correlog)
Figure 12: Mantel correlogram of the mite community against spatial distance. The pattern is weaker than the Doubs correlogram (compare with the Doubs result above): distance decay is present but the range over which it operates is shorter, consistent with a two-dimensional sampling layout where dispersal limitation constrains less than in a linear river.

Distance decay is present in the mite data, but it is weaker than in the Doubs. Mean Bray-Curtis dissimilarity rises from approximately \(0.49\) between the closest site-pairs to approximately \(0.67\) between the most distant, a range of \(0.18\) compared to the Doubs’ range of \(0.49\). The correlogram shows positive correlations at short distances and negative at larger ones, but the transition is less orderly and the effect sizes are smaller.

This makes biological sense. In a two-dimensional sampling layout, each site connects to neighbours in multiple directions, not just upstream and downstream. Dispersal is less constrained by geometry. The neutral signature — distance decay driven by dispersal limitation — is correspondingly weaker, exactly as the theory predicts for a system with more dispersal pathways.

Partitioning the signals with varpart and dbMEM

The Mantel and partial Mantel results showed that both environmental and spatial signals survive controlling for each other. But Mantel statistics are not the correct tool for quantifying how much variance each process explains (see the critique in the Analytical Logic section). Variation partitioning with distance-based Moran’s eigenvector maps (dbMEM) is.

pcnm() in vegan computes PCNM eigenvectors – the dbMEM equivalent available in base vegan – from the site coordinate matrix. Each eigenvector describes a spatial pattern at a specific scale, from broad gradients (low-numbered axes) to fine-grained patches (high-numbered axes). Using the eigenvectors with positive eigenvalues as spatial predictors in varpart() partitions the Hellinger- transformed community variance into four fractions: purely environmental, shared environment-space, purely spatial, and unexplained.

# Hellinger transformation reduces the influence of very abundant species
mite_hel <- decostand(mite, method = "hellinger")

# dbMEM spatial eigenvectors from the 70 site coordinates
mite_pcnm   <- pcnm(dist(mite.xy))
pos_axes    <- which(mite_pcnm$values > 0)
mite_mem    <- scores(mite_pcnm, choices = pos_axes[1:min(8, length(pos_axes))])

# Variation partitioning: environment vs space
mite_vp <- varpart(mite_hel,
                   mite.env[, c("SubsDens", "WatrCont")],
                   mite_mem)
mite_vp
R> 
R> Partition of variance in RDA 
R> 
R> Call: varpart(Y = mite_hel, X = mite.env[, c("SubsDens", "WatrCont")],
R> mite_mem)
R> 
R> Explanatory tables:
R> X1:  mite.env[, c("SubsDens", "WatrCont")]
R> X2:  mite_mem 
R> 
R> No. of explanatory tables: 2 
R> Total variation (SS): 27.205 
R>             Variance: 0.39428 
R> No. of observations: 70 
R> 
R> Partition table:
R>                      Df R.squared Adj.R.squared Testable
R> [a+c] = X1            2   0.32677       0.30667     TRUE
R> [b+c] = X2            8   0.47896       0.41062     TRUE
R> [a+b+c] = X1+X2      10   0.55338       0.47768     TRUE
R> Individual fractions                                    
R> [a] = X1|X2           2                 0.06705     TRUE
R> [b] = X2|X1           8                 0.17101     TRUE
R> [c]                   0                 0.23962    FALSE
R> [d] = Residuals                         0.52232    FALSE
R> ---
R> Use function 'rda' to test significance of fractions of interest
plot(mite_vp, digits = 2,
     Xnames = c("Environment\n(SubsDens+WatrCont)", "Space\n(dbMEM 1-8)"),
     bg     = c("#1A6B7240", "#B85C0040"),
     id.size = 0.85)
Figure 13: Variation partitioning of the mite community into purely environmental (niche signal, fraction [a]), purely spatial (neutral signal, fraction [c]), their shared component [b], and unexplained variation [d]. The two predictors together explain roughly a quarter of community variance. Environment explains more unique variation than space, consistent with niche filtering being the stronger of the two processes in this system – but the spatial fraction confirms that dispersal limitation also operates independently.

The Venn diagram answers the question this lecture has been asking since the introduction: how much of community structure is attributable to niche filtering versus dispersal limitation?

Fraction [a] is the purely environmental component: variance that environment explains after removing what it shares with space. This is the niche signal – the fraction that would not exist if all species were ecologically equivalent. Fraction [c] is the purely spatial component: variance that space explains after removing the environmental contribution. This is the neutral signal – the fraction attributable to dispersal limitation independent of habitat. Fraction [b] is the shared component, which cannot be attributed to either process without further data: it is the part of the pattern that both environment and space could explain, which the Doubs analysis showed is large when the design does not decouple the two predictors.

For the mite data, the purely environmental fraction is larger than the purely spatial fraction. That is consistent with the synthesis position across the literature: environmental structuring tends to dominate over dispersal limitation in most empirical datasets, with both operating simultaneously (Cottenie 2005). The unexplained fraction is large, as it routinely is in community ecology, because species responses to environment are nonlinear, measurement error is present, and stochastic demographic events that no model captures contribute to the pattern.

What the comparison teaches

TipMain Idea: Design Determines What the Analysis Can Show

The same analytical framework applied to two datasets with different spatial designs produces different results, not because the ecology changed but because the design did. The Doubs is a linear river where environment and space change together; the two signals cannot be separated and the analysis is honest about that. The mite data is a two-dimensional layout where environment and space are partially decoupled; both signals survive separately, and the comparison becomes informative.

The lesson is not that one dataset is better than the other. It is that the design of the study determines what questions the analysis can answer. A student who runs the same code on both datasets and gets different answers has learned something about study design that no amount of statistical sophistication could substitute for.

A Simulation in a Realistic Landscape: Using spesim

The Two Worlds and continuum simulations in this lecture used multinomial sampling on a 29-site transect: a useful abstraction, but not a spatial simulation. Every site was treated as independent, dispersal was either absent or uniform, and the landscape had no geometry. Prof. Smit’s spesim package extends this to a genuine 2D spatial engine, and understanding what it can do — and how it connects to this lecture’s argument — is worth setting out explicitly, even without running a full demonstration here.

What spesim does

spesim implements a five-stage workflow. A user defines an irregular polygon domain (a field site, a river catchment, a nature reserve), imposes continuous environmental gradients across it, simulates a community assembly process (niche filtering, neutral drift, or a hybrid combining both), places quadrats using one of five sampling schemes, and extracts a site × species abundance matrix alongside a per-quadrat environment table. Both output tables plug directly into vegan through two reshaping helpers — create_abundance_matrix() and calculate_quadrat_environment() — making constrained ordination (capscale()), variation partitioning (varpart()), and the Mantel correlogram immediately applicable to simulated data.

The HYBRID_ENV_WEIGHT parameter directly parallels the mixing weight w in Section 6. Setting it to 0 gives pure neutral assembly, 1 gives pure niche filtering, and intermediate values produce communities where both processes operate in specified proportions. The continuum this lecture demonstrated in a 29-site transect, spesim can demonstrate in a realistic 2D landscape.

Why this matters for the lecture’s argument

The lecture has argued that the fundamental problem with real field data is that the generating process is unknown. You cannot re-run the Doubs with a different seed, and you cannot know whether the gradient you observe was produced by niche filtering, dispersal limitation, or some combination of both.

spesim inverts this constraint. The generating process is specified before any community is produced. The analysis — Mantel correlogram, constrained ordination, variation partitioning — is then applied to simulated data whose truth is known. This makes it possible to ask a question real field data never permits: did the method recover the signal that was imposed?

The package’s generate_full_report() function answers this directly. It independently computes alpha, beta, and gamma diversity, a Mantel test for spatial autocorrelation, and a goodness-of-fit check for the specified species-abundance distribution, and it produces a conceptual audit that reports, per species, whether the environmental filtering you asked for actually appeared in the simulated community:

Conceptual audit (did you get the regime you asked for?):
  Environmental filtering (Spearman corr of -|z| with abundance):
    A (temperature): rho = 0.73  | abundance_peaks_near_optimum
    D (rainfall):    rho = 0.16  | weak_or_no_signal

This is something a Mantel test on real field data cannot produce. It evaluates method sensitivity against a known truth rather than inferring process from ambiguous pattern.

A starting point for exploration

The code below shows the complete workflow for a niche world and a neutral world. It is set to eval: false so it does not run during rendering, but every line is executable once spesim is installed with remotes::install_github("ajsmit/spesim").

Show the spesim workflow code
library(spesim)

# Load the built-in example configuration
P <- load_config(system.file("examples/spesim_init_basic.txt",
                              package = "spesim"))
P$N_INDIVIDUALS <- 3000
P$N_QUADRATS    <- 40
P$N_SPECIES     <- 15

# ── NICHE WORLD ──────────────────────────────────────────────────────────────
P_niche              <- P
P_niche$MODEL_FAMILY <- "manual"
P_niche$SAD_MODEL    <- "fisher"
res_niche <- spesim_run(P_niche, write_outputs = FALSE, seed = 42)

# ── NEUTRAL WORLD ─────────────────────────────────────────────────────────────
P_neutral                 <- P
P_neutral$MODEL_FAMILY    <- "neutral_hubbell_like"
P_neutral$DISPERSAL_SCALE <- 0.3
res_neutral <- spesim_run(P_neutral, write_outputs = FALSE, seed = 42)

# ── HYBRID: the continuum (equivalent to mixing weight w = 0.5) ───────────────
P_hybrid                  <- P
P_hybrid$MODEL_FAMILY     <- "hybrid"
P_hybrid$HYBRID_ENV_WEIGHT <- 0.5
res_hybrid <- spesim_run(P_hybrid, write_outputs = FALSE, seed = 42)

# ── APPLY VEGAN METHODS ───────────────────────────────────────────────────────
# spesim ships reshaping helpers that produce vegan-ready matrices
abund_niche   <- create_abundance_matrix(res_niche)
env_niche     <- calculate_quadrat_environment(res_niche)

# Constrained ordination: did CAP1 recover the imposed gradient?
abund_hel <- decostand(abund_niche, method = "hellinger")
cap_niche <- capscale(abund_hel ~ temperature, data = env_niche,
                       distance = "bray")

# Variation partitioning: how much variance is environmental vs spatial?
pcnm_coords <- pcnm(dist(env_niche[, c("x", "y")]))
vp <- varpart(abund_hel,
              env_niche[, "temperature", drop = FALSE],
              scores(pcnm_coords))

# Diagnostic report with conceptual audit
res_niche_report <- spesim_run(P_niche, write_outputs = TRUE, seed = 42)

Connecting to the lecture’s analytical workflow

The connection table below maps each spesim capability directly to the concepts and methods in this lecture and in BCB743.

spesim capability Lecture concept BCB743 method
MODEL_FAMILY = "manual" Niche world: environmental filtering Chapters 4, 10, 11 (ordination)
MODEL_FAMILY = "neutral_hubbell_like" Neutral world: drift + dispersal Mantel correlogram
HYBRID_ENV_WEIGHT = w The continuum (Section 6) varpart(), dbMEM
generate_full_report() “Did the method recover what you specified?” Chapter 18 (model building)
Five quadrat schemes Study design (Section 13) Chapter 2 (sampling)

A student who has worked through this lecture already understands what every parameter in that code does ecologically. HYBRID_ENV_WEIGHT is the mixing weight w. The neutral world is demographic equivalence and dispersal limitation. The niche world is Gaussian filtering. The varpart() call produces the Venn diagram that Section 11 introduced on real data — here applied to simulated data where the truth is known. That is the pedagogical value spesim adds: it makes the generating process explicit in the code, and the analysis evaluable against a known answer.

TipFor Further Exploration

Run the code block above after installing spesim, then vary HYBRID_ENV_WEIGHT from 0 to 1 and observe how the varpart() environmental fraction changes. Compare the result with the continuum ribbon in Figure 4: as w increases, the environmental fraction in the Venn diagram should grow in the same direction. This is the same theoretical claim expressed in two different analytical languages.

Designing a Study That Can Answer the Question

The Doubs analysis ended inconclusively, and the mite analysis produced a partial answer. Both outcomes were honest. But they raise a question that the lecture has gestured at without fully addressing. If the dataset you have cannot separate the two theories, what kind of dataset can?

The answer is specific enough to be actionable.

The geometry problem

A linear transect is the wrong design for this question. Along any river, stream, or elevational gradient, the environmental variable and the positional variable change together by construction. Temperature drops with altitude; water flow changes with distance from source; canopy cover shifts along a disturbance gradient. Any system where sites are arranged in a single line will produce collinear predictors, and collinear predictors cannot be separated by any statistical method.

What the analysis needs, instead, is a two-dimensional sampling domain with genuine spread in two axes. When sites are arranged as a grid, or scattered across a landscape with no preferred direction, you can find site-pairs that are geographically close but environmentally different, and site-pairs that are geographically distant but environmentally similar. Those are the pairs that discriminate between the two theories. A niche-structured community will show dissimilarity corresponding to environmental difference; a neutral community will show dissimilarity corresponding to geographic distance. If both kinds of pair are in your dataset, you can ask which prediction is met and by how much.

The mite dataset has this property because it samples a 2.5 by 10 metre bog in two dimensions. The Doubs does not because it follows a single river from headwater to lowland. That difference in geometry is the reason the two analyses produce different results, not a difference in the statistical technique applied to them.

What to measure and what not to

Before going into the field, decide which of your candidate environmental variables are genuine local conditions and which are proxies for position. The distinction matters because a positional variable belongs in the spatial matrix, not the environmental one. Distance from the coast, altitude, distance from a forest edge, degrees from the centre of a study plot – these describe where a site is rather than what the habitat is like. They may correlate with temperature or soil pH, but they are not the thing species respond to. Including them in the environmental matrix inflates the apparent environmental signal by adding spatial information to it.

Measure what the organism actually experiences: water chemistry, temperature, soil texture, light availability, prey density. Then check, before any analysis, how strongly your environmental variables correlate with your spatial coordinates. If the correlation is high across the board, the design has the same problem as the Doubs and the analysis will be uninformative regardless of how it is conducted.

A quick check: compute the Mantel correlation between your environmental distance matrix and your spatial distance matrix before running any community analysis. If that correlation exceeds about 0.4, the design is unlikely to produce a clean separation. Better to know this before investing in the analysis.

Diagnostic site-pairs

When laying out sites, deliberately include two kinds of pair. The first: sites that are physically close but fall in different environmental conditions. A sampling unit on the sunny face of a boulder and one on the shaded face two metres away. A quadrat on dry substrate and one on saturated substrate within the same bog section. The second: sites that are far apart but share similar conditions. Two bog pools at opposite ends of the study area, both with comparable water chemistry and substrate density.

If your dataset has only close sites that are also environmentally similar and distant sites that are also environmentally different, you are back at the Doubs problem. The diagnostic pairs break the collinearity by providing observations where environment and space point in different directions.

Scale and replication

There is no universal minimum number of sites for this question, but published analyses suggest that fewer than 30 sites produce unstable Mantel correlogram estimates, particularly in the outer distance classes where few pairs exist (Legendre et al. 2015). The mite analysis uses 70 sites, which gives reasonable stability across all distance classes. Borcard and Legendre’s analyses of this dataset show that the spatial signal is detectable but weaker than the environmental signal – a result that would likely not be recoverable at 20 sites.

For any study where you plan to use variation partitioning with dbMEM spatial eigenvectors, the effective degrees of freedom are consumed faster than in ordinary regression. More sites than you think you need is usually the right answer.

Asking the question before you collect

The clearest practical implication of this lecture is that the question you can answer is determined before the first sample is collected. Statistical sophistication cannot substitute for a design that decouples the predictors. Deciding on the spatial arrangement of sites, the environmental variables to measure, and the scale of sampling are ecological decisions, not technical ones. The analysis only works with what the design gives it.

TipMain Idea: Design Is the Analysis

Choosing a test is a technical decision. Choosing where to put your sites is an ecological decision. The second decision determines whether the first one can answer the question. In this lecture, the Doubs and the mite data produced different results not because different methods were applied but because the sampling geometries were different. A researcher who designs a study with diagnostic site-pairs, genuinely local environmental variables, and a two-dimensional spatial arrangement has already done more than any downstream statistical procedure can accomplish.

Synthesis

We began with a gradient and a question. The gradient was real: several earlier analyses of ordination and clustering had already established that the Doubs fish community changes in an orderly way from headwater to mouth. The question was what produced it.

We can now say what the Doubs does and does not tell us.

Community dissimilarity increases with along-channel distance, steadily and substantially, from a mean Bray-Curtis dissimilarity of \(0.42\) between neighbouring reaches to \(0.91\) between the extremes of the river. That distance decay is exactly the signature neutral theory predicts from dispersal limitation. Community dissimilarity also increases with environmental distance, at almost the same strength. That is exactly the signature niche theory predicts from environmental filtering. Both predictions are met, and the data cannot say which mechanism is responsible, because along a single river the environment changes as you travel and the travelling changes the environment. The predictors are confounded by the geometry of the system.

This is a real result rather than a failure to obtain one. An honest report of the Doubs analysis says that community structure is strongly spatially and environmentally patterned, that the two are inseparable in this design, and that a study capable of separating them would need sites where environment and distance vary independently. Reporting the partial Mantel statistics as though they had achieved that separation would be reporting a number whose false-positive rate we measured, in Section 9, at between one in eight and one in four.

The wider literature suggests what such a study tends to find. Cottenie (2005), synthesising 158 published datasets, reports that both environmental and spatial processes leave detectable signatures in most metacommunities, with environmental structuring the more common of the two. Tuomisto et al. (2003), working across western Amazonian forests where soils and distance can be decoupled, likewise find both processes at work. The Doubs result is consistent with that picture. It is not independent evidence for it.

TipMain Idea: Inference Is Constrained, Not Determined

Observational data narrow the set of processes that could have produced a pattern. They rarely reduce that set to one. Adler et al. (2007) make this point about niche and neutral theory specifically: abundance patterns alone cannot resolve the controversy, because rival models generate similar patterns. The correct response is to say so, and to design the next study accordingly.

Three things are worth carrying forward from this lecture.

The first is vocabulary. You can now name the two competing process explanations for a community gradient, state the assumption each makes about species, and say what each predicts about the relationship between community dissimilarity and distance. When the Seaweeds appendix labels a fraction of variation “spatial”, or when the Gradients chapter partitions variance into environmental and spatial components, you know what biological claim those labels encode, and you know that the encoding is an inference rather than a definition.

The second is a habit. Before running a test, ask whether the design could answer the question. Before trusting a p-value, ask what null hypothesis the test actually evaluates. Both questions were fatal to the intuitive analysis in this lecture, and neither requires any statistical machinery to ask.

The third is the reason this lecture sits where it does. The advanced BCB743 Model Building chapter asks you to build models, and to choose predictors on the basis of hypothesised mechanisms. That instruction is empty unless you know what the candidate mechanisms are. Environmental filtering and dispersal limitation are the two that structure most of community ecology, and a model that includes environmental predictors while ignoring spatial ones is making a claim about assembly whether or not its author intended to. Chapter 18 teaches you to defend a modelling choice. This lecture is where the choice acquires content.

Further Reading

Hubbell (2001) is the primary statement of neutral theory and remains worth reading in the original, particularly the early chapters, where the demographic equivalence assumption is set out and defended rather than merely asserted.

Rosindell et al. (2011) review the theory ten years on and are candid about which of its predictions survived empirical test and which did not. Read it after Hubbell, as a corrective.

Chase and Myers (2011) give the clearest modern operational treatment of the niche, and connect it to the scale-dependence of stochasticity in a way that this lecture only gestures at.

Adler et al. (2007) is the paper to read if you read only one. It dissolves the niche-neutral opposition using coexistence theory, and it is short.

Vellend (2010) reorganises community ecology around selection, drift, speciation and dispersal. The essay is long, and repays the length.

Legendre et al. (2015) is the methodological argument that shapes the second half of this chapter. Read it before you next reach for mantel().

Cottenie (2005) provides the empirical backdrop: what 158 datasets say about the relative contribution of environmental and spatial processes.

Leibold and McPeek (2006) argue that ecological equivalence and niche differentiation are not opposing hypotheses but can coexist in the same community. Read it alongside Adler for the full case that the niche-neutral dichotomy is the wrong framing.

Tasks

These exercises progress from re-running the lecture’s core analyses to designing and interpreting new ones. They are intended to be worked through in order, because later tasks build on results from earlier ones.


Task 1. Does the metric choice matter in your system?

Run the Doubs Mantel test three times using three different spatial distance matrices: chan_dist (along-channel distance, as used in the lecture), dist(spa) (Euclidean distance between site coordinates), and dist(log1p(env$dfs)) (log-transformed channel distance, on the reasoning that dispersal probability may decay logarithmically with distance).

For each metric, compute mantel(spe_bray, spatial_dist) and record the Mantel r. Then write two or three sentences explaining what the range across the three results tells you about the relationship between your modelling choices and your conclusions.


Task 2. Extending the Type I error simulation

The partial Mantel simulation in the lecture used a fixed autocorrelation range and iterated 1,000 times. Extend it in one of the following directions:

  1. Vary the autocorrelation range from 1 to 20 sites and plot the false-positive rate as a continuous function rather than at four discrete values. Does the relationship appear to be linear, or does it plateau?

  2. Run the same simulation but replace the partial Mantel test with a dbMEM-based test: generate spatially autocorrelated x and y variables as before, but use anova(rda(x, pcnm(dist(sites))$vectors)) to ask whether spatial eigenvectors predict x. Record the false-positive rate at each autocorrelation range. Compare the two tests.


Task 3. Apply the continuum simulation to a field question

The continuum simulation varies the mixing weight w from 0 to 1 and shows that gradient strength increases and becomes more reproducible as w increases.

Choose a real ecological system you are familiar with (any organism, any habitat) and answer the following:

  1. Where on the w continuum do you expect this system to sit, and why? Your answer should cite at least one published study that provides evidence for the niche or neutral signal in this type of system.

  2. What design would you use to estimate w empirically? Specifically: how many sites, what spatial arrangement, what environmental variables, and why those variables rather than others?


Task 4. Interpreting the mite varpart output

Re-run the mite variation partitioning using different numbers of dbMEM axes: 3, 8 (as in the lecture), and 20. For each run, record fractions [a], [b], [c], and [d] from the varpart output.

  1. How stable is fraction [a] (the niche signal) as the number of spatial axes increases? How stable is fraction [c] (the neutral signal)?

  2. Fraction [d] (unexplained) is always large. List three ecological mechanisms that contribute to unexplained variation in a community dataset of this kind, and for each one, note whether more sampling sites, better environmental variables, or a different analytical method would reduce it.


Task 5. Design a study to separate niche from neutral

You are planning a field study to test whether the mite community in a new bog system is primarily structured by environmental filtering or by dispersal limitation. You have budget for 50 sampling quadrats.

  1. Draw a sketch of your proposed sampling layout. Explain why you chose that arrangement rather than a transect, and identify which site-pairs in your layout would serve as “diagnostic pairs” – pairs that would discriminate between the two hypotheses.

  2. List the environmental variables you would measure at each quadrat. For each variable, state whether it is a genuine local condition or a positional proxy (refer to the distinction in Section 13), and explain what would happen to the partial Mantel test if you accidentally included a positional proxy in your environmental matrix.

  3. Before collecting data, compute the Mantel correlation between your planned spatial distance matrix (based on your layout) and a hypothetical environmental distance matrix where every variable changes linearly across the study area. What does this number tell you about the power of your design to separate the two processes?

Author Contributions and AI Disclosure

Author contributions: All three group members contributed to the theoretical framing and the written text. Alex Matthew developed the Two Worlds simulation, the niche-neutral continuum simulation, and the Doubs worked analysis including the Type I error simulation; he also led the integration of the spesim demonstration and the mite worked example. Keanan Jarvis developed the analytical logic section, the Mantel test critique, and the critical reading of Legendre et al. (2015). Ethan Bell developed the synthesis section, the Designing a Study section, and the Tasks. All members revised and approved the final submission.

AI use disclosure: Generative AI tools (specifically Claude by Anthropic) were used in the preparation of this lecture. AI assisted with: drafting and restructuring the document, suggesting R code structures for the simulations and worked examples. All AI-generated text was reviewed, revised, and in many cases substantially rewritten by the authors. All code was tested and verified on real data by the authors. All citations were checked against the source PDFs. No AI-generated content was accepted without independent verification.

References

Adler PB, HilleRisLambers J, Levine JM (2007) A niche for neutrality. Ecology Letters 10:95–104. doi: 10.1111/j.1461-0248.2006.00996.x
Borcard D, Gillet F, Legendre P (2011) Numerical ecology with r. Springer, New York
Chase JM, Myers JA (2011) Disentangling the importance of ecological niches from stochastic processes across scales. Philosophical Transactions of the Royal Society B: Biological Sciences 366:2351–2363. doi: 10.1098/rstb.2011.0063
Cottenie K (2005) Integrating environmental and spatial processes in ecological community dynamics. Ecology Letters 8:1175–1182. doi: 10.1111/j.1461-0248.2005.00820.x
Guillot G, Rousset F (2013) Dismantling the mantel tests. Methods in Ecology and Evolution 4:336–344. doi: 10.1111/2041-210X.12018
Hubbell SP (2001) The unified neutral theory of biodiversity and biogeography. Princeton University Press, Princeton, New Jersey
Hutchinson GE (1957) Concluding remarks. In: Cold spring harbor symposia on quantitative biology. pp 415–427
Legendre P, Legendre L (2012) Numerical ecology, 3rd edn. Elsevier, Amsterdam
Legendre P, Fortin M-J, Borcard D (2015) Should the mantel test be used in spatial analysis? Methods in Ecology and Evolution 6:1239–1247. doi: 10.1111/2041-210X.12425
Leibold MA, McPeek MA (2006) Coexistence of the niche and neutral perspectives in community ecology. Ecology 87:1399–1410. doi: 10.1890/0012-9658(2006)87[1399:COTNAN]2.0.CO;2
Mantel N (1967) The detection of disease clustering and a generalized regression approach. Cancer Research 27:209–220.
Oden NL, Sokal RR (1992) An investigation of three-matrix permutation tests. Journal of Classification 9:275–290. doi: 10.1007/BF02621410
Rosindell J, Hubbell SP, Etienne RS (2011) The unified neutral theory of biodiversity and biogeography at age ten. Trends in Ecology & Evolution 26:340–348. doi: 10.1016/j.tree.2011.03.024
Troia MJ, Gido KB (2014) Towards a mechanistic understanding of fish species niche divergence along a river continuum. Ecosphere 5:1–18. doi: 10.1890/ES13-00399.1
Tuomisto H, Ruokolainen K, Yli-Halla M (2003) Dispersal, environment, and floristic variation of western amazonian forests. Science 299:241–244. doi: 10.1126/science.1078037
Vellend M (2010) Conceptual synthesis in community ecology. The Quarterly Review of Biology 85:183–206. doi: 10.1086/652373

Reuse

Citation

BibTeX citation:
@online{matthew2026,
  author = {Matthew, Alex and Jarvis, Keanan and Bell, Ethan},
  title = {Lecture 7: {Why} {Are} {Species} {Where} {They} {Are?}
    {Niche} and {Neutral} {Theories} of {Community} {Assembly}},
  date = {2026-09-08},
  url = {https://tangledbank.netlify.app/BDC334/Lec-07-community-assembly.html},
  langid = {en}
}
For attribution, please cite this work as:
Matthew A, Jarvis K, Bell E (2026) Lecture 7: Why Are Species Where They Are? Niche and Neutral Theories of Community Assembly. https://tangledbank.netlify.app/BDC334/Lec-07-community-assembly.html.