Back to feed

The Curve After ReLU: Why GELU, SiLU and SwiGLU Promise Better Generalization

This StatQuest episode walks through how three modern activation functions differ from ReLU and how they can all be derived from the idea of input-dependent dropout; I turned the narrative into an article and stress-tested it against independent sources.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 2FaI2Fen1mQ
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

StatQuest host Josh Starmer sets out to explain three activation functions, GELU, SiLU and SwiGLU, to anyone who already knows ReLU, and the headline is simple: the first two differ from ReLU only subtly, the third can morph into far wilder shapes, yet all three share a family trait of resisting the urge to memorize training data.

The visible differences are small but precise: where ReLU flatlines at zero for every negative input and kinks sharply at the origin, GELU and SiLU dip slightly below zero between roughly minus two and zero, round off the corner into a smooth differentiable bend, and return values a touch smaller than the input on the positive side.

SwiGLU plays a different game: it can mimic a smoothed ReLU, but its extra parameters let it bend into shapes its cousins cannot reach, which is exactly why the episode treats it as the flexible heavyweight of the trio rather than a mere variant.

The payoff the episode promises is generalization: a large ReLU network draped over scattered training points hugs every dot including the noise, while the same capacity built on the new activations traces a calmer middle path that is likelier to predict unseen points well.

The story rewinds to before 2010, when sigmoid ruled: its zero-to-one output echoed a neuron either firing or silent, its S-curve mimicked the transition between those states, and its smoothness everywhere made backpropagation straightforward, so shallow networks trained happily.

Depth broke the spell: far from zero the sigmoid slope collapses toward nothing, gradients shrink to whispers, each optimization step crawls, and a network that should have traced the ideal wavy fit stalls into a timid curve that misses the data, the classic underfit.

Around 2010 ReLU took over with a bent line almost suspiciously simple: negatives become zero, positives pass through unchanged, the slope is either zero or one, so large inputs keep gradients alive and deep stacks finally become trainable, a run crowned in 2017 inside the first Transformer.

But capacity became the new trap: a giant ReLU network does not underfit anymore, it overfits, wrapping every quirk of the sample, and shrinking the network is no real answer since nobody can test a hundred thousand sizes one unit at a time to find the sweet spot.

Dropout was the first systematic fix the episode recalls: during training random subsets of units are silenced, each step effectively trains a thinner network, and at prediction time the full model behaves like an average over all those thin networks, pulling the wobbly ReLU fit closer to the ideal; but the fix has two costs, machinery bolted onto training and blindness to signal strength, since dropout kills units without asking whether their input is large or small, which motivates making survival odds depend on the input itself.

The derivation starts from a toy one-input network under fifty-percent dropout, where an input of minus two averages to minus one, then replaces the coin flip with a curve: the Gaussian cumulative curve maps minus two to a two-percent survival chance for an average of minus zero point zero four, minus one to sixteen percent for minus zero point one six, with zero, zero point eight four at one, and one point nine six at two falling on the emerging GELU shape, compactly written as the input times its own survival probability.

Swapping the bell-curve mapping for the familiar sigmoid yields SiLU, a near-twin of GELU that thrives in similar settings and leaves practitioners choosing mostly by ecosystem, while SwiGLU adds a learned weight, a Swish-like branch and a second projection multiplied together through a gating junction, a design whose own paper offers no causal story, only results, plus the host's two hunches that costlier units force leaner networks and that trainable shapes let the network decide for itself, a bet Meta's model family visibly took.

Visualization: nodesdaily AI

AI commentary

"I used to treat activation functions as plumbing between layers, but this episode changed my mind: seeing dropout turn into a curve, and that curve turn into GELU, made the whole modern LLM stack feel less like magic and more like a chain of reasonable bets."

AI assessment

The strongest case for the old guard is cost and sufficiency, and I take it seriously: scaled-ReLU studies show tuned ReLU variants can still train vision transformers competitively, and inference work that strips GELU out of models exists precisely because smooth activations are more expensive at serving time, so the upgrade is not free.

What the episode cannot give me is a number: its overfitting story is told with dot clouds and hand-drawn curves, while the hard evidence lives in the papers, where the GELU study reports gains across vision, language and speech tasks and practitioners note the SiLU-versus-GELU gap is usually small enough that engineering convenience decides.

For verifiability I weigh the ecosystem over the anecdote: the SwiGLU paper itself declines to explain why the design works, so I treat industry adoption, from Llama and Mistral to Qwen and DeepSeek standardizing on gated SiLU blocks, as the real confirmation, and any single overfitting claim as something to re-test before deciding.

My practical read is this: I would default to SwiGLU-style blocks for a new transformer, keep GELU where a codebase already runs well on it, and reach for plain ReLU only when inference budget or hardware dictates, because the smooth family buys robustness while the bill lands in compute.

Sources

10 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

gelu · silu · swiglu · activation function · relu · dropout · statquest

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…