The Eddington Luminosity in Spherical Accretion and Beyond

The Eddington luminosity is the point at which the radiation coming out of an object pushes on the surrounding gas as hard as the object’s gravity pulls it in. Above that luminosity, material near the source is pushed away instead of falling in. It sets a rough ceiling on how bright an accreting object can be, a rough ceiling on how massive a star can get, and a rough floor on how long a black hole needs to grow. The formula is short enough to memorise, which is a shame, because the derivation takes about ten lines and every one of those lines contains an assumption you can later break on purpose.

Set up the simplest possible situation. A compact object of mass $M$ sits at the origin and radiates isotropically with luminosity $L$. Around it is fully ionised hydrogen, which means free protons and free electrons and nothing else. The gas is optically thin, so a photon travelling outward has a small chance of interacting and mostly just leaves. Ask what happens to one electron-proton pair sitting at radius $r$.

Start with the radiation. Luminosity is energy per unit time, and if the source is isotropic that energy is spread evenly over a sphere of area $4\pi r^2$, so the energy flux at radius $r$ is $$F=\frac{L}{4\pi r^{2}}.$$ That’s energy per unit area per unit time. What we actually need is momentum, because forces come from momentum transfer.

A photon of energy $E$ carries momentum $p=E/c$. That follows from the relativistic energy-momentum relation $E^2=p^2c^2+m^2c^4$ with $m=0$, and it’s also what classical electromagnetism gives you for the ratio of momentum density to energy density in a plane wave. So a stream carrying energy flux $F$ also carries momentum flux $F/c$, and momentum flux has units of pressure. This quantity, $F/c$, is the radiation pressure available to push on anything that gets in the way.1It helps to see $p=E/c$ from more than one direction. In classical electromagnetism it comes from the Poynting vector and the Maxwell stress tensor: a plane wave has momentum density $S/c^2$ and energy density $S/c$, so their ratio is $1/c$. In quantum terms each photon carries $E=h\nu$ and $p=h\nu/c=h/\lambda$. Both give the same momentum flux, which is why the classical derivation of the Eddington limit needs no quantum mechanics at all.

Now we need to know how much of that momentum a single electron actually receives, which means we need a cross section. Calling a cross section an area is useful, but it shouldn’t be taken literally. The electron isn’t a small disk of area $6.65\times10^{-25}\ \mathrm{cm^2}$. A cross section packages an interaction probability into the dimensions of an area: if a beam delivers a certain number of photons per square centimetre per second, multiplying that flux by $\sigma$ gives the rate at which one electron scatters them. The area belongs to the interaction, and it tells you nothing about how big the electron is. With that understood, a particle with momentum-transfer cross section $\sigma$ sitting in a stream of momentum flux $F/c$ receives net momentum at a rate $\sigma F/c$, and a rate of momentum gain is a force.

The relevant process here is Thomson scattering, the elastic scattering of low-energy photons off free electrons, and it’s worth deriving rather than looking up because the derivation tells you exactly when it stops being valid. Take an electromagnetic wave with electric field amplitude $E_0$ passing a free electron. The field accelerates the electron with $a=eE_0\cos\omega t/m_e$. An accelerating non-relativistic charge radiates according to the Larmor formula, $P=2e^2a^2/3c^3$, and time-averaging $\cos^2$ gives a factor of $1/2$, so $$\langle P\rangle=\frac{2e^{2}}{3c^{3}}\cdot\frac{e^{2}E_0^{2}}{2m_e^{2}}=\frac{e^{4}E_0^{2}}{3m_e^{2}c^{3}}.$$ The incident flux is the time-averaged Poynting vector, $\langle S\rangle=cE_0^2/8\pi$. The cross section is the ratio of power scattered to flux incident, $$\sigma_T=\frac{\langle P\rangle}{\langle S\rangle}=\frac{8\pi}{3}\left(\frac{e^{2}}{m_ec^{2}}\right)^{2}=\frac{8\pi}{3}r_e^{2}=6.65\times10^{-25}\ \mathrm{cm^{2}},$$ where $r_e=e^2/m_ec^2=2.82\times10^{-13}$ cm is the classical electron radius.

That derivation hands us something useful for free. The cross section scales as $1/m^2$, so the same calculation applied to a proton gives a cross section smaller by $(m_e/m_p)^2=(1/1836)^2\approx3\times10^{-7}$. Protons scatter light about three million times less effectively than electrons. So when we say the radiation force acts on the electrons and ignore the protons, that’s a conclusion with a number behind it.

So the outward radiation force on one free electron is $$f_{\rm rad}=\sigma_T\,\frac{F}{c}=\frac{\sigma_T L}{4\pi r^{2}c}.$$

LrF = L / 4πr²momentum flux F/cσ_Tone free electronforcef_rad = σ_T F / c
The geometry of the calculation. Luminosity $L$ spreads over a sphere of area $4\pi r^2$ to give an energy flux $F$, and since each photon carries momentum $E/c$ the same stream carries a momentum flux $F/c$. A free electron intercepts that stream over an effective area $\sigma_T$ and receives net momentum at a rate that constitutes a force. Nothing in this picture is quantum mechanical.

There’s something slightly suspicious about what comes next. The radiation force we just wrote contains an electron cross section. The gravitational force is about to contain a proton mass. Why are we allowed to balance two forces that act on two different particles?

The answer is electrostatics. Gravity pulls on mass, and essentially all the mass is in the protons since $m_p/m_e\approx1836$, while radiation pushes on the electrons. If the electrons started to drift outward relative to the protons, that would immediately create a charge separation, which creates an electric field, which pulls the electrons back and drags the protons out. The question is how much separation is needed, and the answer is almost none, because the Coulomb interaction is absurdly stronger than gravity. For a single proton-electron pair, $$\frac{e^{2}}{G m_p m_e}=2.3\times10^{39}.$$ A fractional charge imbalance far too small to measure supplies an electric field far larger than anything gravity or radiation is doing. The plasma therefore moves as a neutral fluid, and the correct bookkeeping is to treat a proton-electron pair as the unit: it feels the radiation force of one electron and the gravitational pull of one proton.2You can also make this concrete. To transmit a gravitational force $GMm_p/r^2$ to an electron by electrostatic means, you need a field $E$ with $eE=GMm_p/r^2$, which for a spherical geometry corresponds to a net enclosed charge of $Q=GMm_p/e$. For a solar mass that’s about $10^{21}$ elementary charges, which sounds like a lot until you notice that a single gram of ionised hydrogen contains $6\times10^{23}$ electrons. The required fractional imbalance is tiny, and the plasma establishes it essentially instantly.

photonsepCoulomb couplingradiation pushes the electron outgravity pulls the proton ine² / (G m_p m_e) = 2.3 × 10³⁹, so the pair cannot separate
Why the mass in the formula is the proton mass while the cross section is the electron’s. Radiation pushes electrons, gravity holds protons, and the Coulomb attraction between them is roughly $10^{39}$ times stronger than their mutual gravity. The pair is welded together, so the balance is written per pair rather than per particle.

The inward gravitational force per pair is then $$f_{\rm grav}=\frac{GM(m_p+m_e)}{r^{2}}\simeq\frac{GMm_p}{r^{2}}.$$ Setting the two forces equal gives $$\frac{\sigma_T L}{4\pi r^{2}c}=\frac{GMm_p}{r^{2}},$$ and the radius cancels from both sides, leaving $$L_{\rm Edd}=\frac{4\pi GMm_pc}{\sigma_T}.$$

That cancellation is easy to read past, and it’s the reason the Eddington limit comes out as a statement about luminosity instead of a statement about some particular radius. Both forces fall off as $1/r^2$: gravity because it’s an inverse square law, and radiation because the flux from a point source dilutes as $1/r^2$. Moving farther away weakens gravity and radiation pressure in exactly the same proportion. Plotted against radius on log axes, both are straight lines of slope $-2$, which means they’re parallel and never cross. The balance is either satisfied at every radius or at no radius. There’s no critical distance inside which radiation wins. There’s a critical luminosity, and above it radiation wins at every radius at once.

log rlog forceL > L_EddgravityL < L_Eddevery line has slope −2, so they are parallel and never cross
Why the result is a critical luminosity rather than a critical radius. Gravity and radiation pressure both dilute as $1/r^2$, so on logarithmic axes both are straight lines of slope $-2$. Changing the luminosity slides the radiation line up and down without ever tilting it, so it either sits entirely above the gravity line or entirely below it.

Putting numbers in, $$L_{\rm Edd}=1.26\times10^{38}\left(\frac{M}{M_\odot}\right)\ \mathrm{erg\,s^{-1}}=3.3\times10^{4}\left(\frac{M}{M_\odot}\right)L_\odot.$$ The Sun’s actual luminosity is about $3\times10^{-5}$ of this value, so it sits very far from the balance point and nothing in its outer layers is close to being pushed off. Massive stars are a different matter, and we’ll come back to that.

Before breaking anything, rewrite the result in a form that makes the assumptions easier to isolate. Define the opacity $\kappa$ as cross section per unit mass, so that the force per gram of material is $\kappa F/c$. For pure ionised hydrogen there’s one electron per proton, so $\kappa_{\rm es}=\sigma_T/m_p=0.40\ \mathrm{cm^2\,g^{-1}}$, and the general statement becomes $$L_{\rm Edd}=\frac{4\pi GMc}{\kappa}.$$ The composition and the physics of the interaction are now quarantined inside a single symbol. Many of the corrections below can be understood either as changing $\kappa$, or as breaking one of the assumptions that let us write this simple force balance in the first place.

Check the units of that expression once, because it shows why $\kappa$ has to be an area per unit mass and not just an area: $$[G][M][c]\,[\kappa]^{-1}=\left(\frac{\mathrm{cm^{3}}}{\mathrm{g\,s^{2}}}\right)(\mathrm{g})\left(\frac{\mathrm{cm}}{\mathrm{s}}\right)\left(\frac{\mathrm{g}}{\mathrm{cm^{2}}}\right)=\frac{\mathrm{g\,cm^{2}}}{\mathrm{s^{3}}}=\mathrm{erg\,s^{-1}}.$$ Notice that the mass in $GM$ and the mass hidden inside $\kappa$ don’t cancel. That’s gravity pulling on the gas and radiation pushing on its cross section, showing up in the units.

Composition is the easiest thing to fix. What matters for scattering is the number of free electrons per gram, and hydrogen is unusual because it supplies one electron per nucleon while everything else supplies roughly one per two nucleons. For a fully ionised mix with hydrogen mass fraction $X$ and helium mass fraction $Y=1-X$, the electron density is $n_e=(\rho/m_p)(X+Y/2)=(\rho/m_p)(1+X)/2$, so $$\kappa_{\rm es}=\frac{\sigma_T}{m_p}\frac{1+X}{2}\simeq0.20\,(1+X)\ \mathrm{cm^{2}\,g^{-1}}.$$ For solar composition, $X=0.7$, this gives $0.34\ \mathrm{cm^2\,g^{-1}}$ and $L_{\rm Edd}=1.5\times10^{38}(M/M_\odot)\ \mathrm{erg\,s^{-1}}$. For hydrogen-free material the opacity halves and the Eddington limit doubles, purely because there are half as many electrons per gram to push on.3The commonly quoted $1.26\times10^{38}(M/M_\odot)$ assumes pure hydrogen. Papers on accreting neutron stars and X-ray bursts often use the solar or the helium value instead, so a factor of two discrepancy between two sources usually means they made different composition assumptions rather than that one of them is wrong. So check which $\kappa$ a quoted Eddington luminosity is built on before comparing two of them.

There’s a distinction buried in the derivation that causes a lot of confusion later, so let’s name it now. The force balance we wrote is local. It compares two forces on one parcel of gas at one place. It only turned into a statement about the total luminosity of the source because spherical symmetry let us write the local flux as $L/4\pi r^2$ everywhere at once. So $L_{\rm Edd}$ is a global luminosity scale derived from a local condition, and the translation between the two depends entirely on the geometry. When a source is described as super-Eddington in the literature, it’s often worth asking which of the two is meant, because a source can exceed the spherical Eddington luminosity as measured by a distant observer while the radiation force on every individual parcel of gas remains sub-Eddington.

That distinction is the first structural assumption to break. The derivation put a point source at the centre of a spherical shell of gas, so radiation and gravity were both radial and could be compared directly. Real accretion usually happens through a disc, and radiation escaping from an accretion flow doesn’t come out isotropically. If the gas arrives along the equatorial plane while the radiation leaves through low-density funnels along the rotation axis, the two never have to fight over the same material, and the luminosity inferred by a distant observer can exceed $L_{\rm Edd}$ while the accretion flow itself stays locally sub-Eddington.

discdiscradiation leavesthrough the funnelgas comes inalong the midplanegas comes inalong the midplanethe two directions are decoupled, so the global limit no longer applies
How anisotropy breaks the translation from local to global. The derivation compared an outward radiation flux with an inward gravitational pull along the same radial direction. In a disc geometry the inflow and the outflow occupy different solid angles, so the spherical formula no longer converts the local force balance into a single luminosity. Geometric beaming and anisotropic accretion are important ingredients in models of ultraluminous X-ray sources.

The second structural assumption is that the situation can be treated through an instantaneous force balance. The derivation compares two forces acting on the gas at a given moment, and by itself says nothing about how the gas responds with time. A transient can therefore exceed the limit for as long as it takes the surrounding material to accelerate, expand, or otherwise adjust. Classical novae and X-ray bursts on neutron stars both show episodes consistent with locally super-Eddington fluxes, and the associated observational signature is photospheric radius expansion, in which the envelope is lifted, expands, and settles back as the luminosity drops.

The third assumption, and the one that fails most often in practice, is that electron scattering is the only relevant opacity. For soft photons in a hot, fully ionised plasma, electron scattering provides a useful baseline: it’s what remains when the gas is too hot to have any bound electrons left to absorb anything. Bound-free, free-free, line and dust interactions can all add substantially to it, and since $L_{\rm Edd}\propto1/\kappa$, every addition lowers the true limit below the value the standard formula gives. But even this baseline isn’t universal, and we’ll see shortly that once photon energies approach $m_ec^2$ the Thomson approximation itself fails and the scattering opacity falls below $\kappa_{\rm es}$.

Spectral lines are the clearest example of an addition. A bound electron in an ion can absorb a photon whose energy matches a transition, and the cross section at line centre can exceed the Thomson cross section by many orders of magnitude. In the ultraviolet spectrum of a hot star there are very large numbers of such transitions from iron-group ions, and although each covers a narrow frequency range, together they intercept a substantial fraction of the flux. The effective opacity for a hot star’s wind can be tens to hundreds of times $\kappa_{\rm es}$, which is how O stars can drive strong winds while their continuum Eddington ratios remain well below unity. The star is sub-Eddington for electron scattering and super-Eddington for lines.4This is the basis of the Castor, Abbott and Klein theory of line-driven winds from 1975, which parametrises the extra push through a force multiplier applied to the electron-scattering radiation force. The multiplier depends on the ionisation state and on the velocity gradient in the wind, because a wind that’s accelerating Doppler-shifts each line into fresh continuum photons and keeps the lines from saturating. The velocity gradient is therefore part of the opacity, which is an unusual feature of the problem.

Dust is the more extreme addition. A grain absorbs across a broad continuum rather than in narrow lines, and for gas with an ordinary interstellar dust-to-gas ratio the ultraviolet opacity per gram of gas is of order a few hundred $\mathrm{cm^2\,g^{-1}}$, which is roughly a thousand times $\kappa_{\rm es}$. In dusty gas the relevant Eddington limit therefore drops by about three orders of magnitude, and objects that look comfortably sub-Eddington by the standard formula can be strongly super-Eddington for the material around them. This is why radiation pressure on dust matters for massive star formation and for feedback in galactic nuclei, and why the electron-scattering formula is a poor guide in those settings.5The estimate comes from the standard interstellar extinction per hydrogen column, $A_V/N_H\approx5\times10^{-22}$ mag cm$^2$, converted to an opacity per gram using a mean mass per hydrogen nucleus of about $1.4m_p$, which gives roughly $200\ \mathrm{cm^2\,g^{-1}}$ in the visual and two to three times that in the ultraviolet. It scales linearly with the dust-to-gas ratio, so it drops in low-metallicity environments and vanishes wherever grains have been destroyed.

0.111010010000.19ionised He, Fesolar mix, 0.34pure hydrogen, 0.40in hot star windsUV line opacity≈ 500dust in the UVopacity κ, cm² g⁻¹L_Edd ∝ 1/κ, so everything to the right of electron scatteringlowers the real limit below the textbook value
Opacity for a gram of material, on a logarithmic scale. Electron scattering sits at the left because it’s what remains when the gas is hot enough to have no bound electrons. Lines and dust sit far to the right, and since the Eddington luminosity is inversely proportional to opacity, they lower the relevant limit by factors of tens to thousands. The electron-scattering value is a convenient reference scale rather than a universal ceiling.

The fourth assumption is that the photons are soft. The Thomson derivation treated the electron as a non-relativistic charge oscillating in a classical wave, which requires the photon energy to be small compared to $m_ec^2=511$ keV. Above that, the photon recoils the electron appreciably, the scattering becomes inelastic, and the Klein-Nishina formula replaces the Thomson result. The cross section falls, roughly as $\ln(2x)/x$ for $x=h\nu/m_ec^2\gg1$. Hard photons are therefore less effective at pushing gas, and this is the one correction that pushes the limit up rather than down: a source radiating mainly above a few hundred keV can exceed the naive Eddington luminosity simply because its photons couple weakly to electrons.

photon energyσ / σ_T1 keV101001 MeV1010.5511 keV = m_e c²Thomson regime
The scattering cross section as a function of photon energy, in units of the Thomson value. The Thomson result is the flat portion on the left, valid while $h\nu\ll m_ec^2$. By $100$ keV the cross section is down by about a quarter, at $511$ keV it’s below half, and at several MeV it’s an order of magnitude smaller. This is the reason electron scattering cannot be treated as a hard floor on the opacity.

The fifth assumption is Newtonian gravity, which is uncomfortable given that the objects people apply this to are usually black holes and neutron stars. In the Schwarzschild metric two things change for a static observer at radius $r$. The local gravitational acceleration picks up a factor $(1-r_s/r)^{-1/2}$, and the flux from a source of luminosity $L_\infty$ as measured at infinity picks up a factor $(1-r_s/r)^{-1}$ from gravitational redshift and time dilation together. Balancing them gives $$L_{\rm Edd}^{\infty}=\frac{4\pi GMc}{\kappa}\sqrt{1-\frac{r_s}{r}}.$$ The two corrections partly cancel, which is why the Newtonian answer survives as well as it does, but not completely. At $3r_s$ the limit is $82$ percent of the flat-space value and at $1.5r_s$ it’s $58$ percent. Note also that the answer now depends on $r$, so the radius no longer cancels and the earlier statement about parallel lines fails.

The sixth assumption is that the gas is ordinary matter with protons in it. Within the same simple force-balance picture, an electron-positron pair plasma changes the bookkeeping completely: there are no protons carrying most of the mass, so gravity acts on a mass of order $m_e$ per scattering particle rather than $m_p$, while the scattering cross section remains of order $\sigma_T$. The corresponding Eddington scale therefore drops by roughly $m_e/m_p\approx1/1836$. Real pair plasmas involve additional physics such as pair creation, annihilation and more complicated radiative transfer, so this factor is what you get by pushing the original argument into a new regime, and the real problem carries more than that. Pair plasmas are thought to be relevant in the inner regions of some accretion flows and in gamma-ray burst fireballs, where corrections of this size are not small.

The seventh assumption is that the gas is smooth. The derivation implicitly treats the medium as uniform, so every gram sees the same flux. If the gas is clumpy, radiation escapes preferentially through the low-density channels between the clumps and deposits less momentum than a uniform calculation predicts. The effective opacity of a porous medium is lower than that of a smooth one with the same mean density, which raises the effective limit. This idea, usually called porosity, is one of the mechanisms proposed for the eruptions of luminous blue variables, which appear to radiate above their nominal limits for extended periods.

With the limit itself understood, two more quantities follow, and these are what the Eddington luminosity actually gets used for in practice. If an accreting object converts a fraction $\eta$ of the rest mass energy of the infalling material into radiation, so that $L=\eta\dot Mc^2$, then the accretion rate corresponding to the Eddington luminosity is $$\dot M_{\rm Edd}=\frac{L_{\rm Edd}}{\eta c^{2}}=\frac{4\pi GM}{\eta\kappa c}\approx2.2\times10^{-8}\left(\frac{0.1}{\eta}\right)\left(\frac{M}{M_\odot}\right)M_\odot\,\mathrm{yr^{-1}}.$$ The efficiency $\eta$ is about $0.057$ for a non-rotating black hole and up to about $0.42$ for a maximally rotating one, so this number carries a factor of several of uncertainty depending on spin.

The second derived quantity needs one piece of care that’s easy to miss. If $\dot M$ is the rest mass supplied to the hole and a fraction $\eta$ of it is radiated away, then the mass actually retained is $\dot M_{\rm BH}=(1-\eta)\dot M$. Carrying that through, the e-folding time for the black hole mass under Eddington-limited accretion is $$t_S=\frac{\eta}{1-\eta}\,\frac{\kappa c}{4\pi G}\approx50\ \mathrm{Myr}\quad(\eta=0.1),$$ rather than the $45$ Myr you get by dropping the $(1-\eta)$. The distinction is small at low efficiency and large at high efficiency: for a maximally spinning hole with $\eta=0.42$ it changes the answer from $189$ to $326$ Myr.6The $45$ Myr form is the one usually quoted, and for order-of-magnitude arguments the difference rarely matters. Check which version a paper is using when the argument turns on a factor of two, which it sometimes does in discussions of early black hole growth.

That number is behind a well-known problem. Growing a $10\ M_\odot$ seed into a $10^9\ M_\odot$ quasar takes about $18$ e-foldings, which at $50$ Myr each is roughly $900$ Myr of continuous Eddington-limited accretion with no idle periods at all. Quasars of that mass are observed less than a billion years after the Big Bang, so the budget is uncomfortably tight, which is why heavier seeds and episodes of super-Eddington accretion are both actively discussed.

The limit bites hardest in massive stars. Main sequence luminosity rises steeply with mass, faster than linearly, while the Eddington luminosity rises exactly linearly, so the ratio $\Gamma=L/L_{\rm Edd}$ climbs with mass. Using standard zero-age main sequence luminosities, a $10\ M_\odot$ star has $\Gamma\approx0.03$ and a $100\ M_\odot$ star roughly $0.3$ on the continuum opacity alone. Add line opacity and the effective ratio approaches unity, which is connected both to the large mass loss rates of the most massive stars and to the existence of an upper limit on stellar masses. For these objects the Eddington limit is doing real work on the structure of the star.

References and Footnotes

  • 1
    It helps to see $p=E/c$ from more than one direction. In classical electromagnetism it comes from the Poynting vector and the Maxwell stress tensor: a plane wave has momentum density $S/c^2$ and energy density $S/c$, so their ratio is $1/c$. In quantum terms each photon carries $E=h\nu$ and $p=h\nu/c=h/\lambda$. Both give the same momentum flux, which is why the classical derivation of the Eddington limit needs no quantum mechanics at all. ↩︎
  • 2
    You can also make this concrete. To transmit a gravitational force $GMm_p/r^2$ to an electron by electrostatic means, you need a field $E$ with $eE=GMm_p/r^2$, which for a spherical geometry corresponds to a net enclosed charge of $Q=GMm_p/e$. For a solar mass that’s about $10^{21}$ elementary charges, which sounds like a lot until you notice that a single gram of ionised hydrogen contains $6\times10^{23}$ electrons. The required fractional imbalance is tiny, and the plasma establishes it essentially instantly. ↩︎
  • 3
    The commonly quoted $1.26\times10^{38}(M/M_\odot)$ assumes pure hydrogen. Papers on accreting neutron stars and X-ray bursts often use the solar or the helium value instead, so a factor of two discrepancy between two sources usually means they made different composition assumptions rather than that one of them is wrong. So check which $\kappa$ a quoted Eddington luminosity is built on before comparing two of them. ↩︎
  • 4
    This is the basis of the Castor, Abbott and Klein theory of line-driven winds from 1975, which parametrises the extra push through a force multiplier applied to the electron-scattering radiation force. The multiplier depends on the ionisation state and on the velocity gradient in the wind, because a wind that’s accelerating Doppler-shifts each line into fresh continuum photons and keeps the lines from saturating. The velocity gradient is therefore part of the opacity, which is an unusual feature of the problem. ↩︎
  • 5
    The estimate comes from the standard interstellar extinction per hydrogen column, $A_V/N_H\approx5\times10^{-22}$ mag cm$^2$, converted to an opacity per gram using a mean mass per hydrogen nucleus of about $1.4m_p$, which gives roughly $200\ \mathrm{cm^2\,g^{-1}}$ in the visual and two to three times that in the ultraviolet. It scales linearly with the dust-to-gas ratio, so it drops in low-metallicity environments and vanishes wherever grains have been destroyed. ↩︎
  • 6
    The $45$ Myr form is the one usually quoted, and for order-of-magnitude arguments the difference rarely matters. Check which version a paper is using when the argument turns on a factor of two, which it sometimes does in discussions of early black hole growth. ↩︎
Posted in Expository | Tagged , , , , , , , | 1 Comment

How a Molecular Cloud Collapses Into a Star

Making a star means compressing gas by something like twenty-four orders of magnitude. The warm neutral interstellar medium, the ordinary atomic gas between the stars, holds roughly a third of a hydrogen atom per cubic centimetre at a temperature of several thousand kelvin, which puts its density near $10^{-24}$ grams per cubic centimetre. The mean density of the Sun is $1.41$ grams per cubic centimetre. The Galaxy takes material from one end of that range and pushes it to the other, currently at a rate of one or two solar masses per year, and it has been doing so for most of its life. The question is what makes a patch of thin interstellar gas stop being gas and become a star.

density, g cm⁻³10⁻²⁴10⁻²⁰10⁻¹⁶10⁻¹²10⁻⁸10⁻⁴1warm neutral medium, n ≈ 0.3 cm⁻³, 8000 Kmolecular cloud, n ≈ 100 cm⁻³, 10 Kdense core, n ≈ 10⁴ cm⁻³pre-stellar core, n ≈ 10⁶ cm⁻³first hydrostatic core, R ≈ 5 AUthe star24 ordersof magnitude
The problem as a density ladder. Each rung is a physically different state of the gas. Gravity matters throughout, but what resists it changes as the collapse proceeds. Thermal pressure, turbulence and magnetic fields dominate at low density. Cooling and radiative transfer become increasingly important as the gas gets denser. Eventually the gas becomes opaque, heats up, dissociates molecular hydrogen and forms a protostar.

Gravity is the only force available that can gather matter over large distances, and it has to compete with pressure. Pressure doesn’t know in advance that a cloud is trying to collapse. It reacts to compression by launching sound waves. So the first question is a race: can gravity make a density enhancement grow faster than pressure can communicate across it and smooth it away? Jeans turned that into an equation in 1902.

Take an isothermal self-gravitating fluid, $$\frac{\partial\rho}{\partial t}+\nabla\!\cdot(\rho\mathbf{v})=0, \qquad \frac{\partial\mathbf{v}}{\partial t}+(\mathbf{v}\cdot\nabla)\mathbf{v}=-\frac{\nabla P}{\rho}-\nabla\Phi, \qquad \nabla^{2}\Phi=4\pi G\rho,$$ with $P=c_s^{2}\rho$. Take a uniform static background of density $\rho_0$, perturb it as $\rho=\rho_0+\rho_1$, $\mathbf v=\mathbf v_1$, $\Phi=\Phi_0+\Phi_1$, and keep only terms first order in the perturbations.1There is a genuine inconsistency hidden here, traditionally called the Jeans swindle. An infinite medium with constant density cannot simultaneously satisfy $\nabla^2\Phi_0=4\pi G\rho_0$ and have no background gravitational acceleration. The usual calculation ignores the gravity of the uniform background and keeps the gravity of the perturbation. More careful treatments using finite systems or expanding backgrounds recover essentially the same instability criterion, so the answer survives even though the original setup is not mathematically self-consistent.

Differentiate the continuity equation with respect to time, take the divergence of the momentum equation, and use Poisson’s equation to eliminate $\Phi_1$. The three equations become one, $$\frac{\partial^{2}\rho_1}{\partial t^{2}}-c_s^{2}\nabla^{2}\rho_1-4\pi G\rho_0\rho_1=0,$$ and looking for a plane wave $\rho_1\propto\exp[i(\mathbf k\cdot\mathbf x-\omega t)]$ collapses the whole problem to $$\omega^{2}=c_s^{2}k^{2}-4\pi G\rho_0.$$

Take the terms one at a time. Without gravity this is $\omega=c_sk$, an ordinary sound wave. Gravity subtracts the same amount from $\omega^2$ at every wavelength, while pressure’s contribution grows as $k^2$ and so becomes increasingly effective on small scales. At large $k$, meaning short wavelengths, pressure wins, $\omega^2>0$, and the perturbation oscillates. At small $k$ gravity wins, $\omega^2<0$, $\omega$ becomes imaginary, and one of the two solutions grows exponentially.

ω² ω² = c s ²k² − 4πGρ₀ −4πGρ₀ k J ² unstableω² < 0exponential growthsound wavesω² > 0 long wavelengths: λ > λ J short wavelengths: λ < λ J
The Jeans dispersion relation plotted against $k^2$. Gravity contributes the constant term $-4\pi G\rho_0$, while pressure contributes $c_s^2k^2$. At $k=k_J$ the two exactly cancel. Longer wavelength perturbations have $\omega^2<0$ and grow exponentially, while shorter wavelength perturbations remain ordinary sound waves.

The boundary sits at $k_J=\sqrt{4\pi G\rho_0}/c_s$, or in wavelength $$\lambda_J=\frac{2\pi}{k_J}=c_s\sqrt{\frac{\pi}{G\rho_0}}.$$ In this idealised medium, perturbations larger than $\lambda_J$ are gravitationally unstable and smaller ones behave like sound waves. Real molecular clouds are neither uniform nor static nor unmagnetised, so $\lambda_J$ shouldn’t be treated as a sharp boundary in nature. What survives is the competition it exposes: pressure grows stronger on small scales, while self-gravity doesn’t weaken when you look at a larger piece of gas.

Turning the Jeans length into a mass gives the more useful quantity. Taking a sphere whose diameter is $\lambda_J$, $$M_J=\frac{4\pi}{3}\rho_0\left(\frac{\lambda_J}{2}\right)^{3}=\frac{\pi^{5/2}}{6}\frac{c_s^{3}}{G^{3/2}\rho_0^{1/2}}\approx2.9\,\frac{c_s^{3}}{G^{3/2}\rho_0^{1/2}}.$$ The prefactor matters less than the scaling, $M_J\propto c_s^3\rho^{-1/2}\propto T^{3/2}\rho^{-1/2}$, and you can recover that scaling without any perturbation theory at all. The only combination of $G$, $c_s$ and $\rho$ with units of mass is $c_s^3G^{-3/2}\rho^{-1/2}$. The full calculation tells you why that combination matters and supplies the numerical factor.

Put in numbers for a cold dense core. Molecular gas with helium has a mean molecular weight $\mu\simeq2.33$, so at $T=10$ K the sound speed is $c_s=\sqrt{k_BT/\mu m_H}\simeq0.19\ \mathrm{km\,s^{-1}}$. At $n=10^4\ \mathrm{cm^{-3}}$ this gives $\lambda_J\simeq0.21$ pc and $M_J\simeq2.9\,M_\odot$. A few solar masses spread over a few tenths of a parsec isn’t an absurd theoretical scale. It’s close to the scale of the dense structures actually found inside nearby molecular clouds.

The temperature dependence matters even more. Since $M_J\propto T^{3/2}$, going from $8000$ K to $10$ K lowers the characteristic unstable mass by $(800)^{3/2}\approx2\times10^4$. Warm diffuse gas is very hard to make gravitationally unstable on stellar mass scales, and cold gas is much easier, which ties star formation directly to cooling.

The details are a little more interesting than saying gas must become molecular in order to cool. Atomic species such as ionised carbon cool neutral gas efficiently, and molecular hydrogen is not always the dominant coolant in ordinary Galactic molecular clouds. What really matters is shielding from the interstellar ultraviolet field. Once enough column density builds up, photoelectric heating weakens, molecules survive, dust becomes cold, and temperatures near $10$ K become possible. Becoming cold and becoming molecular happen together, though one is not simply the cause of the other.

The Taurus molecular cloud imaged in the far infrared by Herschel, showing a network of filaments with compact dense structures
The Taurus molecular cloud in far infrared dust emission. Real molecular clouds look nothing like the uniform medium used in the Jeans calculation. Much of their dense gas is arranged into filaments, and dense cores are often found along those filaments. Herschel observations led to the suggestion that many nearby filaments have an inner width around $0.1$ pc. Whether that width is truly universal is still debated, because resolution, distance, tracer choice and the way a filament is defined can all change the measured width.2Image credit: ESA/Herschel/PACS, SPIRE/Gould Belt survey Key Programme. The characteristic $0.1$ pc width was emphasised in the Herschel Gould Belt work, but later studies have questioned how universal it really is. The safe statement is that filaments are common and closely connected to dense core formation. The exact distribution of their widths remains an active observational problem.

There’s another way to phrase the same stability problem that sits closer to what an observed core actually looks like. A real core is finite and can be confined by the pressure of the gas around it, and an isothermal pressure-confined sphere has a maximum stable mass, the Bonnor-Ebert mass, $$M_{\rm BE}\simeq1.18\,\frac{c_s^{4}}{G^{3/2}P_{\rm ext}^{1/2}}.$$ Above that, no hydrostatic configuration exists for the assumed temperature and external pressure. The details differ from the infinite-medium Jeans calculation, but the message is nearly the same. A cold enough and massive enough concentration of gas cannot find a pressure-supported equilibrium.

Suppose it crosses that line. Strip away pressure completely, start with a uniform sphere, and a shell at initial radius $r_0$ feels only the mass inside it, so $\ddot r=-GM/r^2$. Integrating once gives $\tfrac12\dot r^2=GM(1/r-1/r_0)$, and a further integration using $r=r_0\cos^2\xi$ gives the free-fall time $$t_{\rm ff}=\sqrt{\frac{3\pi}{32G\rho_0}}.$$ The radius has disappeared. For a uniform sphere the collapse time depends only on density, so every shell arrives at the centre together. At $n=10^4\ \mathrm{cm^{-3}}$ this is $0.34$ Myr, and at the more typical mean density of a molecular cloud, $n\simeq100\ \mathrm{cm^{-3}}$, about $3.4$ Myr.

Those numbers fail badly against observation. The Milky Way contains of order $10^9\,M_\odot$ of molecular gas, and if all of it turned into stars on a cloud free-fall time the star formation rate would be roughly $$\dot M_\star\sim\frac{M_{\rm mol}}{t_{\rm ff}}\sim\frac{10^9\,M_\odot}{3.4\ \mathrm{Myr}}\sim300\ M_\odot\,\mathrm{yr^{-1}}.$$ The real Galactic rate is around $1$ to $2\,M_\odot\,\mathrm{yr^{-1}}$, so the naive estimate misses by roughly two orders of magnitude.

A convenient way to write the discrepancy is the efficiency per free-fall time, $\epsilon_{\rm ff}\equiv\dot M_\star t_{\rm ff}/M_{\rm gas}$. On molecular-cloud scales the inferred values are of order one percent, often somewhere between $0.3$ and $3$ percent, with substantial scatter from cloud to cloud and through a cloud’s lifetime. The point isn’t that nature picked exactly $0.0100$. It’s that the answer is nowhere near unity, and something is preventing most molecular gas from falling straight in.3The low value of $\epsilon_{\rm ff}$ is one of the central empirical constraints on star formation theory. Modern cloud-lifecycle work also makes clear that individual clouds need not form stars at a perfectly constant rate. A cloud can evolve through quiet and active phases, so an average value near one percent does not mean every cloud converts exactly one percent of its mass every free-fall time.

The usual bookkeeping device is the virial theorem. The full scalar version contains kinetic energy, gravitational energy, magnetic terms, surface pressure and time-dependent terms, but ignore most of those and approximate the cloud as a uniform sphere. Then $W=-\tfrac35 GM^2/R$ and $T=\tfrac32 M\sigma^2$, with $\sigma$ the one-dimensional velocity dispersion, and their ratio gives the virial parameter $$\alpha_{\rm vir}\equiv\frac{5\sigma^{2}R}{GM}.$$ For this simple model $\alpha_{\rm vir}=1$ corresponds to $2T=|W|$.

Values of order unity tell us kinetic and gravitational energies are comparable. They don’t prove a cloud is sitting in static equilibrium. A cloud can be accreting, contracting, losing material, confined by external pressure or threaded by a dynamically important magnetic field and still have an order-unity virial parameter. That matters because observed molecular clouds commonly have $\alpha_{\rm vir}$ of order one or a few. They aren’t behaving like cold pressureless spheres in free fall, but the virial parameter alone doesn’t tell us what they are doing.

The most obvious extra ingredient is turbulence. At $10$ K the sound speed is only $0.19\ \mathrm{km\,s^{-1}}$, yet molecular lines are routinely much broader than that, and the non-thermal dispersion increases with the size of the region measured. Larson identified this in 1981, and later CO surveys found relations roughly of the form $$\sigma_v\sim0.7\left(\frac{R}{1\ \mathrm{pc}}\right)^{1/2}\ \mathrm{km\,s^{-1}},$$ although neither the slope nor the normalisation is universal and surface density and environment both matter. Taken literally at $R=10$ pc it gives a few kilometres per second, more than ten times the sound speed of cold molecular gas.

Supersonic turbulence then does two things at once. It resists global collapse by supplying kinetic energy, and it also creates the structures that collapse first, because supersonic flows form shocks and shocks compress gas. In idealised isothermal turbulence without self-gravity the density distribution is approximately lognormal, and once self-gravity becomes important the densest collapsing regions develop a power-law tail. Those are the regions where the local Jeans mass becomes small and collapse runs ahead. Turbulence moves mass around, builds filaments and sheets, produces rare high-density regions and changes where gravity first gets to win, which is why many theories of the star formation rate begin by asking what fraction of a turbulent density field becomes gravitationally unstable.

Magnetic fields complicate the picture again, because a field doesn’t resist all compression equally. Gas moves much more easily along a field than across it, and under ideal flux freezing matter and flux move together, so the useful quantity is the mass-to-flux ratio $M/\Phi$. For a simple sheet-like geometry the critical value is of order $$\left(\frac{M}{\Phi}\right)_{\rm crit}\simeq\frac{1}{2\pi\sqrt G}, \qquad M_\Phi\simeq\frac{\Phi}{2\pi\sqrt G}\sim\frac{R^{2}B}{2\sqrt G}.$$ At $R=1$ pc and $B=10\,\mu\mathrm{G}$ that is about $93\,M_\odot$.

If a cloud is magnetically subcritical, the field can prevent collapse across the field lines in the idealised flux-frozen limit. If it’s supercritical, gravity is strong enough that the field can’t hold the cloud up forever on its own. That doesn’t make the field irrelevant, since a supercritical field can still change the geometry, the collapse rate, the fragmentation and the angular momentum of the gas.

subcritical: M / Φ < (M / Φ)critsupercritical: M / Φ > (M / Φ)critgas can move along Bcross-field collapse is resistedgravity pulls gas across Band drags the field inward
Why the mass-to-flux ratio matters. Gas can move relatively easily along magnetic field lines, so even a subcritical cloud can become flattened. What the field strongly resists is collapse across the field. In a supercritical region gravity is strong enough to pull material across the field lines, dragging them inward as the collapse proceeds.

Observationally the situation is messy. Zeeman splitting gives the line-of-sight field strength, dust polarisation gives the field orientation projected on the sky, and neither gives the full three-dimensional field without assumptions. The broad picture is that magnetic fields are dynamically important and that dense star-forming gas is generally not enormously subcritical. Exactly how close clouds and cores sit to the critical value is harder to establish than a single number makes it sound.

Flux freezing isn’t exact either. Molecular gas is only weakly ionised, so charged particles couple directly to the field while neutrals drift relative to them. This is ambipolar diffusion, and it lets matter move inward without dragging the full flux along, so the mass-to-flux ratio of a dense region can increase. For a long time slow ambipolar diffusion through a magnetically supported cloud was one of the standard explanations for inefficient star formation. It still matters, but that clean picture has been replaced by a mixed one in which turbulence, gravity, ambipolar diffusion, Ohmic dissipation, the Hall effect, accretion from the surrounding cloud and stellar feedback all operate somewhere in the problem, with the dominant one depending on density and environment.

Once a core does collapse, another idealised calculation becomes useful. A uniform sphere collapses everywhere at once because every shell shares a free-fall time, but real dense cores are centrally concentrated. The classic analytic model is Shu’s inside-out collapse of a singular isothermal sphere, which starts from $\rho(r)=c_s^2/2\pi Gr^2$. Once collapse begins at the centre an expansion wave moves outward at the sound speed, gas inside the wave falls inward, gas outside hasn’t yet responded, and the accretion rate is constant: $$\dot M=0.975\,\frac{c_s^{3}}{G}\simeq1.5\times10^{-6}\,M_\odot\,\mathrm{yr^{-1}}\quad(T=10\ \mathrm{K}).$$ A solar mass at that rate takes several hundred thousand years to assemble.

The scaling follows from dimensions again, since the only variables in the isothermal problem are $c_s$ and $G$ and the combination with units of mass per time is $c_s^3/G$. Warmer gas therefore has a much larger characteristic accretion rate. The number $0.975$ belongs to one very specific initial condition, though. Real cores don’t begin as exact singular isothermal spheres and real accretion isn’t constant. Simulations and observations both find time dependence, bursts and strong effects from rotation, magnetic fields and the surrounding environment. Shu’s solution is useful because it sets a clean scale, not because protostars follow it like a timetable.

The infalling material also doesn’t jump straight from molecular-cloud density to stellar density. At first the gas radiates away most of the energy released by compression, so the collapse stays close to isothermal. Eventually the central region becomes optically thick to its own infrared radiation, cooling can no longer keep up, the temperature rises, pressure stiffens, and collapse temporarily stops. A small hydrostatic object a few astronomical units across appears: the first hydrostatic core, predicted by Larson in 1969. Finding one has been much harder, since the phase is short-lived and deeply buried, and although several objects have been proposed, an unambiguous observational identification remains difficult.4Modern radiation-hydrodynamic calculations give a detailed picture of the first and second core stages, but the observational signature of a genuine first hydrostatic core is not unique. Several good candidates exist. It is safer to call them candidates than to say the phase has been securely observed as a class of objects.

The first core doesn’t survive. As its centre approaches roughly $2000$ K, molecular hydrogen begins to dissociate, and breaking an $\mathrm{H_2}$ molecule costs $4.48$ eV. Energy from compression now goes into dissociation instead of efficiently raising the temperature, the effective adiabatic index drops, pressure support weakens, and the gas enters a second collapse. That continues until most of the hydrogen in the centre has become atomic and the equation of state stiffens again. What’s left behind is the second hydrostatic core, which is what we normally call the protostar.

This change in thermodynamics also matters for fragmentation. While gas cools efficiently a collapsing region can keep breaking into smaller unstable pieces, and once it becomes opaque and starts heating rapidly, further fragmentation becomes much harder. That produces an opacity-limited minimum fragment scale. The exact mass depends on the thermodynamics, the rotation and the later accretion, so it shouldn’t be confused with a hard lower limit on the final mass of a brown dwarf. It is the first hint that microscopic physics can leave a preferred mass scale inside a problem that looked scale-free when we started with Jeans.

The Pillars of Creation in the Eagle Nebula imaged in the near infrared by JWST
The Pillars of Creation in M16 seen by JWST’s NIRCam. The pillars are dense parts of a molecular cloud being eroded by ultraviolet radiation from nearby massive stars. Young stars are still forming inside them, and some of the small knots and wave-like structures are associated with outflows from embedded young objects. Star formation and destruction are happening in the same piece of cloud at the same time.5Image credit: NASA, ESA, CSA, STScI; image processing by Joseph DePasquale, Anton M. Koekemoer and Alyssa Pagan. JWST NIRCam image released in October 2022.

There’s another problem waiting for the collapsing gas, and it has nothing to do with whether gravity is strong enough. The gas rotates. A core with radius $0.1$ pc and a characteristic rotational speed of $0.1\ \mathrm{km\,s^{-1}}$ has a specific angular momentum of order $j_{\rm core}\sim Rv\sim3\times10^{21}\ \mathrm{cm^2\,s^{-1}}$, while the specific angular momentum actually stored in the Sun’s spin is only of order $10^{15}\ \mathrm{cm^2\,s^{-1}}$. The precise comparison depends on the internal rotation profile and on how you define the core’s rotation, but the mismatch is around six orders of magnitude, and angular momentum can’t simply be destroyed.

10¹⁴10¹⁶10¹⁸10²⁰10²²Sun’s spin≈ 10¹⁵Jupiter’s orbit≈ 10²⁰dense core≈ 3 × 10²¹specific angular momentum, cm² s⁻¹logarithmic scale≈ 3 × 10⁶ in specific angular momentum has to be redistributed
The angular momentum problem on a logarithmic scale. A dense molecular core contains millions of times more specific angular momentum than the spin of the star it can eventually produce. That angular momentum cannot disappear, so it has to be transferred elsewhere. The Solar System gives a familiar example: the Sun contains almost all the mass, while the planetary orbits contain most of the angular momentum. During star formation, discs and outflows provide the main routes for that redistribution.

The first part of the solution is a disc. Gas with too much angular momentum to fall directly onto the centre reaches a centrifugal barrier and settles into rotation, with $\Omega(R)=\sqrt{GM_\star/R^3}$ for a Keplerian disc. Making a disc doesn’t solve anything by itself, because if every ring keeps its angular momentum forever the disc just sits there. For matter to move inward, angular momentum has to move outward.

Shakura and Sunyaev gave a way to describe that transport without knowing its cause, writing the effective viscosity as $\nu=\alpha c_sH$ with $H$ the disc scale height and $\alpha$ packaging the unknown physics into one dimensionless number. For years the most famous candidate for supplying it was the magnetorotational instability. Take two pieces of a differentially rotating disc and connect them weakly with a field line. The inner piece orbits faster, magnetic tension transfers angular momentum from the inner piece to the outer one, the inner piece falls further inward while the outer one moves further out, and the increased differential motion lets the instability grow.

The difficulty is that a protoplanetary disc isn’t an ideal conducting fluid. Much of it is cold and only weakly ionised, and Ohmic resistivity, ambipolar diffusion and the Hall effect all modify the magnetic coupling, suppressing vigorous MRI turbulence over large parts of the disc. The modern picture puts much more weight on magnetised disc winds, where large-scale fields threading the disc carry angular momentum vertically away from it and let gas accrete without strong turbulence throughout. MRI probably still matters in some regions and gravitational torques matter in young massive discs, but there’s no single universal $\alpha$ mechanism.6The history is worth noting. Velikhov and Chandrasekhar studied the instability decades before Balbus and Hawley recognised its importance for accretion discs in 1991, after which MRI became almost synonymous with angular momentum transport in disc theory. Non-ideal MHD changed that. Current work increasingly treats magnetic winds as a major and possibly dominant route for angular momentum loss in protoplanetary discs, with the balance between winds, MRI, hydrodynamic instabilities and gravitational torques depending on radius and evolutionary stage.

ALMA millimetre image of the protoplanetary disc around HL Tauri showing bright rings and dark gaps
The disc around HL Tauri in millimetre continuum emission from ALMA. The bright rings and dark gaps showed that strong disc substructure can appear astonishingly early. Planets are one possible explanation for some of these gaps, though not the only one, since dust growth, pressure structures and other disc processes can also produce rings. The important point here is that the smooth-disc approximation breaks down very quickly in real systems.7Image credit: ALMA (ESO/NAOJ/NRAO), released in 2014 from ALMA’s long-baseline campaign.

A young disc can also become gravitationally unstable on its own, and the disc version of the Jeans argument is Toomre’s criterion. For a thin gas disc with surface density $\Sigma$, sound speed $c_s$ and epicyclic frequency $\kappa$, define $$Q=\frac{c_s\kappa}{\pi G\Sigma},$$ with $\kappa=\Omega$ for a Keplerian disc. Axisymmetric disturbances become gravitationally unstable when $Q$ drops below order unity.

Instability doesn’t automatically mean fragmentation. A disc with $Q\sim1$ can develop spiral structure and use gravitational torques to transport angular momentum without breaking into separate objects. To fragment it also has to shed the heat generated by compression and shocks quickly enough, which is why disc thermodynamics matters as much as the value of $Q$. Young massive discs are therefore the most vulnerable, and if fragmentation does occur it can make stellar or brown-dwarf companions. Forming giant planets this way is possible in some parts of parameter space, especially at large radii, but it isn’t the inevitable outcome of $Q<1$.[mfn]A common way to express the cooling requirement is through $t_{\rm cool}\Omega$. Fragmentation is favoured when the cooling time is only a few orbital times or less, although the exact critical value is not universal and depends on the equation of state, irradiation and numerical setup.[/mfn] The other half of the angular momentum story sits above and below the disc, where young stars launch jets and winds. The fastest collimated jets reach hundreds of kilometres per second and extend for parsecs, and their connection to accretion isn't an accident, because a magnetic field can extract angular momentum from rotating disc material and put it into an outflow. The cleanest toy model is Blandford and Payne's bead on a wire: take a field line anchored to the disc at radius $r_0$, pretend it's rigid, and force a parcel of gas to slide along it while rotating with the footpoint. In the rotating frame the effective potential is $$\Phi_{\rm eff}(r,z)=-\frac{GM}{\sqrt{r^{2}+z^{2}}}-\frac12\Omega_0^{2}r^{2},$$ and expanding it along the field line near the disc shows that if the line leans more than $30^\circ$ from the rotation axis the potential falls as you move outward along it. Gas can then accelerate away from the surface without thermal pressure pushing it over a barrier. The wind doesn't get its energy for free: the energy and angular momentum come from the rotating accretion flow, and the magnetic field is the lever that transfers them. Further out the gas passes the Alfvén point and the field can no longer hold it in exact corotation, but by then the outflow has taken angular momentum with it, which is what the accreting material needed.

discfootpoint, r₀parallel to rotation axis30° thresholdfield line, θ = 55°30°no launchmagnetocentrifugal launch
The magnetocentrifugal launching criterion in the idealised bead-on-a-wire picture. The angle is measured from the direction parallel to the rotation axis. A field line inclined by more than $30^\circ$ provides a downhill path in the effective potential, allowing gas to accelerate away from the disc. The outflow draws its energy and angular momentum from the rotating accretion flow.

For most of the history of astronomy almost all of this happened where we couldn’t see it, because a young protostar sits inside an optically thick dusty envelope and visible light tells you very little. Infrared and submillimetre radiation do much better, since dust absorbs short-wavelength radiation and reradiates the energy at longer wavelengths. That led to classifying young stellar objects by the shape of their spectral energy distributions, using the infrared spectral index $$\alpha=\frac{d\log(\lambda F_\lambda)}{d\log\lambda}$$ measured across the near and mid infrared.

A Class I source has a rising infrared SED, roughly $\alpha>0.3$. Flat-spectrum sources sit in the range $-0.3\lesssim\alpha\lesssim0.3$, Class II sources have $-1.6\lesssim\alpha\lesssim-0.3$, and Class III sources have still more negative slopes. Class 0 is different. It was added later and isn’t properly defined by the same infrared slope, because a genuine Class 0 source can be almost invisible at the wavelengths used to measure $\alpha$. It’s identified instead from its very cold SED and from the fact that a large fraction of its luminosity emerges in the submillimetre, with most of the system’s mass still in the envelope.

The sequence roughly tracks evolution. Class 0 and I objects are embedded protostars, by Class II the envelope is mostly gone and a circumstellar disc produces the infrared excess, and by Class III the excess is weak and most of the primordial disc has disappeared. Class and age aren’t the same thing, though. Turn a disc edge-on and it can look far more embedded than it is, and extra foreground extinction changes the slope again. The classes are useful observational labels, and geometry can make one evolutionary stage masquerade as another.

Class 0envelope dominatedcold, submm SEDClass Iα > 0.3embeddedFlat−0.3 < α < 0.3Class II−1.6 < α < −0.3discClass IIIα < −1.6disc mostly gone
The observational classes of young stellar objects. The sequence broadly follows the removal of the envelope and then the disc, but it isn’t a perfect clock. Class 0 is defined using the cold, envelope-dominated SED rather than the same near-to-mid-infrared slope used for Classes I, Flat, II and III. Inclination and extinction can also move a source between apparent classes.

Once the main accretion phase ends, a low-mass object still isn’t on the main sequence. It’s larger than an ordinary star of the same mass and is still contracting, with gravitational energy of order $E_{\rm grav}\sim GM^2/R$ available, so dividing by the luminosity gives the Kelvin-Helmholtz time $$t_{\rm KH}\sim\frac{GM^{2}}{RL}\sim3\times10^{7}\ \mathrm{yr}$$ for present-day solar values.

A young solar-mass star first moves down a Hayashi track in the Hertzsprung-Russell diagram, staying at roughly constant effective temperature while its luminosity falls, and it’s mostly convective. As a radiative core develops it moves onto a Henyey track and heads toward the main sequence. Deuterium burning starts earlier, at central temperatures of order $10^6$ K, and acts as a temporary thermostat, while sustained hydrogen burning through the proton-proton chain needs of order $10^7$ K. Once nuclear burning replaces the energy radiated from the surface, contraction largely stops. Below roughly $0.075$ to $0.08\,M_\odot$, depending slightly on composition, the centre never gets hot enough, electron degeneracy pressure becomes important first, and the result is a brown dwarf that spends the rest of its life cooling.

So far this has been about how one collapsing region makes one central object. The harder question is why it makes the masses it does. Count newly formed stars by mass and you get the initial mass function, which above roughly a solar mass approaches the Salpeter form $dN/dM\propto M^{-2.35}$, or equivalently $dN/d\log M\propto M^{-1.35}$. Below a solar mass the distribution flattens and turns over, with a characteristic mass of a few tenths of a solar mass.

The broad shape is remarkably similar across many nearby star-forming regions, which is one of the strangest facts in the subject, since clouds vary in density, temperature, turbulence, radiation field and chemical composition and the stellar mass distribution doesn’t respond as violently as you’d expect. That doesn’t mean universality has been proved everywhere. There are continuing claims of variation, especially in extreme environments and in unresolved old stellar populations. The safer statement is that large variations are surprisingly hard to produce in the environments where we can measure individual young stars well.

dN / d log M0.010.1110100characteristic mass ≈ 0.2–0.3 M⊙Salpeter taildN / d log M ∝ M⁻¹·³⁵0.075 M⊙brown dwarfsstellar mass, M / M⊙
A schematic initial mass function plotted as number per logarithmic mass interval. The distribution turns over at a characteristic mass of a few tenths of a solar mass. Below about $0.075\,M_\odot$ lie the brown dwarfs, while at high masses the IMF approaches the Salpeter form, $dN/d\log M\propto M^{-1.35}$. Low-mass stars dominate the number count, while the rare massive stars dominate the luminosity, ionising radiation and eventual supernova feedback.

Why the IMF has this shape is still not settled. For the turnover, one line of argument has become more specific in recent years. The Jeans mass by itself isn’t enough, because it changes continuously as density and temperature change. But the thermodynamics of collapse contains a point where the gas stops behaving nearly isothermally and begins to heat strongly as it becomes opaque, and that change produces the first hydrostatic core and introduces a characteristic mass tied to dust opacity and molecular hydrogen physics. Since those microphysical scales don’t change wildly across ordinary Galactic star-forming environments, they may help explain why the peak is fairly stable.8This is not the only proposed explanation for the IMF peak. Earlier arguments often tied the characteristic mass to the density at which gas and dust thermally couple, while modern calculations increasingly emphasise the thermodynamic transition associated with opacity and the first hydrostatic core. The broader point is the same: some piece of non-scale-free microphysics has to enter if the IMF is going to acquire a preferred mass.

The high-mass power law probably has a different origin. Gravity and turbulence are nearly scale-free over a large range and both naturally generate power-law structures, while turbulent fragmentation, continuing accretion, interactions between neighbouring collapsing objects and stellar feedback can all change the final mass spectrum. Current simulations produce IMFs that look impressively realistic, but there’s still no short derivation that starts from cloud properties and predicts the whole function.

That becomes clearer once massive stars enter. A massive protostar doesn’t finish accreting and then switch on. Its Kelvin-Helmholtz time is so short that it can begin sustained hydrogen burning while material is still falling onto it, so strong radiation, ionisation and outflows appear while the star is still being assembled. That used to look like a serious obstacle to making massive stars at all, since radiation pressure might simply blow away the infalling gas. The answer is that the accretion flow isn’t spherical: dense disc-fed accretion can continue while radiation and ionised gas escape preferentially through lower-density directions. Observations and simulations now suggest more continuity between low-mass and high-mass star formation than older pictures implied, though massive stars live in denser, more strongly accreting and more strongly clustered environments.

The newly formed stars then start destroying the cloud that made them. Take an O star producing roughly $Q\sim10^{49}$ ionising photons per second. In uniform hydrogen gas, ionisations balance recombinations at the Strömgren radius $$R_S=\left(\frac{3Q}{4\pi n^{2}\alpha_B}\right)^{1/3}\sim3\ \mathrm{pc}$$ for $n=100\ \mathrm{cm^{-3}}$ and a standard case-B recombination coefficient. The gas inside is heated to roughly $10^4$ K, far above the temperature of the molecular material around it, so the H II region expands and drives a shock into the surrounding cloud. Radiation pressure, stellar winds and protostellar outflows add more momentum, and a few million years later the most massive stars begin to explode.

Feedback doesn’t explain the entire low value of $\epsilon_{\rm ff}$ by itself. Turbulence, magnetic fields, cloud geometry and the continual assembly and dispersal of gas matter before the first O star appears. What feedback does effectively is put a clock on the process, because once enough stars have formed, especially massive ones, the cloud begins losing the reservoir from which further stars could have formed. That helps resolve something which otherwise sounds contradictory. Star formation is slow measured against the molecular gas available, while individual dense structures can still collapse quickly. The Galaxy doesn’t need every dense core to collapse at one percent of free fall. It needs only a small fraction of the molecular gas to be in a rapidly collapsing state at any one time, with feedback and dynamics recycling or dispersing the rest.

The timescales then form a rough hierarchy. Giant molecular clouds live for something like several to a few tens of millions of years, depending on environment and on how a cloud is defined. Dense prestellar cores evolve over hundreds of thousands of years. The deeply embedded Class 0 and I phases together last of order half a million years for nearby low-mass populations. Protoplanetary discs commonly survive a few million years with a broad spread. A solar-mass star then takes tens of millions of years to settle onto the main sequence, while a $20\,M_\odot$ object reaches it while still accreting and starts ionising and disturbing its birth environment before the formation process around it has finished.

Put the chain back together and the original twenty-four orders of magnitude stop looking like one collapse. The gas first has to become cold and concentrated enough for self-gravity to beat pressure, which is the Jeans condition. Then the obvious prediction fails, because clouds don’t turn themselves into stars in one free-fall time, and turbulence, magnetic fields, cloud structure and feedback enter. Once a dense core collapses, radiative transfer changes the equation of state and produces the first and second hydrostatic cores. Rotation creates a disc, the disc loses angular momentum through gravitational torques, magnetic stresses and winds, and the forming star launches outflows. Massive stars eventually ionise and disrupt the larger cloud. Somewhere inside all of that, the process also produces nearly the same broad distribution of stellar masses again and again.

Some parts of that chain are on firm ground. The Jeans dispersion relation follows directly from the fluid equations, the free-fall time follows from Newtonian gravity, the need to transport angular momentum is unavoidable, and the first and second collapse follow from the thermodynamics of molecular hydrogen and radiative transfer. Other parts are much less settled. Star formation is inefficient on molecular-cloud scales, and there’s no agreed partition of responsibility between turbulence, magnetic fields, cloud evolution and feedback. Discs move angular momentum, and the relative roles of magnetised winds, MRI, hydrodynamic instabilities and gravitational torques vary from one part of a disc to another. The IMF is measured very well, and its full shape still can’t be derived from first principles.

References and Footnotes

  • 1
    There is a genuine inconsistency hidden here, traditionally called the Jeans swindle. An infinite medium with constant density cannot simultaneously satisfy $\nabla^2\Phi_0=4\pi G\rho_0$ and have no background gravitational acceleration. The usual calculation ignores the gravity of the uniform background and keeps the gravity of the perturbation. More careful treatments using finite systems or expanding backgrounds recover essentially the same instability criterion, so the answer survives even though the original setup is not mathematically self-consistent. ↩︎
  • 2
    Image credit: ESA/Herschel/PACS, SPIRE/Gould Belt survey Key Programme. The characteristic $0.1$ pc width was emphasised in the Herschel Gould Belt work, but later studies have questioned how universal it really is. The safe statement is that filaments are common and closely connected to dense core formation. The exact distribution of their widths remains an active observational problem. ↩︎
  • 3
    The low value of $\epsilon_{\rm ff}$ is one of the central empirical constraints on star formation theory. Modern cloud-lifecycle work also makes clear that individual clouds need not form stars at a perfectly constant rate. A cloud can evolve through quiet and active phases, so an average value near one percent does not mean every cloud converts exactly one percent of its mass every free-fall time. ↩︎
  • 4
    Modern radiation-hydrodynamic calculations give a detailed picture of the first and second core stages, but the observational signature of a genuine first hydrostatic core is not unique. Several good candidates exist. It is safer to call them candidates than to say the phase has been securely observed as a class of objects. ↩︎
  • 5
    Image credit: NASA, ESA, CSA, STScI; image processing by Joseph DePasquale, Anton M. Koekemoer and Alyssa Pagan. JWST NIRCam image released in October 2022. ↩︎
  • 6
    The history is worth noting. Velikhov and Chandrasekhar studied the instability decades before Balbus and Hawley recognised its importance for accretion discs in 1991, after which MRI became almost synonymous with angular momentum transport in disc theory. Non-ideal MHD changed that. Current work increasingly treats magnetic winds as a major and possibly dominant route for angular momentum loss in protoplanetary discs, with the balance between winds, MRI, hydrodynamic instabilities and gravitational torques depending on radius and evolutionary stage. ↩︎
  • 7
    Image credit: ALMA (ESO/NAOJ/NRAO), released in 2014 from ALMA’s long-baseline campaign. ↩︎
  • 8
    This is not the only proposed explanation for the IMF peak. Earlier arguments often tied the characteristic mass to the density at which gas and dust thermally couple, while modern calculations increasingly emphasise the thermodynamic transition associated with opacity and the first hydrostatic core. The broader point is the same: some piece of non-scale-free microphysics has to enter if the IMF is going to acquire a preferred mass. ↩︎
Posted in Expository | Tagged , , , , , , , , | Leave a comment

Why $R_{\mu\nu}=0$ does not imply $R_{\mu\nu\rho\sigma}=0$

Almost everyone makes this mistake once, and it’s a good one to make. You learn the Einstein field equations. You learn that in vacuum, with no matter and no energy and no cosmological constant, they collapse down to something very simple: $$R_{\mu\nu} = 0.$$ So you think: $R_{\mu\nu}$ is the curvature, it’s zero, so there’s no curvature out there, spacetime is flat, nothing is happening. Then somebody points out that the entire outside of a black hole is vacuum. Every single point out there satisfies $R_{\mu\nu}=0$ exactly. And things fall into it.

So the reasoning broke somewhere, and I want to find out where, because once you see which step failed you’ll understand what the Ricci tensor actually measures, what the rest of the curvature is doing, and why gravity has waves at all. Short version: $R_{\mu\nu}$ isn’t the curvature. It’s a contraction of the curvature, and contracting throws information away.

Before any tensors show up, though, I want to convince you that you already believe this and have believed it since your first mechanics course. Newtonian gravity obeys Poisson’s equation, $$\nabla^2 \Phi = 4\pi G \rho.$$ Where there’s no matter, $\rho = 0$, and you get Laplace’s equation, $\nabla^2\Phi = 0$. Now, does anybody think Laplace’s equation means the gravitational field is zero? No. The potential outside a spherical mass is $\Phi = -GM/r$, you can check for yourself that $\nabla^2(-GM/r) = 0$ everywhere except the origin, and $\Phi$ is obviously not zero, its gradient isn’t zero, and things fall. The air above your head is vacuum. Laplace’s equation holds up there. You’re still stuck to the floor.

r Φr = 0Φ → −∞Φ = −GM/rVacuum region: ρ = 0, so ∇²Φ = 0.But Φ ≠ 0, ∇Φ ≠ 0, and its tidal Hessian is nonzero.
The same puzzle in Newtonian language. Outside the source, Poisson’s equation becomes Laplace’s equation. That sets only the trace of the Hessian $\partial_i\partial_j\Phi$ to zero; it does not set the potential, the field, or the trace-free tidal Hessian to zero.

That last bit is the whole answer already, so let me slow down on it. The second derivatives of the potential form a symmetric matrix, the Hessian $\partial_i \partial_j \Phi$, and in three dimensions that has six independent entries. Laplace’s equation sets exactly one number to zero: the trace, $\sum_i \partial_i\partial_i \Phi$. Five entries are left completely alone. Those five are the tidal field, the thing that stretches you along the line toward the Earth and squeezes you sideways, and they’re what gravity does in empty space.

General relativity does the same thing with better bookkeeping. The full curvature of spacetime is the Riemann tensor $R^\rho{}{\sigma \mu \nu}$, and the Ricci tensor isn’t some separate object, it’s what you get by contracting Riemann on a pair of indices: $$R{\mu\nu} = R^\alpha{}{\mu\alpha\nu}.$$ You’re summing over one direction. That’s an averaging operation, and averaging loses information. If I tell you a list of numbers averages to zero, you don’t conclude the numbers were all zero, because ${+1, -1}$ averages to zero too. So $R{\mu\nu} = 0$ says that certain averages of the curvature vanish. It doesn’t say the curvature vanishes.

You can make that exact by counting. Riemann has a lot of symmetry: antisymmetric in the first two indices, antisymmetric in the last two, and symmetric when you swap the two pairs. In four dimensions an antisymmetric pair takes $\binom{4}{2}=6$ values, so the pair symmetry makes Riemann a symmetric $6\times 6$ matrix with $\tfrac{6\cdot 7}{2}=21$ entries, and the first Bianchi identity $R_{\mu[\nu\rho\sigma]}=0$ knocks off one more. So the Riemann tensor of a four dimensional spacetime has $$20 \ \text{independent components.}$$ The Ricci tensor is symmetric, so it has $\tfrac{4\cdot 5}{2}=10$. Setting $R_{\mu\nu}=0$ puts ten conditions on twenty numbers.1Worth doing this count yourself once. The first Bianchi identity $R_{\mu[\nu\rho\sigma]}=0$ looks like it should impose a lot of conditions, but nearly all of them already follow from the pair symmetries. The only new content is the totally antisymmetric part $R_{[\mu\nu\rho\sigma]}=0$, and in four dimensions there’s exactly $\binom{4}{4}=1$ way to pick four antisymmetrised indices, so exactly one new constraint. In general the answer is $\frac{n^2(n^2-1)}{12}$, which gives 1 for $n=2$, 6 for $n=3$, 20 for $n=4$, and 50 for $n=5$. Ten are left over. Whatever those ten are, they’re the curvature that empty space is allowed to have.

Riemann curvature in four dimensions: 20 independent components Ricci-determined part10 curvature componentsWeyl part10 curvature components impose Rμν = 0Ricci-determined partvanishesWeyl may remain nonzero10 componentsAlgebraically free of the Ricci condition; its propagation is constrained by the Bianchi identities.
At a point, the vacuum condition removes the ten Ricci-determined components of the four-dimensional Riemann tensor while permitting ten independent Weyl components to remain. “Permitting” does not mean unconstrained dynamics: the Bianchi identities govern how the Weyl curvature propagates through vacuum.

The ten leftovers have a name. There’s exactly one way to split the Riemann tensor into a piece built out of Ricci and a remainder that’s completely traceless, and the remainder is the Weyl tensor $C_{\mu\nu\rho\sigma}$. In $n$ dimensions the split reads $$R_{\mu\nu\rho\sigma} = C_{\mu\nu\rho\sigma} + \frac{2}{n-2}\Bigl(g_{\mu[\rho}R_{\sigma]\nu} – g_{\nu[\rho}R_{\sigma]\mu}\Bigr) – \frac{2R}{(n-1)(n-2)}\, g_{\mu[\rho}g_{\sigma]\nu},$$ where the square brackets mean antisymmetrise. Every trace of $C$ vanishes by construction, $C^\alpha{}{\mu\alpha\nu}=0$, and that’s exactly why the field equations can’t see it. So in vacuum, where $R{\mu\nu}=0$ and therefore $R=0$ as well, the whole thing collapses to $$R_{\mu\nu\rho\sigma} = C_{\mu\nu\rho\sigma}.$$ The curvature of empty space is pure Weyl. Not mostly Weyl, not approximately Weyl. Exactly and only Weyl.2The Weyl tensor has $\frac{n(n+1)(n+2)(n-3)}{12}$ independent components in $n$ dimensions. It’s also conformally invariant: rescale the metric by any positive function, $g_{\mu\nu} \to \Omega^2 g_{\mu\nu}$, and the Weyl tensor with one index raised doesn’t change. That’s why it sometimes gets called the conformal curvature tensor, and why a spacetime is conformally flat exactly when its Weyl tensor vanishes.
Now stare at that component formula for a second, because there’s something odd hiding in it. The factor $(n-3)$ means the Weyl tensor has zero components whenever $n \le 3$. Not zero for some particular metric, zero identically, always, for every three dimensional geometry there is. Which means the Riemann tensor in three dimensions is completely fixed by the Ricci tensor. Which means that in three dimensions, $R_{\mu\nu}=0$ really does force spacetime to be flat.

dimension nRiemannRicci data*Weyl 211*03660 4201010 5501535 *In 2D, “1” means one independent curvature degree of freedom carried by Ricci:Rμν = (R/2)gμν. An arbitrary symmetric 2-tensor would have three components.
Independent local curvature data by dimension. In two and three dimensions the Riemann tensor contains no independent Weyl part, so Ricci-flat implies flat. Four dimensions are the first in which curvature can remain after all Ricci curvature has vanished. The 2D Ricci entry counts curvature information, not the components of an arbitrary symmetric rank-two tensor.

This still surprises me. The wrong intuition isn’t wrong because anyone reasoned badly. It’s wrong because it’s reasoning about a universe with one dimension too few. In three dimensional gravity there’s no curvature outside a source, no light bending around a star, no structure in a black hole exterior, and no gravitational radiation, because there’s simply nowhere for that information to sit. The theory is topological. Our fourth dimension is what opens up ten free slots in the curvature and turns gravity into a field with a life of its own.3Three dimensional gravity isn’t boring, and the qualification matters: with a negative cosmological constant you get the BTZ black hole, with a horizon and a temperature and an entropy. But that isn’t a counterexample. BTZ is locally identical to anti de Sitter space everywhere, with constant curvature set entirely by $\Lambda$; what makes it a black hole is a global identification of points, which is topology rather than local geometry. There are still no local gravitational degrees of freedom in three dimensions and no gravitational waves.

So Weyl is what survives. The question that matters is what it does, and here the cleanest statement in the whole subject is waiting. Take a small ball of test particles, freely falling, starting at rest relative to each other. Coffee grounds released in a spacecraft will do fine. Watch the ball. Its volume is controlled by the expansion $\theta = \nabla_\mu u^\mu$ of the congruence of worldlines, which obeys the Raychaudhuri equation $$\frac{d\theta}{d\tau} = -\frac{1}{3}\theta^2 – \sigma_{\mu\nu}\sigma^{\mu\nu} + \omega_{\mu\nu}\omega^{\mu\nu} – R_{\mu\nu}u^\mu u^\nu,$$ with $\sigma_{\mu\nu}$ the shear and $\omega_{\mu\nu}$ the vorticity. Start the ball at rest with no shear and no rotation, so at the first instant $\theta$, $\sigma$ and $\omega$ are all zero, and every term on the right dies except the last one. You’re left with $$\left.\frac{\ddot V}{V}\right|{\tau=0} = -R{\mu\nu}u^\mu u^\nu.$$ The Ricci tensor is the thing, and the only thing, that changes the volume of a ball of freely falling matter.4Getting from $\theta$ to volume: $\theta = \dot V/V$, so $\dot\theta = \ddot V/V – (\dot V/V)^2$, and at the initial instant where $\dot V=0$ that’s just $\ddot V / V$. Feed in Einstein’s equation and the geometric statement becomes a statement about matter: for a perfect fluid you get $\ddot V/V = -4\pi G(\rho + 3p/c^2)$, which is a very direct way to see why pressure gravitates in general relativity, and why no equation of state stiff enough can save a star from collapsing.

Now put the ball in vacuum. There $R_{\mu\nu}u^\mu u^\nu = 0$ identically, so the volume doesn’t change at all. But something obviously does happen to coffee grounds falling toward the Earth, because the ones nearer the Earth fall faster and the ones off to the sides fall along slightly converging lines. The ball deforms. It just deforms at fixed volume. In vacuum the geodesic deviation equation reads $$\frac{D^2 \xi^\mu}{d\tau^2} = -R^\mu{}{\alpha\nu\beta},u^\alpha \xi^\nu u^\beta = -E^\mu{}\nu, \xi^\nu, \qquad E_{\mu\nu} \equiv C_{\mu\alpha\nu\beta}u^\alpha u^\beta,$$ where $E_{\mu\nu}$ is called the electric part of the Weyl tensor. It’s symmetric, and since every trace of Weyl vanishes it’s traceless too: $E^\mu{}\mu = 0$. A traceless matrix of relative accelerations is exactly the statement that volume is preserved while shape isn’t. That’s the division of labour, and it’s a theorem, not a slogan. Ricci changes volume. Weyl changes shape.5Weyl also has a magnetic part, $B{\mu\nu} = \tfrac{1}{2}\epsilon_{\mu\alpha\beta\gamma}C^{\alpha\beta}{}_{\nu\delta}u^\gamma u^\delta$, and the two together account for all ten components, five each, both symmetric and traceless. The electric and magnetic language isn’t a loose analogy either: $E$ and $B$ obey equations closely parallel to Maxwell’s, and it’s the two of them oscillating into each other that makes a gravitational wave. A purely electric Weyl field is just the static tidal field of a mass. You need the magnetic part to get radiation.

instantaneous response of an initially round, comoving ball Ricci focusing examplevacuum: pure Weyl tidal field Ricci controls the traceso V̈/V can be nonzero initially Weyl contributes traceless distortionV̈(0) = 0, while the shape changesmatching transverse squeeze also occurs into/out of the page
For an initially round, comoving ball of freely falling particles, Ricci curvature controls the trace of the relative-acceleration tensor and therefore the initial volume acceleration. In vacuum that trace vanishes, while Weyl curvature can immediately produce shear. Thus $\ddot V(0)=0$; the figure should not be read as claiming exact volume conservation at all later times, because Weyl-generated shear can subsequently enter the Raychaudhuri equation.

You can check the whole story with nothing but calculus, and I’d recommend actually doing it, because it makes the abstraction concrete. Take $\Phi = -GM/r$ and compute the Hessian: $$\partial_i \partial_j \Phi = GM\left(\frac{\delta_{ij}}{r^3} – \frac{3 x_i x_j}{r^5}\right).$$ The trace is $GM(3/r^3 – 3r^2/r^5) = 0$, which is Laplace’s equation turning up as promised. The eigenvalues are easy to read off: along the radial direction you get $-2GM/r^3$, and in each of the two transverse directions you get $+GM/r^3$. Since the relative acceleration of nearby particles is $\ddot\xi^i = -\partial_i\partial_j\Phi\,\xi^j$, the ball gets pulled apart radially at a rate $2GM/r^3$ and squeezed inward in each transverse direction at $GM/r^3$. One stretch, two squeezes, and because $2 = 1+1$ the volume stays put. That’s the traceless condition, sitting there in numbers you can work out in a minute.

M r +2GM/r³−GM/r³second transverse eigenvalue, into/out of page: −GM/r³trace of relative acceleration: +2GM/r³ − GM/r³ − GM/r³ = 0therefore the initial volume acceleration vanishes
The Newtonian tidal field of a point mass. The relative-acceleration eigenvalues are $+2GM/r^3$ radially and $-GM/r^3$ in each of the two transverse directions. Their trace is zero, so an initially comoving small ball has zero initial volume acceleration even though it immediately begins to distort.

Some numbers help here. Near the Earth’s surface the tidal rate is $2GM_\oplus/R_\oplus^3 \approx 3.1\times 10^{-6}\ \mathrm{s^{-2}}$, so across the two metres of a standing person the difference in acceleration between head and feet is about $6\times 10^{-6}\ \mathrm{m\,s^{-2}}$, which is why you’ve never noticed it. Go to a neutron star of $1.4$ solar masses and stand $100$ kilometres away and the same expression gives $3.7\times 10^{5}\ \mathrm{s^{-2}}$, so the head to foot difference is around seventy five thousand times Earth gravity and you get pulled into a thread. Both of those places are perfect vacuum. Both have $R_{\mu\nu}=0$ exactly. The difference between an effect you can’t feel and one that kills you is entirely a difference in the Weyl tensor, which the field equations never mention.

If you want a proof rather than an argument that Schwarzschild isn’t flat, compute a curvature invariant, because scalars can’t be changed by any coordinate transformation. The Kretschmann scalar is $$K = R_{\mu\nu\rho\sigma}R^{\mu\nu\rho\sigma} = \frac{48G^2M^2}{c^4 r^6}.$$ That’s nonzero at every finite radius. If spacetime were flat then $R_{\mu\nu\rho\sigma}=0$ and $K$ would have to vanish, and it doesn’t. There’s no clever choice of coordinates that turns a nonzero scalar into a zero one. So the outside of a star is curved, curved despite being empty, and since Riemann and Weyl coincide there we can also write $K = C_{\mu\nu\rho\sigma}C^{\mu\nu\rho\sigma}$.6The same scalar settles the other classic Schwarzschild confusion. At the horizon $r = r_S$ the metric components go haywire, but $K = 48G^2M^2/(c^4 r_S^6)$ is perfectly finite there, which tells you the horizon is a problem with the coordinates and not with the geometry. As $r \to 0$ the scalar blows up, and that singularity is real and no choice of coordinates removes it.

The most extreme version of all this is gravitational radiation. A gravitational wave crossing interstellar space solves $R_{\mu\nu}=0$ at every point along its path. No matter anywhere in it, nothing on the right hand side of the field equations, nothing at all in the naive sense. It’s pure Weyl curvature moving at the speed of light, made entirely out of the components the vacuum equations decline to constrain. And when one arrives it does exactly what we derived above: changes the shape of things while leaving their volume alone, stretching along one axis and squeezing along the perpendicular one. That’s precisely what LIGO measures, a difference in the lengths of two perpendicular arms, and in September 2015 that difference was about $4\times 10^{-18}$ metres over a four kilometre baseline, roughly a thousandth of the width of a proton. Empty space, curved, and curved strongly enough to detect.

So the corrected statement is this. $R_{\mu\nu}=0$ doesn’t say spacetime is flat. It says spacetime is Ricci flat, which is a real and restrictive condition, but one that leaves half the curvature untouched in four dimensions. The half it kills is the half that responds directly to local matter and changes the volumes of falling objects. The half it leaves is the tidal, shape distorting, freely propagating half, and that half isn’t sourced locally at all. It’s set by matter somewhere else and carried through the vacuum by the Bianchi identities, in exactly the way that $\nabla^2\Phi=0$ in the air above your head is perfectly compatible with the Earth underneath pulling on you.7The mechanism by which distant matter fixes Weyl in a vacuum region is the second Bianchi identity, $\nabla_{[\lambda}R_{\mu\nu]\rho\sigma}=0$. Contract it in vacuum and you get $\nabla^\mu C_{\mu\nu\rho\sigma}=0$, a wave equation for the Weyl tensor where the matter enters only through boundary conditions. That’s the precise sense in which Weyl curvature is the free gravitational field, and the parallel with the source free Maxwell equations $\nabla^\mu F_{\mu\nu}=0$ is very close.

One last thing, for anyone drifting toward pure mathematics. Ricci flat manifolds aren’t something physicists invented to describe black holes, they’re a central object in Riemannian geometry, and $R_{\mu\nu}=0$ is a hard nonlinear equation with a big space of solutions. Calabi conjectured and Yau proved that every compact Kähler manifold with vanishing first Chern class carries a Ricci flat metric, and the resulting Calabi-Yau manifolds are highly curved things that happen to have no Ricci curvature anywhere. The K3 surface is a four dimensional example: compact, Ricci flat, and very much not flat. When a geometer says Ricci flat they never mean flat, and it wouldn’t occur to them that anyone might. It only occurs to us because we met the Ricci tensor for the first time in an equation with the word vacuum sitting next to it.

There’s a general lesson in here somewhere, and it’s the thing I’d want you to keep. When a theory hands you an equation of the form “some contraction of the interesting object is zero,” the first question isn’t what the equation forbids rather it’s what it leaves alone. Count the components. Find the part the equation never touches. In general relativity that part is the Weyl tensor, it has ten components, it exists only because we live in four dimensions instead of three, and it’s where black holes and tides and gravitational waves all live.

References and Footnotes

  • 1
    Worth doing this count yourself once. The first Bianchi identity $R_{\mu[\nu\rho\sigma]}=0$ looks like it should impose a lot of conditions, but nearly all of them already follow from the pair symmetries. The only new content is the totally antisymmetric part $R_{[\mu\nu\rho\sigma]}=0$, and in four dimensions there’s exactly $\binom{4}{4}=1$ way to pick four antisymmetrised indices, so exactly one new constraint. In general the answer is $\frac{n^2(n^2-1)}{12}$, which gives 1 for $n=2$, 6 for $n=3$, 20 for $n=4$, and 50 for $n=5$. ↩︎
  • 2
    The Weyl tensor has $\frac{n(n+1)(n+2)(n-3)}{12}$ independent components in $n$ dimensions. It’s also conformally invariant: rescale the metric by any positive function, $g_{\mu\nu} \to \Omega^2 g_{\mu\nu}$, and the Weyl tensor with one index raised doesn’t change. That’s why it sometimes gets called the conformal curvature tensor, and why a spacetime is conformally flat exactly when its Weyl tensor vanishes. ↩︎
  • 3
    Three dimensional gravity isn’t boring, and the qualification matters: with a negative cosmological constant you get the BTZ black hole, with a horizon and a temperature and an entropy. But that isn’t a counterexample. BTZ is locally identical to anti de Sitter space everywhere, with constant curvature set entirely by $\Lambda$; what makes it a black hole is a global identification of points, which is topology rather than local geometry. There are still no local gravitational degrees of freedom in three dimensions and no gravitational waves. ↩︎
  • 4
    Getting from $\theta$ to volume: $\theta = \dot V/V$, so $\dot\theta = \ddot V/V – (\dot V/V)^2$, and at the initial instant where $\dot V=0$ that’s just $\ddot V / V$. Feed in Einstein’s equation and the geometric statement becomes a statement about matter: for a perfect fluid you get $\ddot V/V = -4\pi G(\rho + 3p/c^2)$, which is a very direct way to see why pressure gravitates in general relativity, and why no equation of state stiff enough can save a star from collapsing. ↩︎
  • 5
    Weyl also has a magnetic part, $B{\mu\nu} = \tfrac{1}{2}\epsilon_{\mu\alpha\beta\gamma}C^{\alpha\beta}{}_{\nu\delta}u^\gamma u^\delta$, and the two together account for all ten components, five each, both symmetric and traceless. The electric and magnetic language isn’t a loose analogy either: $E$ and $B$ obey equations closely parallel to Maxwell’s, and it’s the two of them oscillating into each other that makes a gravitational wave. A purely electric Weyl field is just the static tidal field of a mass. You need the magnetic part to get radiation. ↩︎
  • 6
    The same scalar settles the other classic Schwarzschild confusion. At the horizon $r = r_S$ the metric components go haywire, but $K = 48G^2M^2/(c^4 r_S^6)$ is perfectly finite there, which tells you the horizon is a problem with the coordinates and not with the geometry. As $r \to 0$ the scalar blows up, and that singularity is real and no choice of coordinates removes it. ↩︎
  • 7
    The mechanism by which distant matter fixes Weyl in a vacuum region is the second Bianchi identity, $\nabla_{[\lambda}R_{\mu\nu]\rho\sigma}=0$. Contract it in vacuum and you get $\nabla^\mu C_{\mu\nu\rho\sigma}=0$, a wave equation for the Weyl tensor where the matter enters only through boundary conditions. That’s the precise sense in which Weyl curvature is the free gravitational field, and the parallel with the source free Maxwell equations $\nabla^\mu F_{\mu\nu}=0$ is very close. ↩︎
Posted in Expository, Notes | Tagged , , , , , , , , , | Leave a comment

Derivation of the Cosmic-Ray Diffusion Coefficient from Pitch-Angle Scattering

Cosmic rays arrive at Earth from every direction almost perfectly evenly, to about one part in a thousand. That’s a disaster if you wanted to know where they came from, because they’re charged and the Galaxy is magnetised, so whatever direction they started in was erased long ago. The usual response is to say fine, we’ll describe them statistically instead, with a diffusion coefficient. Then somebody hands you a formula, $D \approx 10^{28}\ \mathrm{cm^2\,s^{-1}}$ scaling as $E^{1/3}$, and you’re expected to nod.

I don’t want to do that. A diffusion coefficient isn’t a number you look up, it’s something you derive, and the derivation is genuinely beautiful. It starts from the Lorentz force with nothing else assumed, and along the way the resonance condition that everyone quotes falls out as a delta function rather than being asserted, and the mysterious factor of one third in $D = \lambda v / 3$ turns out to be $\int_{-1}^{1}(1-\mu^2)\,d\mu = 4/3$ divided by 8. I want to walk the whole chain, because once you’ve seen it you never have to memorise any of it again. You can reconstruct it.

So here’s the plan. Write the exact equation of motion for a particle in a mean field plus a fluctuation. Read off how the pitch angle changes. Turn that into a diffusion coefficient in pitch angle using a correlation integral. Discover the resonance. Then turn pitch angle diffusion into spatial diffusion by solving a transport equation. Then, and only then, put in numbers.

Step one. A charged particle in a magnetic field obeys

$$
\frac{d\mathbf{p}}{dt}=\frac{q}{c}\mathbf{v}\times\mathbf{B},
$$

and because the magnetic force is always perpendicular to $\mathbf{v}$, the speed $v$ never changes. Only the direction does. So the natural variable isn’t the velocity, it’s the pitch angle cosine

$$
\mu\equiv\frac{v_z}{v},
$$

measuring how much of the motion is along the mean field, which I’ll put along $\hat{z}$.

Write the field as $\mathbf{B}=B_0\hat{z}+\delta\mathbf{B}$. Since $v$ is fixed,

$$\frac{d\mu}{dt}=\frac{1}{v}\frac{dv_z}{dt}=\frac{q}{\gamma mcv}(\mathbf{v}\times\mathbf{B})_z=\frac{q}{\gamma mcv}(v_xB_y-v_yB_x).$$

Now look hard at that last expression, because it already tells you something important. It only involves the transverse components of $\mathbf{B}$. The mean field is purely along $\hat{z}$, so $B_{0x}=B_{0y}=0$ and the mean field contributes nothing at all. Only $\delta\mathbf{B}$ can change the pitch angle.

1This is worth sitting with, because it’s the reason the whole subject exists. A uniform field cannot scatter anything. It rotates the velocity around $\hat{z}$ at a constant rate and leaves the angle to $\hat{z}$ untouched forever, so a particle in a uniform field slides along its field line for eternity with perfect memory of where it was going. Every bit of transport physics that follows comes from the difference between the real field and a uniform one.
B₀ vθ a kick from δBmoves v off the coneB₀ alone keeps von this cone foreverμ = cos θ = v_z / v
The unperturbed motion carries the velocity vector around a cone of fixed opening angle, so $\mu$ is a constant of the motion. Nothing in a uniform field can change it. Only the transverse part of $\delta\mathbf B$ appears in $d\mu/dt$, which is why the entire transport problem is a problem about the fluctuating field and not about the mean one.

Let’s make that explicit. Write the gyrofrequency $\Omega = qB_0/\gamma m c$, and use the unperturbed orbit for the velocity, which is legitimate as long as the fluctuation is a small perturbation over one turn. On that orbit $v_x = v_\perp\cos\psi$ and $v_y = -v_\perp\sin\psi$ with phase $\psi = \Omega t + \psi_0$ and $v_\perp = v\sqrt{1-\mu^2}$. Substituting, $$\boxed{\ \frac{d\mu}{dt} = \frac{\Omega\sqrt{1-\mu^2}}{B_0}\Big[\cos(\Omega t + \psi_0)\,\delta B_y + \sin(\Omega t + \psi_0)\,\delta B_x\Big].\ }$$ That single line contains everything. Three features of it will each become a piece of physics later, so let me point them out now. The factor $\sqrt{1-\mu^2}$ vanishes when $\mu = \pm 1$, meaning a particle streaming exactly along the field has no transverse velocity for the fluctuation to push against, so it cannot be scattered. The fluctuation enters only through its transverse components. And the whole kick oscillates at the gyrofrequency, which is going to select which fluctuations matter.

Step two is to turn a fluctuating kick into a diffusion coefficient, and the recipe is completely general. Any quantity being pushed around randomly performs a random walk, and the accumulated change over time $t$ is $\Delta\mu = \int_0^t \dot\mu\,dt’$. That has zero mean, but its square doesn’t: $$\left\langle (\Delta\mu)^2 \right\rangle = \int_0^t!!\int_0^t \left\langle \dot\mu(t’)\,\dot\mu(t”) \right\rangle dt’\,dt”.$$ Now suppose the correlation function $C(\tau) = \langle\dot\mu(0)\dot\mu(\tau)\rangle$ depends only on the time difference and dies away after some correlation time $\tau_c$. Then for $t \gg \tau_c$ the double integral collapses: for each $t’$, the inner integral over $t”$ contributes the same finite amount $\int C(\tau)d\tau$, and there are $t$ worth of choices of $t’$. So $\langle(\Delta\mu)^2\rangle$ grows linearly in $t$, which is exactly what diffusion means, and defining $D_{\mu\mu} = \langle(\Delta\mu)^2\rangle/2t$ gives $$D_{\mu\mu} = \frac{1}{2}\int_{-\infty}^{\infty} \left\langle \dot\mu(0)\,\dot\mu(\tau) \right\rangle d\tau.$$ A transport coefficient is the area under a correlation function. This is a Green-Kubo relation, and the same structure gives you electrical conductivity, viscosity and thermal conductivity in other contexts. The physical content is that transport is set by how long a fluctuation remembers itself.2The condition $t \gg \tau_c$ is not a technicality, it’s the entire definition of diffusive behaviour. On timescales shorter than the correlation time the motion is ballistic and $\langle(\Delta\mu)^2\rangle \propto t^2$, not $t$. Everything in this post applies only after the particle has been scattered many times. That’s also why diffusion fails at the highest cosmic ray energies, where a particle crosses the Galaxy before it gets scattered even once.

τ ⟨μ̇(0) μ̇(τ)⟩ the area under this curve is 2 Dμμthe width of it is the correlation timetransport is set by how long a fluctuation remembers itself
The Green-Kubo picture. Take the quantity that’s being kicked around, correlate it with itself at a later time, and integrate. The area is the diffusion coefficient. The same relation gives conductivity, viscosity and every other transport coefficient in physics, which is a good hint that the structure is more fundamental than the application.

Step three is where the resonance appears, and I think it’s the prettiest moment in the whole derivation. Put the expression for $\dot\mu$ into the correlation function. Two things get correlated. The gyration phase gives $\langle\cos\psi(0)\cos\psi(\tau)\rangle$, which after averaging over the initial phase is $\tfrac12\cos\Omega\tau$. The field gives $\langle \delta B(z_0)\,\delta B(z_0 + v\mu\tau)\rangle$, because in time $\tau$ the particle has moved a distance $v\mu\tau$ along the field and is sampling the fluctuation at a new place. Writing that correlation in Fourier space with power spectrum $P(k)$ turns it into $\int dk\, P(k)\,e^{ikv\mu\tau}$. So the thing we have to integrate over all $\tau$ is $$\int_{-\infty}^{\infty} d\tau\ \cos(\Omega\tau)\, e^{i k v\mu \tau} = \pi\Big[\delta(kv\mu – \Omega) + \delta(kv\mu + \Omega)\Big].$$ There it is. A delta function. Not a rule of thumb, not a scaling argument, an exact statement that out of the entire turbulent spectrum the particle feels only the wavenumber satisfying $kv\mu = \Omega$, which rearranges to $$k_{\rm res} = \frac{\Omega}{v\mu} = \frac{1}{r_g \mu}, \qquad r_g \equiv \frac{v}{\Omega} = \frac{pc}{qB_0}.$$ The physical reading is exactly what you’d hope. The particle turns at rate $\Omega$ and drifts along the field at speed $v\mu$. If the wave it’s passing through has a wavelength such that it advances by one wavelength in exactly one gyroperiod, then it meets the same phase of the wave at the same phase of its own orbit every single turn, and the kicks add coherently. Any other wavelength and the phase slips, and over many turns the kicks average to nothing.

phase locked: k = k_resphase slipping: k ≠ k_resgyrationδB seenproduct every kick the same sign, they addkicks alternate sign, they cancel
Where the delta function comes from. The top curve is the particle’s transverse velocity oscillating at $\Omega$, the middle is the fluctuating field it samples as it drifts along the line, and the bottom is their product, which is what actually pushes the pitch angle. When the two are locked in phase the product never changes sign and the kicks accumulate. When they aren’t, the product oscillates about zero and the time integral kills it. That is the entire content of the resonance condition.

Carrying the delta function through the $k$ integral is now just bookkeeping, and it gives $$D_{\mu\mu} = \frac{\pi}{4}\,\Omega\,(1-\mu^2)\ \mathcal{P}(k_{\rm res}), \qquad \mathcal{P}(k) \equiv \frac{k\,P(k)}{B_0^{2}},$$ where $\mathcal P$ is the fraction of the mean field energy sitting in one logarithmic band of wavenumber around $k$, a dimensionless number.3Prefactors of order unity in this formula are convention dependent, because different books normalise the spectrum differently, use one sided or two sided $k$, and include one or both transverse components in $\delta B^2$. I’m using $\int P(k)\,dk = \delta B^2$ with $k$ running over positive values only. If you compare with a textbook and find a factor of two, that’s why, and it doesn’t change any scaling. It’s worth checking this against intuition before going on. It says the pitch angle randomises at a rate equal to the gyrofrequency multiplied by the fractional turbulent power at resonance. If the turbulence at the resonant scale were as strong as the mean field, $\mathcal P \sim 1$, the pitch angle would be scrambled in a single orbit. That’s as fast as scattering can possibly go, and we’ll come back to it.

Step four is to get from pitch angle diffusion to actual spatial transport, and this is the step almost every treatment skips. It shouldn’t, because the answer is not what you’d naively guess. Write the phase space density $f(z,\mu,t)$ and let it obey streaming along the field plus diffusion in pitch angle, $$\frac{\partial f}{\partial t} + v\mu \frac{\partial f}{\partial z} = \frac{\partial}{\partial \mu}\left( D_{\mu\mu} \frac{\partial f}{\partial \mu} \right).$$ Split $f = f_0(z,t) + f_1(z,\mu,t)$ into a nearly isotropic part and a small anisotropy. To leading order the anisotropy is set by balancing the streaming of $f_0$ against the scattering of $f_1$, $$v\mu \frac{\partial f_0}{\partial z} = \frac{\partial}{\partial\mu}\left( D_{\mu\mu}\frac{\partial f_1}{\partial\mu} \right).$$ Integrate once in $\mu$, fixing the constant by demanding nothing blows up at $\mu = \pm 1$, and you get $$\frac{\partial f_1}{\partial \mu} = -\frac{v(1-\mu^2)}{2 D_{\mu\mu}}\frac{\partial f_0}{\partial z}.$$ The flux along the field is $S = \tfrac{v}{2}\int_{-1}^{1}\mu f_1 \,d\mu$, and integrating that by parts to bring in $\partial f_1/\partial\mu$ gives $$S = -\frac{v^2}{8}\left[ \int_{-1}^{1}\frac{(1-\mu^2)^2}{D_{\mu\mu}(\mu)}\,d\mu \right] \frac{\partial f_0}{\partial z}.$$ Comparing with Fick’s law $S = -D_\parallel \,\partial f_0/\partial z$ identifies $$D_\parallel = \frac{v^2}{8}\int_{-1}^{1} \frac{(1-\mu^2)^2}{D_{\mu\mu}(\mu)}\, d\mu.$$

Stop and look at that, because it’s counterintuitive in an important way. The spatial diffusion coefficient contains $D_{\mu\mu}$ in the denominator. Fast pitch angle scattering means slow spatial transport, which makes sense, but more than that, the integral is dominated by whichever pitch angles scatter worst. It’s a bottleneck. If there’s a value of $\mu$ where $D_{\mu\mu}$ becomes small, that region controls the whole answer no matter how efficiently everything else is being scattered, in the same way that a single closed lane sets the travel time on an otherwise empty motorway. And we already know $D_{\mu\mu}$ has a factor $(1-\mu^2)$, so it vanishes at $\mu = \pm1$, and from the resonance it also carries a factor $|\mu|^{q-1}$, so it vanishes at $\mu = 0$ too. Particles moving nearly perpendicular to the field are the bottleneck, and they’re the hardest ones to scatter because the wavelength they resonate with runs off to infinity.

μ Dμμ −10+1 scattering stalls hereand controls everythingD∥ is an integral of 1/Dμμ, so the worst pitch angles set the answer
The pitch angle diffusion coefficient across the range of $\mu$. It dies at $\mu = \pm1$, where a particle aligned with the field has no transverse velocity to push on, and it dies at $\mu = 0$, where a particle moving perpendicular to the field would need an infinitely long wave to resonate with. Because $D_\parallel$ integrates the reciprocal, these minima dominate, which is why so much effort in this field goes into what happens near ninety degrees.

Now we can settle the factor of one third that gets quoted everywhere without justification. Take the simplest case, $D_{\mu\mu} = D_0 (1-\mu^2)$, which is what you get if the resonant power happens to be flat in $\mu$. Then the integral is elementary: $$D_\parallel = \frac{v^2}{8 D_0}\int_{-1}^{1}\left(1-\mu^2\right)d\mu = \frac{v^2}{8D_0}\cdot\frac{4}{3} = \frac{v^2}{6 D_0}.$$ If we define a scattering rate $\nu = 2D_0$, the rate at which $\mu$ loses memory, and a mean free path $\lambda = v/\nu$, this is $$D_\parallel = \frac{v^2}{3\nu} = \frac{1}{3}\lambda v.$$ So the one third isn’t a three dimensional geometry factor and isn’t a hand wave. It’s $\tfrac{4}{3} \div 8 \times 2$. The number came out of an integral over pitch angle, and if $D_{\mu\mu}$ had a different shape you’d get a different number.

Everything so far is exact given the model, and none of it needed a spectrum. Now we choose one. Take a power law $P(k) \propto k^{-q}$ from an outer scale $L$ downwards, normalised so the total power is $\delta B^2$, which fixes $$\mathcal{P}(k) = (q-1)\frac{\delta B^2}{B_0^2}\left( kL \right)^{1-q} \quad\Longrightarrow\quad \mathcal{P}(k_{\rm res}) = (q-1)\frac{\delta B^2}{B_0^2}\left( \frac{r_g|\mu|}{L} \right)^{q-1}.$$ Feed that into $D_{\mu\mu}$, feed that into the $D_\parallel$ integral, and the $\mu$ dependence collects into a single number: $$D_\parallel = \frac{I(q)}{2\pi (q-1)}\ v\, r_g \left(\frac{B_0}{\delta B}\right)^{2}\left(\frac{L}{r_g}\right)^{q-1}, \qquad I(q) = \int_{-1}^{1}(1-\mu^2)|\mu|^{1-q}d\mu = \frac{2}{2-q}-\frac{2}{4-q}.$$ For a Kolmogorov cascade, $q = 5/3$, that gives $I = 36/7$ and a prefactor of $1.23$. Notice that $I(q)$ blows up as $q \to 2$: the bottleneck at $\mu = 0$ becomes fatal and the theory stops giving a finite answer, which is the ninety degree problem showing up as a divergent integral rather than as a footnote.

The energy dependence now falls out with no extra input. Since $r_g \propto E$ for relativistic particles, $$D_\parallel \propto r_g \cdot r_g^{-(q-1)} = r_g^{2-q} \propto E^{2-q},$$ so Kolmogorov gives $E^{1/3}$ and a Kraichnan cascade with $q = 3/2$ gives $E^{1/2}$. Measurements of the boron to carbon ratio, which tell you how much interstellar gas cosmic rays have ploughed through and therefore how long they were confined, want an exponent between about $0.3$ and $0.6$. Both cascades sit inside that, which is honest to report and mildly annoying, because it means the data cannot yet decide which cascade the interstellar medium actually has.

Now, finally, numbers, and I want to stress that these are a check on the derivation rather than facts to carry around. Take a $10$ GeV proton in a mean field of $3\ \mu$G, turbulence as strong as the mean field so $\delta B = B_0$, an outer scale of $L = 100$ pc which is roughly the size of a supernova driven eddy, and Kolmogorov. The gyroradius from $r_g = pc/(qB_0)$ is $1.11\times10^{13}$ cm, about three quarters of an astronomical unit. Then $$D_\parallel = 1.23\, c\, r_g \left(\frac{L}{r_g}\right)^{2/3} \approx 3.8\times 10^{28}\ \mathrm{cm^2\,s^{-1}},$$ with a corresponding mean free path $\lambda = 3D/c \approx 1.2$ pc. The value inferred from cosmic ray composition is a few times $10^{28}\ \mathrm{cm^2\,s^{-1}}$. Three inputs you could look up in an afternoon, one derivation, and the answer lands on the measurement. Note also that $\lambda/r_g \approx 3\times10^5$, so the particle really does complete hundreds of thousands of clean orbits between scatterings, which is exactly the condition that made the perturbative treatment legitimate in the first place. The theory checks its own assumptions.

From here the consequences follow in a line. Confinement in a magnetised halo of half thickness $H \approx 4$ kpc takes $\tau \sim H^2/2D \approx 64$ Myr, and radioactive beryllium-10, with a $1.4$ Myr half life, independently reads tens of millions of years from completely unrelated nuclear physics. In that time the particle covers a path length $c\tau \approx 20$ Mpc while getting $4$ kpc from where it started, a ratio of about five thousand. It crossed the Local Group many times over and never left the Galaxy.

I’ve been writing $D_\parallel$ throughout, and the subscript matters. Scattering randomises motion along the field efficiently, but to move across field lines a particle has to hop between lines, which is much harder. So diffusion is anisotropic, and the diffusion coefficient is not a number but a rank two tensor $$D_{ij} = D_\perp \delta_{ij} + \left(D_\parallel – D_\perp\right) b_i b_j, \qquad \mathbf{b} = \mathbf{B}/|\mathbf{B}|.$$ Contract it with $\mathbf b$ twice and you recover $D_\parallel$; contract it with anything perpendicular and you get $D_\perp$. Which is to say the field direction picks out the principal axes and the tensor is diagonal in that frame with eigenvalues $(D_\parallel, D_\perp, D_\perp)$. Classical transport gives $D_\perp/D_\parallel \approx [1 + (\lambda/r_g)^2]^{-1}$, and with our $\lambda/r_g \approx 3\times10^5$ that would be $10^{-11}$, absurdly small.4It is absurdly small, and it’s wrong, for an instructive reason: that estimate assumes the field lines themselves are straight. They’re not. Turbulent field lines separate from each other exponentially, so a particle sliding along one is carried sideways for free without ever hopping. Including this field line random walk puts the ratio nearer $10^{-2}$. Perpendicular transport remains the least settled part of the theory and the place where numerical simulations have contributed most.

B D∥ D⊥D_ij = D⊥ δ_ij + (D∥ − D⊥) b_i b_j
Because the field breaks the isotropy of space, the diffusion coefficient has to be a tensor, and the field direction is one of its principal axes. Drawing the surface of constant diffusion gives a cigar pointing along the field line. This is the same structure as the polarizability of a crystal: a direction picked out by the medium, and a tensor that becomes diagonal when you use it.

Strong fields are where the derivation earns its keep, because you can see immediately what happens rather than having to look anything up. Every appearance of $B_0$ is inside $r_g = pc/qB_0$, so raising the field shrinks the gyroradius, shrinks the resonant wavelength, and shrinks the mean free path. But there’s a floor. Look back at $D_{\mu\mu} = \tfrac{\pi}{4}\Omega(1-\mu^2)\mathcal P$: the largest $\mathcal{P}$ can be is of order one, which puts the scattering rate at the gyrofrequency, the mean free path at $r_g$, and $$D_{\rm Bohm} = \tfrac{1}{3} r_g v.$$ You cannot scatter a particle in less than one orbit. Bohm diffusion is the tightest confinement physics allows, and it isn’t an extra assumption, it’s the ceiling on $\mathcal P$.

That floor is what makes supernova remnants work as cosmic ray sources, and the calculation is one line. In diffusive shock acceleration a particle gains energy crossing and recrossing the shock, taking a time $t_{\rm acc} \approx 20 D(E)/u_s^2$ for a shock of speed $u_s$. Take $u_s = 5000$ km s$^{-1}$ and ask how long to reach the knee at $10^{15}$ eV. In the ambient $3\ \mu$G field, $r_g = 1.1\times10^{18}$ cm and even Bohm diffusion gives $D = 1.1\times10^{28}\ \mathrm{cm^2\,s^{-1}}$ and $t_{\rm acc} \approx 28{,}000$ years, which is far longer than the remnant stays fast. But cosmic rays streaming ahead of the shock amplify the field they’re moving through, and at $100\ \mu$G the gyroradius drops by a factor of thirty, so $D$ drops by thirty, and $$t_{\rm acc} \approx 840\ \mathrm{years},$$ comfortably inside a young remnant’s life. The entire case for supernovae as the origin of Galactic cosmic rays rests on that factor of thirty, and you can see exactly where it enters: through $r_g$, through $\mathcal P$, through $D$.5The amplification is the non-resonant streaming instability worked out by Bell in 2004. The cosmic ray current ahead of the shock drives a return current in the background plasma, and the resulting force is unstable to transverse modes on scales much shorter than the cosmic ray gyroradius. Thin non-thermal X-ray filaments in remnants like Cassiopeia A independently point to fields of hundreds of microgauss.

Push the field far higher and the derivation tells you it stops applying, which is more useful than a formula that keeps returning numbers. Near a magnetar at $10^{15}$ G a 1 GeV proton has $r_g = 3\times10^{-9}$ cm, smaller than an atom. There is no turbulence at that scale, so $\mathcal{P}(k_{\rm res})$ is essentially zero, so there is no scattering and no diffusion. The particle is welded to one field line and the problem becomes one dimensional motion along a prescribed curve. Diffusion needs the gyroradius to sit comfortably inside the turbulent inertial range, and outside that window the diffusion coefficient stops meaning anything at all.

The same thing happens at the other end, where it’s more famous. As energy climbs, $r_g$ grows until it’s comparable to the system, at which point $t \gg \tau_c$ fails and the particle simply leaves. Requiring $r_g > R$ for a source of size $R$ and field $B$ gives the Hillas criterion, $E_{\max} \approx 0.9\,Z\beta\,(B/\mu\mathrm{G})(R/\mathrm{kpc})$ EeV. The whole Galactic disk manages a few times $10^{18}$ eV. We detect particles at $10^{20}$ eV with gyroradii of about a hundred kiloparsecs, larger than the Galaxy, so they were neither made here nor confined here, and they arrive barely deflected. The highest energy cosmic rays are the only ones that still remember their direction, which is exactly why so much effort goes into catching the rarest particles in the sky.

Let me be honest about what’s shaky, because a derivation you can’t criticise is a derivation you don’t understand. Quasi-linear theory assumes the fluctuation is a small perturbation over one orbit, and we then applied it with $\delta B \approx B_0$, which is outside its formal domain. Real magnetohydrodynamic turbulence is not the isotropic slab we assumed either; in the Goldreich and Sridhar picture the eddies are stretched along the local field, which starves the resonant parallel wavenumber and predicts much weaker scattering than we computed. And cosmic rays are not test particles. They carry an energy density of about $1$ eV cm$^{-3}$, essentially equal to the magnetic field and the turbulent gas, so they generate the very waves that scatter them and the problem is really nonlinear. Every one of those objections attaches to a specific line in the derivation above, which is the point of having done it line by line.

What I’d want you to take from this isn’t the value of $D$. It’s the shape of the argument, because it recurs everywhere. Write the exact microscopic equation of motion. Identify the variable that random walks. Integrate a correlation function to get a transport coefficient. Discover which piece of the environment the system actually couples to, and let a delta function tell you rather than guessing. Then close the loop by asking which part of the phase space is the bottleneck, because that’s what the answer will be controlled by. You can run that program on cosmic rays in the Galaxy, on electrons in a metal, on momentum in a fluid, and on heat in a solid, and it is the same program every time. The particles change and the correlation function changes. Nothing else does.

References and Footnotes

  • 1
    This is worth sitting with, because it’s the reason the whole subject exists. A uniform field cannot scatter anything. It rotates the velocity around $\hat{z}$ at a constant rate and leaves the angle to $\hat{z}$ untouched forever, so a particle in a uniform field slides along its field line for eternity with perfect memory of where it was going. Every bit of transport physics that follows comes from the difference between the real field and a uniform one. ↩︎
  • 2
    The condition $t \gg \tau_c$ is not a technicality, it’s the entire definition of diffusive behaviour. On timescales shorter than the correlation time the motion is ballistic and $\langle(\Delta\mu)^2\rangle \propto t^2$, not $t$. Everything in this post applies only after the particle has been scattered many times. That’s also why diffusion fails at the highest cosmic ray energies, where a particle crosses the Galaxy before it gets scattered even once. ↩︎
  • 3
    Prefactors of order unity in this formula are convention dependent, because different books normalise the spectrum differently, use one sided or two sided $k$, and include one or both transverse components in $\delta B^2$. I’m using $\int P(k)\,dk = \delta B^2$ with $k$ running over positive values only. If you compare with a textbook and find a factor of two, that’s why, and it doesn’t change any scaling. ↩︎
  • 4
    It is absurdly small, and it’s wrong, for an instructive reason: that estimate assumes the field lines themselves are straight. They’re not. Turbulent field lines separate from each other exponentially, so a particle sliding along one is carried sideways for free without ever hopping. Including this field line random walk puts the ratio nearer $10^{-2}$. Perpendicular transport remains the least settled part of the theory and the place where numerical simulations have contributed most. ↩︎
  • 5
    The amplification is the non-resonant streaming instability worked out by Bell in 2004. The cosmic ray current ahead of the shock drives a return current in the background plasma, and the resulting force is unstable to transverse modes on scales much shorter than the cosmic ray gyroradius. Thin non-thermal X-ray filaments in remnants like Cassiopeia A independently point to fields of hundreds of microgauss. ↩︎
Posted in Expository, Projects, Research | Tagged , , , , , | Leave a comment

The rubber sheet gets the geometry right and the mechanism wrong!

You have seen this demonstration. Somebody stretches a rubber sheet over a frame, drops a bowling ball in the middle, and the sheet sags. Then they roll a marble across it and the marble curves toward the bowling ball, maybe even loops around it once or twice before spiralling in. And the narrator says: that is gravity. Mass curves spacetime, and curved spacetime tells matter how to move. Everyone nods. It is on the cover of textbooks. It is in every documentary. I have used it myself when someone asks me at a party what general relativity is about, because it takes ten seconds and it gets the shape of the idea across.

But it is wrong. Not wrong in the sense of being a simplification that leaves out details, which would be fine, but wrong in a specific and interesting way: it is an accurate picture of something that is not the thing causing gravity. And once you see exactly what it has left out, you understand general relativity considerably better than the demonstration was ever going to teach you.

The objection you usually hear is that the demonstration is circular. The marble rolls toward the bowling ball because real gravity is pulling it down into the dip, so you are using gravity to explain gravity. This is true, and it is the weakest objection available. It is an inelegance in the staging, not an error in the physics. You could imagine running the whole thing in free fall with the sheet under tension and a marble constrained to its surface, and you would still have a geometry, and the marble would still follow the shortest path across it. The circularity is a presentational problem. I want to make a harder objection.

Here it is. Take an apple and hold it at rest above the ground. Let go. It falls. Now go back to the rubber sheet and put a marble on it, at rest, on a flat part of the sheet far from the bowling ball. It does not move. To make the marble move on the sheet, you have to give it a velocity, or you have to put it where the sheet is already sloped, at which point real gravity does the work. But the apple was at rest. It had no velocity. It was not sitting on a slope in space. So what curved geometry acted on an object that was not moving?

The rubber sheet cannot answer this, and the reason is that it has deleted the axis along which the apple was moving. An apple at rest in space is not at rest in spacetime. It is moving through the time direction, and moving through it fast. That motion is what the curvature acts on. The rubber sheet shows you two space directions and no time direction at all, which means it has thrown away the entire mechanism.1Sometimes people defend the demonstration by saying the sheet is meant to represent spacetime, not space, with time somehow implicit. It is not. The surface drawn in every version of this demonstration is a two dimensional surface, and both of its dimensions are spatial. There is no axis on the sheet along which a stationary object advances. If there were, a stationary marble would move along it, which is exactly what does not happen.

the sheet you are showntwo dimensions, both spatialwhat is in the picturex, a space directiony, a space directionz, suppressedt, absentThe apple moves along t.It is not in the picture.
The rubber sheet is a two dimensional spatial surface. Every direction it shows is a direction of space. The direction along which a stationary apple is actually travelling, and along which the curvature acts, has been removed from the picture before the demonstration begins.

Let me make that precise, because “the time direction matters more” is the kind of claim that should come with an equation attached. Take a weak, static gravitational field, meaning a Newtonian potential $\Phi$ with $|\Phi|/c^2 \ll 1$, which covers the Earth, the Sun, and essentially every situation in which anyone has ever demonstrated a rubber sheet. To first order in $\Phi/c^2$, the metric of spacetime is $$ds^2 = -\left(1 + \frac{2\Phi}{c^2}\right)c^2 dt^2 + \left(1 – \frac{2\Phi}{c^2}\right)\left(dx^2 + dy^2 + dz^2\right).$$ There are two corrections to flat spacetime here. One sits on the time part, $g_{00}$, and one sits on the space part, $g_{ij}$. They are the same size, both of order $\Phi/c^2$. The rubber sheet is a picture of the second one. Let us find out what the first one does.2This form of the metric is written in what is called the Newtonian gauge or the conformal Newtonian gauge. The two potentials appearing in $g_{00}$ and $g_{ij}$ are in general independent functions, often written $\Phi$ and $\Psi$, and they are only forced to be equal when the stress energy tensor has no anisotropic stress. For ordinary matter, dust and stars and planets, this condition holds, so a single $\Phi$ does the job.

A free particle in general relativity moves along a geodesic, which is the curve satisfying $$\frac{d^2 x^\mu}{d\tau^2} + \Gamma^\mu{}_{\alpha\beta}\, \frac{dx^\alpha}{d\tau}\frac{dx^\beta}{d\tau} = 0,$$ where $\tau$ is proper time along the path. This is just the statement that the particle goes as straight as the geometry allows. Now take our apple, which is moving slowly, so that its spatial velocity is negligible compared to $c$. Then $dx^i/d\tau \approx 0$ while $dx^0/d\tau = c\, dt/d\tau \approx c$, and the double sum over $\alpha$ and $\beta$ collapses to the single term with $\alpha = \beta = 0$: $$\frac{d^2 x^i}{d\tau^2} \approx -\Gamma^i{}_{00}\, c^2.$$ Everything now depends on one Christoffel symbol. For a static metric, $$\Gamma^i{}_{00} = -\tfrac{1}{2} g^{ij} \partial_j g_{00},$$ and substituting $g_{00} = -(1 + 2\Phi/c^2)$ and $g^{ij} \approx \delta^{ij}$ gives $\Gamma^i{}_{00} = \partial_i \Phi / c^2$. Therefore $$\frac{d^2 x^i}{dt^2} = -\partial_i \Phi, \qquad \text{that is,} \qquad \vec{a} = -\nabla\Phi.$$ That is Newton’s law of gravitation, recovered exactly, and I want you to look carefully at where it came from. It came from $g_{00}$. The spatial part of the metric, the part the rubber sheet is drawing, never appeared in the calculation. Not once. It contributed nothing.

You can see why from the structure of the geodesic equation. The spatial components of the metric couple to $dx^i/d\tau$, the rate at which the particle moves through space. For an apple, that rate is tiny. The time component couples to $dx^0/d\tau$, the rate at which the particle moves through time, and that rate is enormous. The two corrections to the metric are the same size, but they get multiplied by wildly different things. The relative contribution of the spatial curvature to the motion of a particle moving at speed $v$ is suppressed by a factor of order $(v/c)^2$. An apple that has fallen a metre is moving at about $4.4$ metres per second, so $(v/c)^2 \approx 2 \times 10^{-16}$. The rubber sheet is showing you the part of the effect that contributes roughly one part in ten thousand million million.

If you want independent confirmation that this accounting is right, look at the bending of starlight, which is the one case where the spatial part genuinely matters. A photon has $v = c$, so the suppression factor is gone and both parts of the metric contribute equally. Compute the deflection using only the time curvature, which is what you get from a naive Newtonian calculation treating light as a fast particle, and you find $\Delta\theta = 2GM/(c^2 b)$ for impact parameter $b$. Compute it in full general relativity and you find $$\Delta\theta = \frac{4GM}{c^2 b},$$ exactly twice as much. That missing factor of two is the spatial curvature, contributing its equal half. For the Sun this is $1.75$ arcseconds against a Newtonian $0.87$, and Eddington’s 1919 expedition measured the larger number, which is what made Einstein famous.3The 1919 measurement was considerably less precise than the confident retellings suggest, with error bars large enough that the two candidate values were not separated as cleanly as one would like. The result has since been confirmed to far better precision by radio interferometry of quasars occulted by the Sun, which now agrees with the general relativistic prediction to better than one part in ten thousand. The factor of two is not in doubt. So the spatial curvature is real, and it is measurable, and for light it accounts for half of everything. For an apple it accounts for essentially nothing, and an apple is what the demonstration claims to be explaining.

None of this means the rubber sheet is a fantasy. It is a picture of a real geometric object, and here it is worth being precise about which one. Take the Schwarzschild solution, freeze the time coordinate at some value $t = \text{const}$, and look at the equatorial plane $\theta = \pi/2$. The geometry of that two dimensional slice is $$dl^2 = \left(1 – \frac{r_S}{r}\right)^{-1} dr^2 + r^2 d\phi^2,$$ and if you ask what surface embedded in ordinary flat three dimensional space has this same intrinsic geometry, you get $z(r) = 2\sqrt{r_S(r – r_S)}$, a paraboloid opening outward. That is the funnel. That is the rubber sheet, and it is not an approximation or an artist’s impression: it is the exact embedding diagram of a spatial slice of Schwarzschild spacetime, computed by Flamm in 1916. The demonstration is drawing a correct picture. It is a correct picture of the wrong thing.

So what is the right thing? The curvature that makes apples fall lives in $g_{00}$, and $g_{00}$ is the thing that tells you how fast clocks run. Write out the proper time for a stationary observer at position $x$: setting $dx = dy = dz = 0$ in the metric gives $$d\tau = \sqrt{1 + \frac{2\Phi}{c^2}}\; dt \approx \left(1 + \frac{\Phi}{c^2}\right) dt.$$ A clock deeper in the potential well, where $\Phi$ is more negative, ticks slower. Near the Earth’s surface, $\Phi = gh$ for height $h$, and the fractional difference in tick rate between two clocks separated vertically by $h$ is $gh/c^2$. This is not a metaphor and it is not a small print correction. Pound and Rebka measured it in 1959 using a $22.5$ metre tower at Harvard, where the predicted fractional shift is $$\frac{gh}{c^2} = \frac{9.81 \times 22.5}{8.99 \times 10^{16}} = 2.46 \times 10^{-15},$$ and they confirmed it to about one percent. Fifty years later, optical lattice clocks at NIST resolved the same effect over a height difference of $33$ centimetres. Your head ages faster than your feet by about half a microsecond over an eighty year life, and that is a real, measurable, in principle observable fact about you.4The half microsecond figure comes from $gh/c^2 \approx 1.9 \times 10^{-16}$ for a height difference of $1.7$ metres, multiplied by roughly $2.5 \times 10^9$ seconds in eighty years. Whether this is worth mentioning at parties depends heavily on the party.

groundslowerfasterhClocks do not agree.Δτ / τ = g h / c²22.5 m gives 2.46 × 10⁻¹⁵Pound and Rebka, 19590.33 m gives 3.6 × 10⁻¹⁷NIST clocks, 2010
The curvature that makes things fall is curvature of the time direction, and it shows up as clocks running at different rates at different heights. This is measured routinely and to high precision. It is also, unlike the sag of a rubber sheet, the actual cause of gravity.

Now, why should a gradient in clock rate make anything move? Here is the principle, and it is one of the most beautiful statements in physics: a free particle follows the path through spacetime that maximises the proper time elapsed along it. Not minimises. Maximises. Of all the ways of getting from one event to another, an unforced particle takes the one on which its own watch reads the longest. Let me show you that this single statement contains Newtonian mechanics whole. For a slow particle in a weak field, $$d\tau = dt\sqrt{1 + \frac{2\Phi}{c^2} – \frac{v^2}{c^2}} \approx dt\left(1 + \frac{\Phi}{c^2} – \frac{v^2}{2c^2}\right),$$ so the total proper time along a path is $$\int d\tau = \int dt + \frac{1}{c^2}\int\left(\Phi – \frac{v^2}{2}\right) dt.$$ The first term is the same for every path between the same two events, so maximising $\int d\tau$ means maximising the second integral, which means minimising $$\int \left(\frac{v^2}{2} – \Phi\right) dt.$$ That integrand is the kinetic energy minus the potential energy, per unit mass. It is the Lagrangian. Extremising proper time is the principle of least action, and running it through the Euler-Lagrange equation gives back $\ddot{x}^i = -\partial_i \Phi$, which is where we came in.5The sign flip between maximising proper time and minimising the action is not a slip. It comes from the overall minus sign in the timelike part of the metric signature, and it means that free particles in relativity extremise a quantity that reduces to the classical action with the conventional sign.

So here is the picture that actually explains gravity. Everything in the universe is moving through spacetime at speed $c$, always. An object sitting still on your desk is not stationary; it is travelling through the time direction at the full speed, and through the space directions at nothing. Gravity does not push on it. What gravity does is make the time direction run at different rates in different places, and a path that keeps going straight through a region where time runs at a gradient does not stay parallel to where it started. Some of the object’s motion gets rotated out of the time direction and into a space direction. That rotation is what you call falling.

And that gives us the analogy I actually want, which is not a sheet at all. It is a cart with two wheels on a common axle, rolling forward. The cart has no steering. Nobody pushes it sideways. If both wheels turn at the same rate, it goes dead straight. Now let the ground under the right wheel be slightly slower, so that the right wheel covers a little less distance per turn than the left one. The cart veers to the right. No sideways force acted on it. The cart was going straight the entire time. It curved because “straight” in a medium with a speed gradient means “bending toward the slow side.”

uniform groundright side slowergoes straightfastslowturns toward the slow side
A cart with no steering, rolling forward. On uniform ground it goes straight. Where the ground under one wheel is slower, it turns toward the slow side without any sideways force acting on it. Replace “distance rolled” with “proper time elapsed” and this is exactly what gravity does to a worldline.

The translation is direct. The cart’s forward motion is the object’s motion through time, which never stops and never slows. The speed gradient across the axle is the gradient in clock rate, $gh/c^2$. The turning is the worldline bending out of the pure time direction and acquiring a spatial component, which is to say the object accelerating downward. And notice what this analogy explains that the rubber sheet cannot. It explains why a stationary object falls, because the cart was never stationary, it was rolling forward the whole time. It explains why gravity is universal and independent of mass or composition, because the turning depends only on the ground, not on what the cart is made of or how heavy it is. And it needs no external gravity to run, because nothing in the setup falls.

It is worth getting a feel for how gentle this bending is. Take the apple falling under $g = 9.81$ and draw its worldline on a spacetime diagram with the time axis measured in metres, so that one second of elapsed time is $c \times 1\,\text{s} = 3 \times 10^8$ metres. In that second the apple moves $4.9$ metres sideways. Its worldline is a parabola $x = gT^2/2c^2$ where $T = ct$, and the radius of curvature of that parabola at its vertex is $$R = \frac{c^2}{g} = \frac{8.99 \times 10^{16}}{9.81} \approx 9.2 \times 10^{15}\,\text{m},$$ which is $0.97$ light years. The worldline of a falling apple is bent on a scale of about one light year. It is almost perfectly straight. Gravity feels overwhelming to us not because the curvature is large but because we persist through enormous stretches of the time direction, and a tiny angle applied over a light year of lever arm still knocks you off a ladder.

drawn to scalecurvature exaggeratedctx1 second = 3 × 10⁸ m4.9 messentially straightctxappleR = c² / g0.97 light yearsbend magnified
A falling apple’s worldline, drawn honestly on the left and exaggerated on the right. In one second the apple advances $3\times10^8$ metres through time and about five metres through space. The curvature is real, and its radius is roughly one light year.

Now let me be honest about where my cart analogy breaks, because an analogy you cannot break is an analogy you do not understand. A uniform speed gradient across the axle is not curvature. It is the flat spacetime of an accelerating observer, which is what the equivalence principle tells you a uniform gravitational field is indistinguishable from. Genuine curvature is what you get when the gradient itself varies from place to place, so that two carts started off parallel do not merely both turn but turn by different amounts and converge. That is the real signature, and it is called geodesic deviation. If $\xi^\mu$ is the separation between two neighbouring geodesics with four velocity $u^\mu$, then $$\frac{D^2 \xi^\mu}{d\tau^2} = -R^\mu{}_{\alpha\nu\beta}\, u^\alpha \xi^\nu u^\beta,$$ where $R^\mu{}_{\alpha\nu\beta}$ is the Riemann tensor. This equation is the honest definition of gravity. In the weak field limit its leading component reduces to $R^i{}_{0j0} \approx \partial_i \partial_j \Phi / c^2$, the Newtonian tidal tensor, which is the thing that stretches you along the radial direction and squeezes you sideways as you fall toward a planet. A uniform field can be transformed away by jumping. Tides cannot, and that is how you know the curvature is real.

ctxreleased from rest, apartnow togetherTwo straight linesthat meet.No force acted.Both are geodesics.That is curvature.
Two objects released from rest side by side, drawn in spacetime with time running upward. Both worldlines are geodesics, and yet they converge. In flat spacetime, straight lines that start parallel stay parallel. The convergence is the curvature, and it is measured by the Riemann tensor. This is the part no uniform gradient can imitate.

So the summary is this. The rubber sheet draws a real object, the spatial geometry of a slice of Schwarzschild spacetime, and it draws it accurately. But that object contributes essentially nothing to why things fall, suppressed relative to the effect that matters by a factor of order $(v/c)^2$, which for anything slower than light is a very small number indeed. The gravity you feel comes from the curvature of the time direction, from the plain fact that clocks run at different rates at different heights, and from the principle that a free object takes the path through spacetime along which its own clock reads the most. Nothing pulls the apple. The apple goes straight. Straight, in a place where time runs at a gradient, means down.

I still use the rubber sheet at parties. It takes ten seconds and it gets across the idea that geometry is doing the work, which is the part most people have never heard. But if the conversation lasts longer than ten seconds, I put the sheet away and start talking about clocks, because that is where the physics actually is. The demonstration answers the question “what does curved space look like.” The question we asked was “why do things fall,” and those turn out to be almost entirely different questions.

References and Footnotes

  • 1
    Sometimes people defend the demonstration by saying the sheet is meant to represent spacetime, not space, with time somehow implicit. It is not. The surface drawn in every version of this demonstration is a two dimensional surface, and both of its dimensions are spatial. There is no axis on the sheet along which a stationary object advances. If there were, a stationary marble would move along it, which is exactly what does not happen. ↩︎
  • 2
    This form of the metric is written in what is called the Newtonian gauge or the conformal Newtonian gauge. The two potentials appearing in $g_{00}$ and $g_{ij}$ are in general independent functions, often written $\Phi$ and $\Psi$, and they are only forced to be equal when the stress energy tensor has no anisotropic stress. For ordinary matter, dust and stars and planets, this condition holds, so a single $\Phi$ does the job. ↩︎
  • 3
    The 1919 measurement was considerably less precise than the confident retellings suggest, with error bars large enough that the two candidate values were not separated as cleanly as one would like. The result has since been confirmed to far better precision by radio interferometry of quasars occulted by the Sun, which now agrees with the general relativistic prediction to better than one part in ten thousand. The factor of two is not in doubt. ↩︎
  • 4
    The half microsecond figure comes from $gh/c^2 \approx 1.9 \times 10^{-16}$ for a height difference of $1.7$ metres, multiplied by roughly $2.5 \times 10^9$ seconds in eighty years. Whether this is worth mentioning at parties depends heavily on the party. ↩︎
  • 5
    The sign flip between maximising proper time and minimising the action is not a slip. It comes from the overall minus sign in the timelike part of the metric signature, and it means that free particles in relativity extremise a quantity that reduces to the classical action with the conventional sign. ↩︎
Posted in Diary, Expository, Notes, Scratch essays | Tagged , , , , , , , , , , , , , , , , | Leave a comment

A Scientist Before “Scientist”

Mary Somerville is so difficult to put into words that it’s funny. To be honest, who was she? Although you would be correct in calling her a mathematician, you would then be required to clarify the astronomy. The word “astronomer” sounds a bit tame when you consider all the physics, geography, and geology involved in her work. This is no longer an issue because the term “scientist” adequately describes those who spend their lives seeking to comprehend the natural world. The only catch is that there were no scientists when Mary Somerville was born. Of course, people were conducting scientific experiments; it’s just that the term “science” didn’t exist, which is a little strange.

Although she was born in Scotland in 1780, Somerville’s upbringing was far from ideal for a mathematician. Her family initially saw her growing interest in mathematics as a mix of a bothersome curiosity and a reason for concern, likely stemming from their belief that girls weren’t supposed to excel in math. Nonetheless, she gained knowledge. Much of her mathematical education was something she had to seek out on her own; there’s a slightly embellished tale about how she studied geometry and algebra in the middle of the night, but the essential idea remains the same. This is my favorite part of her story because it describes the math learning process so realistically. Some symbols appear, and you find yourself confused. This annoys you. You come across an explanation for them. There are five new things in that explanation that you don’t get. Somehow, you’ve picked up a completely new subject a few months down the road. It appears pointless that Somerville accomplished this without Google.

Over time, she excelled at it. Competent enough to tackle Laplace’s Mécanique Céleste, conversant enough to correspond with prominent European scientists, and capable of delving deeply into astronomical mathematics. The mathematical machinery of the heavens, including the gravitational motion of the planets, was something Laplace was attempting to describe. Famously, his writing was not something you could read on a leisurely basis before turning in for the night. The Mechanism of the Heavens was written by Somerville. This, which is sometimes called a translation, is disrespectful to her writing. While it’s true that she was translating the French, she was also filling in the math and reasoning so that others could learn from it. Cambridge decided to use the book because of how popular it was. As someone who was advised against studying algebra in the past, that path seems rather peculiar.

However, the truly peculiar historical accident was brought about by a book. She released On the Connexion of the Physical Sciences in 1834. Her priorities are clear from the title. Not physics or math, but rather the fields that form their connections. Physics, chemistry, physics, heat, light, electricity, and magnetism. These were getting into increasingly complex topics, but Somerville hoped to demonstrate that they were not independent worlds. They were just two perspectives on the same cosmos. Remember that the best ideas don’t always sound obvious until someone has spent years pointing them out; this may sound obvious when you say it now, but it’s worth remembering.

Because it is necessary, we continue to divide information into different categories. Someone has to make a call at some point: schools have schedules, universities have departments, to designate which classroom will be used for physics and which for mathematics. Nature was never sluggish with this setup. All these things, chemistry, electromagnetism, gravity, and nuclear physics, are happening simultaneously inside a star. It is not a chemistry problem when the chemist walks in. Geometry, thermodynamics, quantum physics, and information theory can all be occupied for an afternoon by contemplation of a black hole. When you delve deeply enough into either mathematics or physics, the lines between them start to blur in an odd way. At the same time that Somerville was becoming fascinated by these connections, the scientific community was heading in the opposite social direction, with men developing into specialists.

As a result, an unexpectedly mundane issue arose. Please tell me what you intended to refer to each of these individuals as. Physicists, geologists, astronomers, mathematicians, and chemists were all present. “Man of science” and “natural philosopher” were both in use in the past, but neither of these terms adequately described what was happening in English at the time. When discussing this naming issue in 1834 while reviewing Somerville’s On the Connexion of the Physical Sciences, William Whewell proposed a rather strange new term that was derived from the word “artist”: scientist. At first, I thought the word was hilarious. Obviously so. So long as everyone agrees to stop making fun of it, every word is slightly absurd.

Whewell basically looks at Mary Somerville, realizes that she knows far too many subjects to fit into any existing category, and has to invent the word “scientist” for her on the spot, according to a very alluring online version of this story. While I would love it if that were the whole story, it’s that good, history is usually less accommodating. Whewell was addressing a broader issue with scientific terminology; the term was not coined solely due to Somerville’s gender or the fact that her field of study was a matter of debate. The genuine serendipitous event is nearly more delightful. In a dispute sparked by a book written specifically to unite increasingly fragmented scientific disciplines, the word “scientist” was first used in print. A new term for scientists was introduced, and a book was also released, urging us to view science as an integrated whole.

Somerville also has a surprisingly contemporary vibe to it. This is the age of ridiculous specialization. Most of the time, this will happen. But there is now an overwhelming amount of information that no one individual can adequately process. No matter how much time you devote to studying a specific area of theoretical physics, there will always be some literature that your neighbors assume you are familiar with. If you haven’t figured out how to make the day last much longer than twenty-four hours, the ideal of being mathematically and physically literate is meaningless. It seems, though, that the intriguing concepts did not get the message. But they persist in traversing borders. Geometry deviated into weightlessness. When it entered particle physics, group theory made a mistake. A collision with condensed matter occurred in topology. Black holes were discovered by information theory. Occasionally, two fields develop independently for many decades until someone realizes that they have been discussing remarkably similar topics in different languages.

I suppose that explains my fondness for Mary Somerville. She “refused to pick one thing,” but that’s just for the Instagram slide. This is because, thankfully, she had developed the habit of asking for connections between concepts, which proved to be quite useful. Something is different. Possessing a broad range of interests can make you seem diverse, but it takes actively seeking out connections to make that diversity matter. If you want to be a good physicist, you need to know how to ask certain kinds of questions, and math is a necessary part of that. When you read about astronomy, a mechanical problem takes on a new perspective. You learn the language of a thought that was hazy in one field to another. As time goes on, it becomes increasingly difficult to distinguish between one interest and the next.

Even in her latter years, ninety-one-year-old Mary Somerville continued to pursue mathematics. She passed away in 1872. A lot bigger scientific world was emerging than the one she had just entered, and the little word that Whewell had proposed was gradually becoming the norm. The term “scientist” is so commonplace now that a six-year-old can confidently declare their goal of becoming one without realizing that the same statement would have had no meaning two centuries ago. It has an eerie beauty to it. Mary Somerville’s life happened at the exact moment when our conception of science was shifting, not because she was the sole inventor of the word (though she certainly wasn’t). Amidst all this discord, someone was penning a book about the interconnectedness of all things, and new terms were being coined.

I don’t want to put life through that, but maybe there’s a lesson there. However, if you want to become an expert in a certain field, specialization is inevitable. That is an undeniable fact. However, you shouldn’t allow the label of your field to dictate the limits of your interest. Sitting next to you in the other class can be just what you need sometimes. Each time, it’s five topics up there. Starting with a question can lead you astray from your original intent; it is only after you’ve gone on that you come to appreciate the journey as its most captivating aspect.

Posted in Diary, Notes, Scratch essays | Tagged | Leave a comment

What are Quarternions?

I used to think complex numbers were the end of the line. You start with counting, patch in zero and the negatives, then the fractions, then the irrationals to fill the gaps between them, and finally you throw in $i$ to answer $x^2 = -1$. At that point the number system feels closed. Complete. Nothing left to want. I believed that for a long time.

Then I was talking with a friend about where complex numbers actually show up in the physical world, and he asked, almost in passing, whether I had ever looked into quaternions. I had heard the word. I could have told you they were “like complex numbers but bigger” and left it there, which is another way of admitting I did not understand them at all. He does. He has a PhD in computer science and the rare gift of making a hard idea feel inevitable instead of arbitrary, and over one long conversation, followed by most of a weekend spent chasing down the details he left me with, quaternions turned from a word I recognized into a structure I could see. This post is my attempt to hand you that same shift.

The place to start is not with quaternions but with a small miracle hiding inside the complex numbers you already know, because quaternions exist only to reproduce that miracle one dimension up. Take a point in the plane, written as $z = x + iy$, and multiply it by a unit complex number $w = \cos\phi + i\sin\phi$. Multiply the two out, keeping $i^2 = -1$:

$$ wz = (\cos\phi + i\sin\phi)(x + iy) = (x\cos\phi – y\sin\phi) + i\,(x\sin\phi + y\cos\phi). $$

Look at the two real numbers that came out. The new $x$-coordinate is $x\cos\phi – y\sin\phi$, the new $y$-coordinate is $x\sin\phi + y\cos\phi$, and that is exactly the rotation you met in the vectors post, the same $R_{ij}$ dictionary that turned one observer’s components into another’s. Multiplying a point by a unit complex number rotates that point about the origin by the angle $\phi$. Rotation, one of the most geometric operations there is, turns out to be nothing more than multiplication in disguise.

Rotation of a complex number Multiplication by the unit complex number w rotates z counterclockwise through the angle phi while preserving its distance from the origin. Re Im φ z wz

Figure 1: Multiplying the complex number z by the unit complex number w = cos φ + i sin φ rotates it about the origin through the angle φ while preserving its distance from the origin. The dashed arc shows the circular path of the endpoint.

Two things make this work, and both matter for what comes next. The first is that the length of $z$ is unchanged: a unit $w$ rotates without stretching. That is a statement about a norm, $|wz| = |w|\,|z|$, and squaring it gives the identity $(a^2+b^2)(c^2+d^2) = (ac-bd)^2 + (ad+bc)^2$, which says that the product of two sums of two squares is again a sum of two squares.1This is the Brahmagupta–Fibonacci identity, known in essence to Diophantus. It is exactly the statement that the complex norm is multiplicative, $|w z|^2 = |w|^2 |z|^2$, written out in real coordinates with $w = a+bi$ and $z = c+di$. Keep an eye on the number of squares: two here, and the question of whether the same trick works for three and for four squares is the whole story below. The second is that rotations compose the way multiplication does, and this is worth doing rather than asserting. Rotate by $\phi_1$ and then by $\phi_2$, which means multiply by the two unit numbers in turn:

$$ (\cos\phi_2 + i\sin\phi_2)(\cos\phi_1 + i\sin\phi_1) = (\cos\phi_2\cos\phi_1 – \sin\phi_2\sin\phi_1) + i\,(\sin\phi_2\cos\phi_1 + \cos\phi_2\sin\phi_1). $$

The angle-addition formulas collapse the right-hand side to $\cos(\phi_1+\phi_2) + i\sin(\phi_1+\phi_2)$, a single rotation through the summed angle. Composing two rotations is nothing but multiplying the two numbers that represent them. Hold onto that sentence. It is the one demand quaternions will have to meet in three dimensions, where composing rotations is a famously trickier business than adding angles. In the plane, rotation is easy and complex numbers are the reason.

Now the obvious hunger sets in. We live in three dimensions, not two, and rotations in space are far richer and far more troublesome than rotations in a plane. If a two-dimensional number system hands us plane rotations for free, surely a three-dimensional one would hand us space rotations the same way. This is precisely the thought that seized William Rowan Hamilton in the 1830s, and he chased it for the better part of a decade.2Hamilton was searching for what he called “triples,” three-dimensional numbers $a + bi + cj$ that could be added and, crucially, multiplied while preserving length, so that unit triples would rotate space the way unit complex numbers rotate the plane. He worked at it for years; there is a much-repeated story that his children would ask him at breakfast whether he could multiply triples yet, and he would have to answer that he could still only add and subtract them.

Set up his problem the natural way. A triple is $a + bi + cj$, with two imaginary units obeying $i^2 = j^2 = -1$, and we want a multiplication that keeps the norm $a^2 + b^2 + c^2$ multiplicative, so that unit triples act as rotations. The entire difficulty collapses onto a single question: what is the product $ij$? For the triples to be closed under multiplication, $ij$ has to be another triple, some combination

$$ ij = \alpha + \beta i + \gamma j $$

with real coefficients $\alpha, \beta, \gamma$. Watch what that assumption forces. Multiply both sides on the left by $i$, and use $i^2 = -1$ together with ordinary associativity:

$$ i(ij) = i^2 j = -j, $$

while the right-hand side becomes

$$ i(\alpha + \beta i + \gamma j) = \alpha i + \beta i^2 + \gamma\, ij = -\beta + \gamma\alpha + (\alpha + \gamma\beta)\,i + \gamma^2 j, $$

where in the last step I substituted the assumed form of $ij$ back into itself. These two expressions must be equal, so match the coefficient of $j$ on each side:

$$ -1 = \gamma^2. $$

There is no real number $\gamma$ whose square is $-1$. The assumption that $ij$ is a triple has destroyed itself. No matter how you try to define the product, $ij$ simply refuses to live in the three-dimensional space spanned by $1$, $i$, and $j$; the multiplication insists on pointing somewhere outside. Hamilton spent years trying to force it back inside, and it could not be done. There is even a clean reason from number theory that it can never be done for any definition at all: the product of two sums of three squares is not always a sum of three squares, so no multiplication on triples can keep the norm multiplicative.3Take $3 = 1^2+1^2+1^2$ and $5 = 0^2+1^2+2^2$, both sums of three squares. Their product is $15$, and $15$ is not a sum of three squares: by Legendre’s three-square theorem a positive integer fails to be a sum of three squares exactly when it has the form $4^a(8b+7)$, and $15 = 8(1)+7$. So the three-square analogue of the identity in the previous footnote is simply false, and a length-preserving multiplication on triples cannot exist. The obstruction is not a failure of imagination; it is a theorem.

The way out is the move that looks like a defeat and turns out to be the discovery. If $ij$ will not stay inside three dimensions, then stop fighting it and give it a home. Call it a new unit, $k := ij$, orthogonal to $1$, $i$, and $j$, and accept that the numbers we are building are four-dimensional, not three. This is what Hamilton finally saw on 16 October 1843, walking with his wife along the Royal Canal in Dublin, and he was so certain it was right and so afraid he would forget it that he carved the defining relation into the stone of Broom Bridge with a knife:

$$ \boxed{\; i^2 = j^2 = k^2 = ijk = -1. \;} $$

A quaternion is a number $q = a + bi + cj + dk$, one real coordinate and three imaginary ones, and everything about how they multiply is squeezed out of that one carved formula. The remarkable thing is how little you have to be told. Those four equalities are the entire definition; every product of $i$, $j$, and $k$ follows from them by algebra alone, and it is worth turning the crank once to see that it really does.

Start from $ijk = -1$ and multiply on the right by $k$. Since $k^2 = -1$,

$$ ijk\cdot k = -k \;\Longrightarrow\; ij\,(k^2) = -k \;\Longrightarrow\; ij\,(-1) = -k \;\Longrightarrow\; ij = k. $$

Now multiply the same $ijk = -1$ on the left by $i$, using $i^2 = -1$:

$$ i\cdot ijk = -i \;\Longrightarrow\; (i^2) jk = -i \;\Longrightarrow\; -jk = -i \;\Longrightarrow\; jk = i. $$

The third product comes out by the same style of manipulation. Multiply $ij = k$ on the left by $k$, so that $k(ij) = k^2 = -1$; reading the left side as $(ki)j$ by associativity gives $(ki)j = -1$, and right-multiplying by $j$ with $j^2 = -1$ turns this into $(ki)(j^2) = -j$, that is $-(ki) = -j$, so

$$ ij = k, \qquad jk = i, \qquad ki = j. $$

The three units cycle, each one times its neighbor returning the third, in the order $i \to j \to k \to i$. Now the order of multiplication, which is where quaternions earn their keep. These products do not commute, and that is the point of the construction rather than a blemish on it. Left-multiply $jk = i$ by $j$:

$$ j(jk) = ji \;\Longrightarrow\; (j^2)k = ji \;\Longrightarrow\; ji = -k, $$

the exact negative of $ij$. The same one-line move on the other two, left-multiplying $ij = k$ by $i$ and right-multiplying $ij = k$ by $j$, gives $ik = -j$ and $kj = -i$. Reversing any product flips its sign, so the forward cycle $i \to j \to k \to i$ carries a plus and the backward cycle a minus,

$$ ij = k,\quad jk = i,\quad ki = j, \qquad ji = -k,\quad kj = -i,\quad ik = -j. $$

i j k forward: + reverse: −

Figure 2: The multiplication rule in one picture. Following the arrows, $ij = k$, $jk = i$, $ki = j$. Going against them flips the sign: $ji = -k$, $kj = -i$, $ik = -j$. Order matters, which is exactly the feature that will let quaternions carry rotations that themselves depend on order.

That last sentence deserves weight. In the vectors post we watched two spatial rotations fail to commute: rotate a book about one axis then another, reverse the order, and the book lands somewhere else, while ordinary addition never cares about order. A number system meant to encode rotations therefore cannot be commutative, and the quaternions are not. The noncommutativity that looks at first like a bug is the precise property that fits them to the job. They gave up an old comfort, $xy = yx$, and bought a new capability with it.

Before going further it pays to learn to multiply any two quaternions at once, not just the units, because a single formula will carry the entire rest of the post. Split a quaternion into its real and imaginary parts exactly as you split a complex number, and give the imaginary part the name it deserves: for $q = a + bi + cj + dk$ write $q = a + \mathbf{u}$, where $a$ is the scalar part and $\mathbf{u} = bi + cj + dk$ is a pure quaternion, an honest three-dimensional vector wearing $i, j, k$ as its axes. Multiplying two pure quaternions is now just the multiplication table applied to every cross term. Take $\mathbf{u} = u_1 i + u_2 j + u_3 k$ and $\mathbf{v} = v_1 i + v_2 j + v_3 k$ and expand the nine products. The three like terms collapse by $i^2 = j^2 = k^2 = -1$ into a single real number, $-(u_1 v_1 + u_2 v_2 + u_3 v_3)$. The six unlike terms fold into the cycle: $u_1 v_2\,(ij) + u_2 v_1\,(ji) = (u_1 v_2 – u_2 v_1)\,k$, and its two rotations give the $i$ and $j$ parts. Collecting everything,

$$ \mathbf{u}\,\mathbf{v} = -(u_1 v_1 + u_2 v_2 + u_3 v_3) + (u_2 v_3 – u_3 v_2)\,i + (u_3 v_1 – u_1 v_3)\,j + (u_1 v_2 – u_2 v_1)\,k. $$

Both pieces are old acquaintances in disguise, and I will name them properly at the very end of this post, once they can land with full force. For now, call the scalar part $-\,\mathbf{u}\cdot\mathbf{v}$ and the vector part $\mathbf{u}\times\mathbf{v}$, so that the product of two pure quaternions is

$$ \mathbf{u}\,\mathbf{v} = -\,\mathbf{u}\cdot\mathbf{v} + \mathbf{u}\times\mathbf{v}. $$

One consequence is immediate and we will lean on it hard. Set $\mathbf{v} = \mathbf{u}$: the cross product of a vector with itself vanishes, so $\mathbf{u}^2 = -\,\mathbf{u}\cdot\mathbf{u} = -|\mathbf{u}|^2$, and a pure unit quaternion squares to $-1$, exactly as $i$ does. There is not a single square root of $-1$ living in the quaternions but a whole sphere of them, one for every direction in space. From here the product of two arbitrary quaternions follows by distributing over the split. With $p = a + \mathbf{u}$ and $q = b + \mathbf{v}$,

$$ pq = (a + \mathbf{u})(b + \mathbf{v}) = ab + a\mathbf{v} + b\mathbf{u} + \mathbf{u}\mathbf{v} = (ab – \mathbf{u}\cdot\mathbf{v}) + (a\mathbf{v} + b\mathbf{u} + \mathbf{u}\times\mathbf{v}). $$

The scalar part of a product is $ab – \mathbf{u}\cdot\mathbf{v}$ and the vector part is $a\mathbf{v} + b\mathbf{u} + \mathbf{u}\times\mathbf{v}$. Every calculation left in this post is a special case of that one line.

The norm we were chasing all along now falls out cleanly, and it is genuinely the multiplicative norm. Define the conjugate of $q = a + bi + cj + dk$ by flipping the sign of the imaginary part, $\bar q = a – bi – cj – dk$, in exact analogy with the complex conjugate. Multiply $q$ by $\bar q$ and every cross term cancels in pairs, because $ij$ and $ji$ carry opposite signs, leaving only the squares:

$$ q\,\bar q = a^2 + b^2 + c^2 + d^2 \;\equiv\; N(q), $$

a real, nonnegative number, the squared length of $q$ as a four-dimensional vector. Multiplicativity now rests on a single fact about conjugation, and with the product formula in hand we can prove it outright instead of borrowing it: conjugating a product reverses the order of the factors, $\overline{pq} = \bar q\,\bar p$.4This order reversal is the same phenomenon as $(AB)^\top = B^\top A^\top$ for matrices, and it is not a coincidence: quaternions can be represented by $2\times 2$ complex matrices, with the conjugate becoming the conjugate transpose. It also rhymes with the covariant bookkeeping from the vectors post, where a transforming quantity reverses order relative to the basis. Order reversal under an involution is a recurring signature of noncommutative structures. Conjugation flips the sign of the vector part, so from the general product $pq = (ab – \mathbf{u}\cdot\mathbf{v}) + (a\mathbf{v} + b\mathbf{u} + \mathbf{u}\times\mathbf{v})$ we read off

$$ \overline{pq} = (ab – \mathbf{u}\cdot\mathbf{v}) – (a\mathbf{v} + b\mathbf{u} + \mathbf{u}\times\mathbf{v}). $$

Now multiply the conjugates in the opposite order. Using $\mathbf{v}\mathbf{u} = -\mathbf{v}\cdot\mathbf{u} + \mathbf{v}\times\mathbf{u} = -\mathbf{u}\cdot\mathbf{v} – \mathbf{u}\times\mathbf{v}$, since the dot product is symmetric and the cross product changes sign when its arguments swap,

$$ \bar q\,\bar p = (b – \mathbf{v})(a – \mathbf{u}) = ab – b\mathbf{u} – a\mathbf{v} + \mathbf{v}\mathbf{u} = (ab – \mathbf{u}\cdot\mathbf{v}) – (a\mathbf{v} + b\mathbf{u} + \mathbf{u}\times\mathbf{v}), $$

which is exactly $\overline{pq}$. The order reversal is not an analogy imported from matrices; it is a two-line consequence of how the vector part multiplies. Because $N(p) = p\bar p$ is a real scalar and commutes with everything, the multiplicativity is then three short steps:

$$ N(pq) = (pq)\overline{(pq)} = pq\,\bar q\,\bar p = p\,N(q)\,\bar p = N(q)\,p\bar p = N(p)\,N(q). $$

The norm is multiplicative. Written out in the sixteen real products this becomes Euler’s four-square identity, the four-dimensional cousin of the two-square identity we started with, and the reason four is the magic dimension where three failed.5Euler’s four-square identity, that a product of two sums of four squares is again a sum of four squares, predates Hamilton by nearly a century; he found it in 1748 while working on Lagrange’s four-square theorem. Hamilton’s quaternions are, in a sense, the hidden reason the identity is true: it is just $N(pq) = N(p)N(q)$ in coordinates. The pattern of which dimensions admit such an identity is severe. Sums-of-squares multiply cleanly only in dimensions $1$, $2$, $4$, and $8$, a result of Hurwitz, and the four cases are the real numbers, the complex numbers, the quaternions, and the octonions. There is no five-dimensional or seven-dimensional number system of this kind, and there never will be.

One consequence is worth pulling out on its own, because it is the property that makes quaternions a genuine number system and not merely an interesting array. Since $q\bar q = N(q)$ is a positive real number for any $q \ne 0$, we can divide by it, and

$$ q^{-1} = \frac{\bar q}{N(q)} $$

is a true multiplicative inverse: $q\,q^{-1} = q\bar q / N(q) = 1$. Every nonzero quaternion can be divided by. A system where you can add, subtract, multiply, and divide, with division always possible except by zero, is called a division algebra, and division algebras are extraordinarily rare. Over the real numbers, if you insist on the ordinary law $x(yz) = (xy)z$, there are exactly three: the real numbers, the complex numbers, and the quaternions, and nothing else, ever.6This is Frobenius’s theorem, 1877: the only finite-dimensional associative division algebras over the reals are $\mathbb{R}$, $\mathbb{C}$, and $\mathbb{H}$ (the $\mathbb{H}$ is for Hamilton). Each rung of the ladder costs you a property. Passing from $\mathbb{R}$ to $\mathbb{C}$ you lose ordering: there is no consistent way to say one complex number is less than another. Passing from $\mathbb{C}$ to $\mathbb{H}$ you lose commutativity, as we have seen. If you are willing to give up associativity as well you can reach one dimension further, to the eight-dimensional octonions $\mathbb{O}$, and there the ladder ends for good. This is the same “the language runs out” story that ran through the earlier posts, now with a precise list of exactly how far it can possibly go.

So we have our four-dimensional number system, and it does everything complex numbers did. The question that started Hamilton off is still open, though: does it actually rotate space? It does, and the way it does is stranger and more beautiful than the two-dimensional case, so we should build it carefully rather than assert it.

First, identify ordinary three-dimensional vectors with the quaternions that have no real part. A vector $\mathbf{v} = (v_1, v_2, v_3)$ becomes the pure quaternion $v = v_1 i + v_2 j + v_3 k$, and a quaternion is pure exactly when $\bar v = -v$. These pure quaternions are our three-dimensional space, sitting inside the four-dimensional quaternions as the imaginary part. The recipe for rotating them is the sandwich product: pick a unit quaternion $q$, meaning $N(q) = 1$, and send

$$ v \;\longmapsto\; q\,v\,q^{-1} = q\,v\,\bar q, $$

where the second equality holds because $q^{-1} = \bar q$ when $q$ is a unit quaternion. I claim this map is a rotation of three-dimensional space, and there are three things to check, each of which is a short calculation rather than an act of faith.

It preserves length, because the norm is multiplicative: $N(qv\bar q) = N(q)\,N(v)\,N(\bar q) = 1\cdot N(v)\cdot 1 = N(v)$, using $N(q) = N(\bar q) = 1$. It sends pure quaternions to pure quaternions, so a vector maps to a vector and never sprouts a real part, which you can see by conjugating the result and finding it flips sign:

$$ \overline{q v \bar q} = \overline{\bar q}\;\bar v\;\bar q = q\,(-v)\,\bar q = -\,q v \bar q, $$

where $\bar v = -v$ because $v$ is pure. A length-preserving linear map of three-dimensional space that fixes the origin is either a rotation or a reflection, and connectedness rules out the reflection.7Such a map has determinant $+1$ (a rotation) or $-1$ (a reflection), with nothing in between allowed. The map $v \mapsto q v \bar q$ depends continuously on $q$, and at $q = 1$ it is the identity, whose determinant is $+1$. The unit quaternions form the three-dimensional sphere $S^3$, which is connected, so as $q$ ranges over it the determinant cannot jump from $+1$ to $-1$ without passing through some forbidden value in between. It stays $+1$ throughout, and every sandwich by a unit quaternion is a rotation. The only question left is which rotation, and the quaternion announces its axis and angle directly. Write the unit quaternion in a form that should look suspiciously like Euler’s formula,

$$ q = \cos\tfrac{\theta}{2} + \sin\tfrac{\theta}{2}\,\hat{\mathbf{n}}, $$

where $\hat{\mathbf{n}} = n_1 i + n_2 j + n_3 k$ is a pure unit quaternion built from a unit vector, so that $\hat{\mathbf{n}}^2 = -1$ by the sphere-of-square-roots fact from earlier. Two short calculations settle exactly what the sandwich does, and neither asks you to take anything on faith.

First, the axis does not move. Because $\hat{\mathbf{n}}^2 = -1$, the quaternion $q$ commutes with $\hat{\mathbf{n}}$:

$$ q\,\hat{\mathbf{n}} = \left(\cos\tfrac{\theta}{2} + \sin\tfrac{\theta}{2}\,\hat{\mathbf{n}}\right)\hat{\mathbf{n}} = \cos\tfrac{\theta}{2}\,\hat{\mathbf{n}} – \sin\tfrac{\theta}{2} = \hat{\mathbf{n}}\left(\cos\tfrac{\theta}{2} + \sin\tfrac{\theta}{2}\,\hat{\mathbf{n}}\right) = \hat{\mathbf{n}}\,q, $$

so $q\,\hat{\mathbf{n}}\,\bar q = \hat{\mathbf{n}}\,q\bar q = \hat{\mathbf{n}}$. A vector pointing along $\hat{\mathbf{n}}$ is left precisely where it was, which is what it means for $\hat{\mathbf{n}}$ to be the axis.

Second, and here the half-angle finally explains itself, take a vector $\mathbf{v}$ perpendicular to the axis and watch it turn. Perpendicularity means $\hat{\mathbf{n}}\cdot\mathbf{v} = 0$, so by the product formula $\hat{\mathbf{n}}\mathbf{v} = \hat{\mathbf{n}}\times\mathbf{v}$ while $\mathbf{v}\hat{\mathbf{n}} = \mathbf{v}\times\hat{\mathbf{n}} = -\hat{\mathbf{n}}\times\mathbf{v}$; the two anticommute, $\mathbf{v}\hat{\mathbf{n}} = -\hat{\mathbf{n}}\mathbf{v}$, and a short step gives $\hat{\mathbf{n}}\mathbf{v}\hat{\mathbf{n}} = -\hat{\mathbf{n}}^2\mathbf{v} = \mathbf{v}$. Abbreviate $c = \cos\tfrac{\theta}{2}$ and $s = \sin\tfrac{\theta}{2}$ and expand the sandwich term by term:

$$ q\mathbf{v}\bar q = (c + s\hat{\mathbf{n}})\,\mathbf{v}\,(c – s\hat{\mathbf{n}}) = c^2\mathbf{v} – cs\,\mathbf{v}\hat{\mathbf{n}} + cs\,\hat{\mathbf{n}}\mathbf{v} – s^2\,\hat{\mathbf{n}}\mathbf{v}\hat{\mathbf{n}}. $$

Feed in the two facts just found, $\mathbf{v}\hat{\mathbf{n}} = -\hat{\mathbf{n}}\mathbf{v}$ and $\hat{\mathbf{n}}\mathbf{v}\hat{\mathbf{n}} = \mathbf{v}$. The two middle terms now add instead of cancelling, and the last becomes $-s^2\mathbf{v}$:

$$ q\mathbf{v}\bar q = (c^2 – s^2)\,\mathbf{v} + 2cs\,(\hat{\mathbf{n}}\mathbf{v}) = \cos\theta\,\mathbf{v} + \sin\theta\,(\hat{\mathbf{n}}\times\mathbf{v}), $$

where the double-angle identities $c^2 – s^2 = \cos\theta$ and $2cs = \sin\theta$ did the folding, and $\hat{\mathbf{n}}\mathbf{v} = \hat{\mathbf{n}}\times\mathbf{v}$ because the two are perpendicular. Read the answer. The vector $\mathbf{v}$ has become $\cos\theta\,\mathbf{v} + \sin\theta\,(\hat{\mathbf{n}}\times\mathbf{v})$, which is $\mathbf{v}$ swung through the angle $\theta$ inside the plane perpendicular to the axis, with $\hat{\mathbf{n}}\times\mathbf{v}$ supplying the perpendicular direction it swings toward. That is a rotation by $\theta$, no more and no less. The half-angle is now no mystery at all: the sandwich touches $\mathbf{v}$ with $q$ on the left and $\bar q$ on the right, so the angle stored in $q$ is spent twice, and the double-angle formula is precisely the arithmetic of spending it twice. A general vector splits into a piece along the axis, which we showed is fixed, and a piece perpendicular to it, which turns by $\theta$; together they carry the whole vector through $\theta$ about $\hat{\mathbf{n}}$.8The result $q\mathbf{v}\bar q = \cos\theta\,\mathbf{v} + \sin\theta\,(\hat{\mathbf{n}}\times\mathbf{v})$ for $\mathbf{v}\perp\hat{\mathbf{n}}$, extended to a general vector by adding the fixed axial part, is Rodrigues’ rotation formula, usually derived with far more trigonometry. The quaternion sandwich produces it in four lines. A unit quaternion storing half the angle for a two-sided action is the same bookkeeping as a square root storing half an exponent; both act, in effect, on both sides.

The general formula is easy to trust once you have watched it collapse to numbers. Let me rotate the $x$-axis by ninety degrees about the $z$-axis, which should carry it onto the $y$-axis. Take $v = i$, the $x$-direction, and $\hat{\mathbf{n}} = k$, the $z$-direction, with $\theta = 90^\circ$, so the quaternion uses half of that, forty-five degrees:

$$ q = \cos 45^\circ + \sin 45^\circ\, k = \tfrac{1}{\sqrt 2}\,(1 + k), \qquad \bar q = \tfrac{1}{\sqrt 2}\,(1 – k). $$

Now grind out the sandwich, using $ki = j$ and $ik = -j$ from the table. First the left product,

$$ q v = \tfrac{1}{\sqrt 2}(1+k)\,i = \tfrac{1}{\sqrt 2}\,(i + ki) = \tfrac{1}{\sqrt 2}\,(i + j), $$

and then the right,

$$ q v \bar q = \tfrac{1}{2}(i+j)(1-k) = \tfrac{1}{2}\,(i – ik + j – jk) = \tfrac{1}{2}\,(i + j + j – i) = j. $$

The $x$-axis lands exactly on the $y$-axis. The forty-five degrees inside the quaternion produced a ninety-degree turn in space, and the arithmetic that got us there used nothing but the multiplication table carved on the bridge.

v v′ θ v′ = q v q⁻¹,  q = cos(θ/2) + sin(θ/2) n̂

Figure 3: The sandwich product $q\,v\,\bar q$ turns the vector $v$ about the axis $\hat{\mathbf{n}}$ through the angle $\theta$, while the quaternion itself stores only half that angle. The dashed ellipse is the circle $v$ sweeps out as it rotates around the axis.

This construction also settles the demand we carried over from the complex numbers, that composing rotations should amount to multiplying the numbers that represent them. Rotate first by $q_1$ and then by $q_2$, and stack the sandwiches:

$$ v \;\longmapsto\; q_2\,(q_1 v \bar q_1)\,\bar q_2 = (q_2 q_1)\,v\,(\bar q_1 \bar q_2) = (q_2 q_1)\,v\,\overline{(q_2 q_1)}, $$

where the last step is the order-reversal $\bar q_1 \bar q_2 = \overline{q_2 q_1}$ we proved earlier. The composite rotation is a single sandwich by the product $q_2 q_1$. Composition of rotations is multiplication of quaternions, just as it was for complex numbers, with the one difference that is now the entire point: quaternion multiplication does not commute, so $q_2 q_1 \ne q_1 q_2$ in general, and neither do the rotations. The book that landed in two different places depending on the order of two turns, back in the vectors post, was showing you the noncommutativity of this product before we had the product to show. Associativity does survive, since quaternion multiplication is associative, so long chains of rotations compose without any ambiguity about grouping; only their order matters. That is exactly the structure that rotations of space actually have.

The half-angle has a consequence that reaches straight into physics. Because the quaternion holds $\theta/2$, turning all the way around by $\theta = 360^\circ$ makes the quaternion carry $180^\circ$, at which point $\cos 180^\circ + \sin 180^\circ\,\hat{\mathbf{n}} = -1$. The quaternion has become $-1$, not $+1$, even though space itself has returned exactly to where it started. You have to go around twice, a full $720^\circ$, before the quaternion comes back to $+1$. Two different quaternions, $q$ and $-q$, describe the very same rotation of space, so the map from quaternions to rotations is two-to-one.9This is the famous double cover, written $SU(2) \to SO(3)$ in the language of groups. It is not a defect of the description; it is a real feature of rotations that only becomes visible once you have the right language. The plate trick, or Dirac’s belt trick, makes it physical: hold a belt buckle fixed, rotate the free end by $360^\circ$, and the belt is twisted; rotate by another $360^\circ$, the same way, and the twist comes out. Space returns to itself after $360^\circ$, but the connection to its surroundings does not, and needs $720^\circ$. In quantum mechanics this is exactly the behavior of a spin-$\tfrac12$ particle, whose wavefunction changes sign under a single full rotation and only returns after two, and Hamilton’s $i$, $j$, $k$ are, up to a factor of $i$, the Pauli matrices that describe it. Space forgets the difference between a $360^\circ$ turn and no turn at all; the quaternions remember it. That memory is not an artifact. It is the same double-valuedness that a spinning electron carries in quantum mechanics, and quaternions saw it a century before anyone knew to look for it.

Now the promise I made in the middle of the post, to name the two pieces of the pure-quaternion product properly once they could land with full force. Recall what we found the moment we first multiplied two pure quaternions, two ordinary vectors dressed in $i, j, k$:

$$ \mathbf{u}\,\mathbf{v} = -\,\mathbf{u}\cdot\mathbf{v} + \mathbf{u}\times\mathbf{v}. $$

The scalar part is minus the dot product, the very inner product the vectors post built an entire discussion around. The vector part is the cross product, component for component. Both of the products that structure all of three-dimensional vector algebra are already inside quaternion multiplication, the scalar shadow and the vector shadow of one operation, and we have been using them without comment for the whole second half of this post: they were in the norm, in the order-reversal proof, in the rotation formula. This is not an analogy. It is the historical source. The dot product and the cross product were carved out of exactly this quaternion product by Gibbs and Heaviside in the 1880s, who kept the two halves and discarded the quaternion framework around them, which is why students meet the dot and cross products today as two separate definitions dropped from the sky rather than as the scalar and vector shadows of a single multiplication.10I mentioned in the vectors post that Gibbs and Heaviside carved modern vector analysis out of Hamilton’s quaternions; this identity is precisely what they carved. There was a genuine and rather heated dispute in the 1890s between the quaternion loyalists, led by Peter Tait, and the vector-analysis reformers, over which language physics should speak. The reformers won for practical purposes, and the dot and cross products became standard, but the quaternions never really left. They returned through the back door in quantum mechanics as the algebra of spin, and through the front door in computer graphics, robotics, and spacecraft attitude control, where a rotation stored as a single unit quaternion avoids the gimbal lock and interpolation headaches that plague rotation matrices and Euler angles. Your friend with the computer science PhD almost certainly met them there.

That is where my weekend of reviewing left me, and where I will leave you. Complex numbers were not the end of the number line rather only a rung on it. One dimension up, at the price of commutativity, the quaternions bought us the ability to turn objects in space, a second copy of the rotation miracle that complex numbers perform in the plane, and they folded the dot and cross products into a single act of multiplication as a kind of parting gift. The ladder climbs one rung further, to the octonions, at the price of associativity, and then, by a theorem, it stops. Not because no one has been clever enough to continue it, but because the mathematics has been shown to have no further rung to offer. Some staircases really do have a top, and it is a strange comfort to know exactly where it is.

References and Footnotes

  • 1
    This is the Brahmagupta–Fibonacci identity, known in essence to Diophantus. It is exactly the statement that the complex norm is multiplicative, $|w z|^2 = |w|^2 |z|^2$, written out in real coordinates with $w = a+bi$ and $z = c+di$. Keep an eye on the number of squares: two here, and the question of whether the same trick works for three and for four squares is the whole story below. ↩︎
  • 2
    Hamilton was searching for what he called “triples,” three-dimensional numbers $a + bi + cj$ that could be added and, crucially, multiplied while preserving length, so that unit triples would rotate space the way unit complex numbers rotate the plane. He worked at it for years; there is a much-repeated story that his children would ask him at breakfast whether he could multiply triples yet, and he would have to answer that he could still only add and subtract them. ↩︎
  • 3
    Take $3 = 1^2+1^2+1^2$ and $5 = 0^2+1^2+2^2$, both sums of three squares. Their product is $15$, and $15$ is not a sum of three squares: by Legendre’s three-square theorem a positive integer fails to be a sum of three squares exactly when it has the form $4^a(8b+7)$, and $15 = 8(1)+7$. So the three-square analogue of the identity in the previous footnote is simply false, and a length-preserving multiplication on triples cannot exist. The obstruction is not a failure of imagination; it is a theorem. ↩︎
  • 4
    This order reversal is the same phenomenon as $(AB)^\top = B^\top A^\top$ for matrices, and it is not a coincidence: quaternions can be represented by $2\times 2$ complex matrices, with the conjugate becoming the conjugate transpose. It also rhymes with the covariant bookkeeping from the vectors post, where a transforming quantity reverses order relative to the basis. Order reversal under an involution is a recurring signature of noncommutative structures. ↩︎
  • 5
    Euler’s four-square identity, that a product of two sums of four squares is again a sum of four squares, predates Hamilton by nearly a century; he found it in 1748 while working on Lagrange’s four-square theorem. Hamilton’s quaternions are, in a sense, the hidden reason the identity is true: it is just $N(pq) = N(p)N(q)$ in coordinates. The pattern of which dimensions admit such an identity is severe. Sums-of-squares multiply cleanly only in dimensions $1$, $2$, $4$, and $8$, a result of Hurwitz, and the four cases are the real numbers, the complex numbers, the quaternions, and the octonions. There is no five-dimensional or seven-dimensional number system of this kind, and there never will be. ↩︎
  • 6
    This is Frobenius’s theorem, 1877: the only finite-dimensional associative division algebras over the reals are $\mathbb{R}$, $\mathbb{C}$, and $\mathbb{H}$ (the $\mathbb{H}$ is for Hamilton). Each rung of the ladder costs you a property. Passing from $\mathbb{R}$ to $\mathbb{C}$ you lose ordering: there is no consistent way to say one complex number is less than another. Passing from $\mathbb{C}$ to $\mathbb{H}$ you lose commutativity, as we have seen. If you are willing to give up associativity as well you can reach one dimension further, to the eight-dimensional octonions $\mathbb{O}$, and there the ladder ends for good. This is the same “the language runs out” story that ran through the earlier posts, now with a precise list of exactly how far it can possibly go. ↩︎
  • 7
    Such a map has determinant $+1$ (a rotation) or $-1$ (a reflection), with nothing in between allowed. The map $v \mapsto q v \bar q$ depends continuously on $q$, and at $q = 1$ it is the identity, whose determinant is $+1$. The unit quaternions form the three-dimensional sphere $S^3$, which is connected, so as $q$ ranges over it the determinant cannot jump from $+1$ to $-1$ without passing through some forbidden value in between. It stays $+1$ throughout, and every sandwich by a unit quaternion is a rotation. ↩︎
  • 8
    The result $q\mathbf{v}\bar q = \cos\theta\,\mathbf{v} + \sin\theta\,(\hat{\mathbf{n}}\times\mathbf{v})$ for $\mathbf{v}\perp\hat{\mathbf{n}}$, extended to a general vector by adding the fixed axial part, is Rodrigues’ rotation formula, usually derived with far more trigonometry. The quaternion sandwich produces it in four lines. A unit quaternion storing half the angle for a two-sided action is the same bookkeeping as a square root storing half an exponent; both act, in effect, on both sides. ↩︎
  • 9
    This is the famous double cover, written $SU(2) \to SO(3)$ in the language of groups. It is not a defect of the description; it is a real feature of rotations that only becomes visible once you have the right language. The plate trick, or Dirac’s belt trick, makes it physical: hold a belt buckle fixed, rotate the free end by $360^\circ$, and the belt is twisted; rotate by another $360^\circ$, the same way, and the twist comes out. Space returns to itself after $360^\circ$, but the connection to its surroundings does not, and needs $720^\circ$. In quantum mechanics this is exactly the behavior of a spin-$\tfrac12$ particle, whose wavefunction changes sign under a single full rotation and only returns after two, and Hamilton’s $i$, $j$, $k$ are, up to a factor of $i$, the Pauli matrices that describe it. ↩︎
  • 10
    I mentioned in the vectors post that Gibbs and Heaviside carved modern vector analysis out of Hamilton’s quaternions; this identity is precisely what they carved. There was a genuine and rather heated dispute in the 1890s between the quaternion loyalists, led by Peter Tait, and the vector-analysis reformers, over which language physics should speak. The reformers won for practical purposes, and the dot and cross products became standard, but the quaternions never really left. They returned through the back door in quantum mechanics as the algebra of spin, and through the front door in computer graphics, robotics, and spacecraft attitude control, where a rotation stored as a single unit quaternion avoids the gimbal lock and interpolation headaches that plague rotation matrices and Euler angles. Your friend with the computer science PhD almost certainly met them there. ↩︎
Posted in Expository | Tagged | Leave a comment

So, what is a vector?

In my first year of undergraduate physics, one of my professors walked into the lecture hall, placed his notes on the desk, looked around the room, and asked a question that, at first, seemed almost too easy.

“What is a vector?”

I immediately raised my hand.

“A vector is a quantity that has both magnitude and direction.”

He smiled.

It wasn’t the kind of smile that tells you you’ve answered correctly, nor was it the kind that says you’ve made a mistake. It was the smile of someone who knows there’s a much more interesting answer waiting just around the corner. At the time, I didn’t think much of it. I had given the definition that every mathematics and physics student learns in high school. Surely that was the answer.

After a short pause, he looked at the class and said, “That’s how we often describe vectors. But it isn’t what a vector really is.”

I remember sitting there wondering what he meant. If a vector isn’t simply something with a magnitude and a direction, then what exactly is it? Every example I had ever seen was an arrow. Every textbook drew arrows. Every problem involved arrows. Wasn’t that enough?

Instead of answering the question directly, he told us a story.

He said that mathematics rarely invents new objects for the sake of inventing them. New mathematical ideas almost always appear because the old ones stop being sufficient. People encounter a problem, struggle with it using the mathematics they already have, and eventually realize that their language is no longer rich enough to describe what they are seeing. The new mathematics is born not from imagination, but from necessity.

He asked us to think about the numbers we learn as children. We begin with the natural numbers: $1, 2, 3, \ldots$. For a while, they seem to answer every question we ask. If there are three apples on the table and someone gives you two more, you have five apples. If there are ten students in a classroom and four leave, six remain. The arithmetic is simple, and the natural numbers seem perfectly adequate.

Eventually, however, we encounter questions they cannot answer. What is $3-5$? Within the natural numbers, the expression has no answer. It isn’t that the calculation is wrong; rather, our number system is too limited to express the result. Instead of declaring the question meaningless, mathematics extends the number system. The integers are introduced, and suddenly the answer $-2$ makes perfect sense.

The same story repeats itself. Rational numbers appear because integers cannot always describe parts of a whole. Real numbers appear because rational numbers cannot represent lengths such as $\sqrt{2}$.1The ancient Greek discovery that the diagonal of a unit square cannot be expressed as a ratio of two integers was one of the earliest demonstrations that the rational numbers are incomplete. Complex numbers appear because some perfectly reasonable equations, such as $x^2+1=0$, simply have no real solutions. Every time mathematics grows, it grows because reality, or mathematics itself, demands a richer language.

“Vectors,” my professor said, “were born for exactly the same reason.”

At first, that statement surprised me. I had always treated vectors as though they were just another topic in geometry, sitting alongside triangles, circles, and coordinate systems. It had never occurred to me that they were invented to solve a problem. But once he said it, the obvious question became unavoidable.

What problem?

To answer that question, we have to forget everything we think we know about vectors for a little while. Forget the arrows. Forget the coordinate axes. Forget the ordered pairs like $(3,4)$. Those are all ways of representing vectors, but they are not where the idea begins.

Instead, let us go back to a time before the word “vector” even existed and ask a much simpler question: what kinds of quantities can ordinary numbers actually describe?

Some quantities surrender completely to a single number. The temperature of this room. The mass of an electron. The time on the clock at the back of the lecture hall. Tell me the number and you have told me everything; there is nothing left over. And these numbers have a second property, so obvious that it hides. Everyone in the room agrees on them, no matter which way they happen to be facing. If I turn my chair forty degrees, the room does not get warmer. Quantities like this are called scalars, and for most of history they were simply called quantities, because nobody imagined there could be another kind.

Then try this one. I leave my house and walk five kilometers. Where am I?

You cannot say. I could be five kilometers north, five kilometers southeast, anywhere on an entire circle of possibilities. The number five is not wrong. It is incomplete. A displacement has a size, but the size does not exhaust it; something is left over after the number has done its work, and that something is which way. The same incompleteness infects velocity, force, acceleration, the pull of a magnet on a compass needle. By the time people were seriously composing forces, adding one push to another and asking for the net effect, the problem could no longer be postponed. A single number per quantity was a language that had run out.2The parallelogram rule for composing forces was stated clearly by Simon Stevin in 1586 and appears as a corollary in Newton’s Principia, so the geometric content of vector addition is centuries older than the word. The word itself comes from William Rowan Hamilton in the 1840s, from the Latin vehere, to carry, as part of his quaternions; the modern, stripped-down vector analysis was carved out of the quaternion system by Gibbs and Heaviside toward the end of the nineteenth century. It is a pleasing echo of the professor’s story: Hamilton was searching for a three-dimensional analogue of the complex numbers, which had themselves been invented when the real numbers ran out.

The tempting fix is to use more numbers. Lay down two reference directions, east and north, and report the walk as a pair: three kilometers east, four kilometers north. Write it $(3, 4)$. Two numbers instead of one, the ambiguity is gone, and you know exactly where I ended up. For a while this feels like the whole answer. It is the answer your textbook’s ordered pairs are quietly assuming.

But a friend of mine surveys the same walk, and she has laid her reference directions differently. Her “east” points along the road outside her house, which runs at an angle to mine. She measures carefully and reports the same walk as roughly $(4.6, 2.0)$. Same starting point, same ending point, same physical walk through the same physical mud. Different numbers.

If a vector were a pair of numbers, we would now have a contradiction: one object equal to two different things. The walk cannot be $(3,4)$ and also $(4.6, 2.0)$. So the walk is not the pair of numbers. The pair of numbers is a description of the walk in a particular language, and the language is the choice of axes. Her language differs from mine, so her words differ from mine, and the walk itself does not care about either of us.

x y x′ y′ θ v vx vy v′x v′y

Figure 1: One walk, two languages. The same displacement $v$ read against two sets of axes rotated by an angle $\theta$. The finely dashed drop lines give the components $(v_x, v_y)$ in the first frame; the coarsely dashed lines give $(v’_x, v’_y)$ in the second. The numbers differ. The walk does not.

This is the moment to slow down, because everything that follows grows out of it. The numbers change with the observer; the object does not. If we want the mathematics to describe the object rather than the observer, we have to find out what, in all this shifting of numbers, stays put.

And here the situation is far better than it first appears, because the two descriptions are not unrelated. If I know the single angle $\theta$ between my axes and hers, I can compute her numbers from mine, exactly, without ever leaving my desk. Suppose my components are $v_x$ and $v_y$, and write $|v|$ for the length of the displacement and $\varphi$ for the angle it makes with my east axis, so that $v_x = |v|\cos\varphi$ and $v_y = |v|\sin\varphi$. Her east axis sits at angle $\theta$ from mine, which means the same displacement makes angle $\varphi – \theta$ with her axis, and her components are $|v|\cos(\varphi-\theta)$ and $|v|\sin(\varphi-\theta)$. Expanding the angle differences3The two identities used here are $\cos(\varphi-\theta) = \cos\varphi\cos\theta + \sin\varphi\sin\theta$ and $\sin(\varphi-\theta) = \sin\varphi\cos\theta – \cos\varphi\sin\theta$. They are the whole trigonometric input of this article; everything else is substitution. and substituting $v_x$ and $v_y$ back in gives

$$ v’_x = v_x\cos\theta + v_y\sin\theta, \qquad v’_y = -\,v_x\sin\theta + v_y\cos\theta. $$

Check it on the walk. My components were $(3,4)$, and suppose her axes are rotated by $\theta = 30^\circ$, so $\cos\theta \approx 0.866$ and $\sin\theta = 0.5$. Then $v’_x = 3(0.866) + 4(0.5) \approx 4.60$ and $v’_y = -3(0.5) + 4(0.866) \approx 1.96$. Those are her numbers, the ones she measured in the field, recovered from mine by pure computation. The two languages are connected by a dictionary, and the dictionary is this pair of formulas.4A bookkeeping note: throughout, the walk is held fixed and the axes rotate, which is called a passive transformation. One can instead hold the axes fixed and rotate the object by $\theta$, an active transformation; the formulas are the same with $\theta$ replaced by $-\theta$. Both conventions are common, which is why signs in other books may differ from mine.

But where does the dictionary itself come from? The trigonometric derivation shows that it is correct. It does not yet show that it is forced. So let me derive it a second time, from a different starting point, because the second derivation is the one that says what a vector actually is. Introduce symbols for my two reference directions themselves: let $e_x$ and $e_y$ be the unit displacements, one kilometer east and one kilometer north. Then my report $(3,4)$ was shorthand for an equation about objects,

$$ v = 3\,e_x + 4\,e_y, $$

three copies of the first reference displacement followed by four of the second, and in general $v = v_x\,e_x + v_y\,e_y$. My friend has her own reference displacements, $e’_x$ and $e’_y$. Since hers are ordinary displacements like any others, I can describe them in my language. Her first axis points at angle $\theta$ from mine, so

$$ e’_x = \cos\theta\, e_x + \sin\theta\, e_y, \qquad e’_y = -\sin\theta\, e_x + \cos\theta\, e_y. $$

The crucial demand comes next, and it is a physical demand, not a mathematical one: she and I are describing the same walk. Whatever components she uses must reassemble the identical object,

$$ v_x\, e_x + v_y\, e_y = v’_x\, e’_x + v’_y\, e’_y. $$

Substitute her basis vectors and collect my $e_x$ and $e_y$:

$$ v_x\, e_x + v_y\, e_y = (v’_x\cos\theta – v’_y\sin\theta)\, e_x + (v’_x\sin\theta + v’_y\cos\theta)\, e_y. $$

Two walks reported in the same language, east part and north part, are the same walk only if the east parts agree and the north parts agree.5I am borrowing a fact here: that a displacement in the plane determines its east part and north part uniquely. Geometrically this is clear enough. In the general setting it is exactly the uniqueness of representation in a basis, which I prove properly once bases are defined below, at which point this step stops being a loan. So $v_x = v’_x\cos\theta – v’_y\sin\theta$ and $v_y = v’_x\sin\theta + v’_y\cos\theta$. Solve the pair for the primed components, multiplying the first by $\cos\theta$ and the second by $\sin\theta$, adding, and using $\cos^2\theta + \sin^2\theta = 1$, and out comes precisely the dictionary we derived with angles:

$$ v’_x = v_x\cos\theta + v_y\sin\theta, \qquad v’_y = -\,v_x\sin\theta + v_y\cos\theta. $$

Nothing about the law was a convention. Once you grant that the two of us describe one object, the law is the only possibility.

Two rules are now on the table, one telling how the basis vectors translate and one telling how the components translate, and the relationship between them is the point this whole article has been walking toward. They are inverses of each other. For pure rotations the inverse of the translation array happens to coincide with its transpose, which disguises the opposition, so let me expose it with a cruder change of language. Suppose my friend does not rotate her axes at all but merely adopts a longer measuring stick, so that her reference displacement is twice mine, $e’_x = 2\,e_x$. The walk is the walk: $v = v_x\,e_x = v’_x\,e’_x = 2\,v’_x\,e_x$, so $v’_x = v_x/2$. Her basis vector doubled and her component halved. The two always pull in opposite directions, and what their pulling multiplies out to, the object, sits perfectly still. A vector is a tug of war between basis and components that ends, in every language, in exactly the same draw.6This opposition has standard names: components that transform inversely to the basis are called contravariant, and objects that transform along with the basis, as the gradient does, are called covariant. The raised and lowered indices of the previous post are precisely this bookkeeping, $x^\mu$ upstairs for contravariant components and $\partial_\mu$ downstairs for covariant objects, and the convention exists so that a paired upper and lower index always assembles a language-independent quantity.

Now ask what the dictionary preserves. Take her components, square them, and add:

$$ v_x’^2 + v_y’^2 = (v_x\cos\theta + v_y\sin\theta)^2 + (v_y\cos\theta – v_x\sin\theta)^2. $$

Multiply everything out. The cross terms $2v_xv_y\cos\theta\sin\theta$ appear once with a plus sign and once with a minus sign, and cancel. What remains is $v_x^2$ and $v_y^2$, each multiplied by $\cos^2\theta + \sin^2\theta$, which is $1$. So

$$ v_x’^2 + v_y’^2 = v_x^2 + v_y^2. $$

On the numbers: $3^2 + 4^2 = 25$, and $4.60^2 + 1.96^2 = 21.2 + 3.8 = 25.0$. Twenty-five both times. The individual components shifted, but this particular combination of them refused to move, and it refuses for every angle $\theta$, not just thirty degrees. Its square root is the length of the displacement. So the length is not part of the description; it is a fact about the object, the same in every language. Direction behaves more subtly: the angle my walk makes with my east axis is $\varphi$, while the angle it makes with hers is $\varphi – \theta$, so “the direction” as an angle relative to axes is language-dependent. But take two walks and ask for the angle between them. Both their angles shift by the same $\theta$, and the difference survives untouched. Lengths and mutual angles belong to the objects. Components and axis angles belong to the description.

You can feel the high-school definition being rebuilt on better foundations. “Magnitude and direction” was gesturing at exactly this: the invariant content, the part of the description that all observers share. It just said it sloppily, as though magnitude and direction were the definition rather than the residue that survives translation.

So here is a first serious answer to the professor’s question, the one a physicist tends to give. A vector is not a list of numbers. A vector is a quantity whose lists of numbers, written down by different observers, are tied to one another by exactly the dictionary above. In two dimensions the dictionary is the pair of formulas we derived; in $n$ dimensions it takes the form

$$ v’_i = \sum_{j} R_{ij}\, v_j, $$

where the array $R_{ij}$ encodes the relative orientation of the two frames.7For rotations, $R$ is an orthogonal matrix: its rows are the new axis directions written in the old language, it satisfies $\sum_i R_{ij}R_{ik} = \delta_{jk}$, where $\delta_{jk}$ equals $1$ when $j=k$ and $0$ otherwise. This orthogonality is forced rather than assumed: demand that $\sum_i v_i’^2 = \sum_j v_j^2$ for every vector, substitute the transformation law to get $\sum_{j,k}\left(\sum_i R_{ij}R_{ik}\right)v_j v_k$, and the only way the result can equal $\sum_j v_j^2$ for all choices of the $v_j$ is for the bracketed sum to be $\delta_{jk}$. The two-dimensional cancellation we did by hand is the $n=2$ case. Requiring in addition that the frame is not mirror-reflected fixes $\det R = +1$. Physicists compress this into a slogan: a vector is anything that transforms like a vector. Said quickly, it sounds circular. It is not. It is a behavioral definition, like defining a key as anything that opens the lock. It does not tell you what the object is made of; it tells you how to test whether a given candidate qualifies.

One consequence of the law deserves its own sentence, because the whole of mechanics leans on it. The dictionary is linear: it never squares a component or multiplies two of them together. So if $u$ and $v$ both pass the test, their sum passes automatically,

$$ (u+v)’_i = \sum_{j} R_{ij}\,(u_j + v_j) = u’_i + v’_i, $$

and likewise any multiple $a\,u$. Sums of vectors are vectors, and rescaled vectors are vectors. That is the license behind composing forces at all: add two forces tip to tail and the result is guaranteed to be a quantity that every observer’s dictionary handles correctly.

The law also has to pass a consistency requirement involving three observers instead of two, and it is worth checking rather than trusting. Bring in a third surveyor whose axes sit at angle $\theta_2$ from my friend’s, and therefore at $\theta_1 + \theta_2$ from mine. There are now two ways to compute his components from mine: translate directly, using the dictionary with angle $\theta_1 + \theta_2$, or translate in two hops, first into her frame with $\theta_1$ and then onward with $\theta_2$. If the two routes disagreed, the whole scheme would collapse; a component would depend on which intermediate observers you consulted along the way. So take the two-hop route and compute, substituting her components into his dictionary:

$$ v”_x = v’_x\cos\theta_2 + v’_y\sin\theta_2 = v_x\left(\cos\theta_1\cos\theta_2 – \sin\theta_1\sin\theta_2\right) + v_y\left(\sin\theta_1\cos\theta_2 + \cos\theta_1\sin\theta_2\right). $$

The parentheses are the angle-addition identities run in reverse, the subtraction identities from the first derivation with $\theta$ replaced by $-\theta$, and they collapse:

$$ v”_x = v_x\cos(\theta_1+\theta_2) + v_y\sin(\theta_1+\theta_2). $$

That is exactly the direct dictionary, and the $y$ component checks the same way. Two hops equal one hop. The dictionaries between all possible observers mesh into a single consistent system,8In modern language: rotations compose to give rotations, every rotation has an inverse, and doing nothing counts as a rotation, so the collection of all rotations forms a group, written $SO(2)$ in the plane and $SO(n)$ in $n$ dimensions. “Transforms like a vector” is, said fully, a statement about how an object responds to this entire group at once, and group theory is the systematic study of such consistency requirements. It becomes one of the load-bearing structures of modern physics. and a vector is an object that speaks every one of those languages at once, coherently.

And the test has teeth. Take three perfectly good numbers measured at a point in this room: the temperature, the pressure, and the humidity. Stack them into a triple, $(T, p, h)$. Three numbers, arranged in a column; you could even draw the triple as an arrow in an abstract space if you felt like it. Now rotate your axes. Nothing happens. The temperature does not care which way you face, and neither do the other two; each is a scalar, so the triple you write down after the rotation is the same triple you wrote before. But the components of a genuine vector mix under rotation, each new component a blend of all the old ones, as our formulas demand. This triple refuses to mix. It fails the test. Three numbers in a column are not a vector, any more than three people in an elevator are a family.

The demolition works from the other side too, and this one surprised me more. There are things in physics that possess a perfectly good magnitude and a perfectly good direction and are still not vectors. A finite rotation is one: it has an axis, which is a direction, and an angle, which is a magnitude. Try to add two of them. Hold a book flat in front of you, rotate it $90^\circ$ about the vertical axis and then $90^\circ$ about the horizontal axis pointing away from you, and note where it ends up. Restore it, and apply the same two rotations in the opposite order. The book lands in a visibly different position. But vector addition does not care about order; $a + b$ and $b + a$ are the same walk. Finite rotations, magnitude and direction and all, break the rules of the game.9The failure is a finite-angle effect. Rotations through very small angles do commute, up to corrections of second order in the angles, and this is why angular velocity, which is built from infinitesimal rotations, is a legitimate vector while “the rotation itself” is not. There is a further subtlety I am deliberately postponing: quantities like angular velocity behave correctly under rotations but pick up an extra sign under mirror reflection, which is why they are called axial vectors or pseudovectors. That distinction deserves its own discussion. So magnitude-and-direction is neither sufficient to make something a vector, as the rotations show, nor, as we are about to see, even necessary.

Because now the mathematician takes a turn, and the mathematician’s move is more radical than anything so far. Forget the transformation law for a moment and ask an operational question: in all our dealings with displacements, what did we ever actually do with them? Two things. We added them: walk $u$ and then walk $v$, and the net effect is a single walk written $u + v$, running from the start of the first to the tip of the second. And we scaled them: twice the walk, half the walk, the reverse of the walk. That is the complete inventory. Every construction we performed was built from adding and scaling. And the addition has a property worth drawing rather than asserting: it does not care about order.

u v v u u + v

Figure 2: The parallelogram law. Walking $u$ then $v$ (solid route) or $v$ then $u$ (dashed route) delivers you to the same corner, and the diagonal is the sum $u+v$. This closing of the quadrilateral is exactly the commutativity that finite rotations, with their perfectly good magnitudes and directions, fail to provide.

Both routes around the parallelogram, $u$ then $v$ along the lower flank, $v$ then $u$ along the upper, arrive at the same corner, so $u + v = v + u$. For displacements, the commutativity that finite rotations just failed is a closed quadrilateral you can check with a ruler.

So the mathematician defines: a vector space is any collection of objects that can be added to each other and multiplied by numbers, provided those two operations obey the familiar rules of arithmetic, and a vector is nothing more and nothing less than a member of such a collection.10The rules, stated compactly: addition is commutative and associative, there is a zero element, every element has a negative, scaling distributes over both vector addition and number addition, scaling is compatible with multiplication of numbers, and scaling by $1$ does nothing. Eight axioms in all. The “numbers” doing the scaling can be the real numbers or the complex numbers; quantum mechanics is built on complex vector spaces, and everything in this article goes through in either case. Read the definition twice and you will see what it does not say. Nothing about arrows. Nothing about magnitude. Nothing about direction. Nothing, even, about space. The definition kept the algebra of displacements and threw away the picture.

And the rules are not decoration. They prove things. Take the smallest possible example: what is a vector scaled by the number zero? For arrows the answer is obvious, but the definition never mentions arrows, so the answer must be squeezed out of the rules alone. Since $0 = 0 + 0$ as numbers, distributivity gives

$$ 0\,v = (0+0)\,v = 0\,v + 0\,v. $$

Add to both sides the negative of $0\,v$, which the rules guarantee exists. The left side becomes the zero element, the right side becomes $0\,v$, and so

$$ 0\,v = \mathbf{0}. $$

Scaling anything by zero gives the zero vector, in every vector space there is or ever will be, polynomials and quantum states included, and we proved it without knowing what the vectors are made of. That is what an axiomatic definition buys: theorems with unlimited jurisdiction.

The payoff for that austerity is enormous, because the definition now admits members the picture could never have suggested. Take polynomials. Add two of them,

$$ (1 + x^2) + (3x – x^2) = 1 + 3x, \qquad 2\,(1+x^2) = 2 + 2x^2, $$

and you can check, patiently, that all eight rules hold. Polynomials form a vector space. So the polynomial $1 + x^2$ is a vector, and here is the question that finishes off the high-school definition for good: what is its direction? The question is not hard. It is empty. Nothing in the space of polynomials points anywhere, and yet the addition and the scaling work flawlessly. Direction was never the essence. It was a feature of one famous example.

The same is true of functions at large, and this is where the abstraction starts paying rent in physics. In the previous post we derived the wave equation, $\Box\phi = 0$. Suppose $\phi_1$ and $\phi_2$ are both solutions. Then for any numbers $a$ and $b$,

$$ \Box\,(a\,\phi_1 + b\,\phi_2) = a\,\Box\phi_1 + b\,\Box\phi_2 = 0, $$

so the combination is a solution too. The solutions of the wave equation form a vector space. You have known this fact for years under a different name: superposition. When two ripples cross a pond and pass through each other, the water is performing vector addition in an infinite-dimensional function space. The principle of superposition is the statement that the solution set is a vector space, and it holds for exactly as long as the governing equation is linear.

At this point the two definitions, the physicist’s and the mathematician’s, look like they belong to different books. One is about dictionaries between observers; the other is about the arithmetic of abstract sets. The bridge between them is the idea of a basis, which has in fact been operating informally ever since the reference displacements $e_x$ and $e_y$ first appeared. It deserves to be defined with care rather than gestured at. A set of vectors $e_1, e_2, \ldots, e_n$ is called a basis for the space if it does two jobs at once. First, it must span: every vector $v$ in the space can be written as some combination

$$ v = v_1\, e_1 + v_2\, e_2 + \cdots + v_n\, e_n. $$

Second, it must be linearly independent, which means the reference vectors carry no internal redundancy: the only combination of them that produces the zero vector is the trivial one,

$$ c_1\, e_1 + c_2\, e_2 + \cdots + c_n\, e_n = \mathbf{0} \quad\Longrightarrow\quad c_1 = c_2 = \cdots = c_n = 0. $$

Spanning guarantees that a description exists. Independence guarantees something better, that the description is unique, and the proof takes three lines. Suppose some vector had two descriptions, $v = a_1 e_1 + \cdots + a_n e_n$ and also $v = b_1 e_1 + \cdots + b_n e_n$. Subtract them:

$$ (a_1 – b_1)\,e_1 + (a_2 – b_2)\,e_2 + \cdots + (a_n – b_n)\,e_n = \mathbf{0}. $$

Independence forces every coefficient to vanish, so $a_i = b_i$ for each $i$, and the two descriptions were the same description all along. This also repays the loan from earlier: matching the east parts and north parts of two equal walks was exactly this argument, run in the plane. The numbers $v_1, \ldots, v_n$ are the components of $v$, and now we can say precisely what they are: not properties of $v$ alone, but properties of the pair, the vector together with the chosen basis. Nor is any of this confined to arrows. The monomials $1, x, x^2, x^3, \ldots$ form a basis for the polynomials,11Independence here means: if $c_0 + c_1 x + \cdots + c_n x^n$ is the zero vector of this space, which is the polynomial equal to zero for every value of $x$, then all the $c_i$ vanish. Set $x=0$ and $c_0$ falls. Differentiate once, set $x=0$ again, and $c_1$ falls. Keep differentiating; each coefficient drops in turn. So the monomials carry no redundancy, and the components below are the unique ones. and in that basis the polynomial $1 + 3x$ from a few paragraphs ago has components $(1, 3, 0, 0, \ldots)$: one part constant, three parts $x$, nothing else. What looked like algebra was a coordinate representation all along.12A basis is far from unique, but a theorem guarantees that every basis of a given space contains the same number of elements, and that shared number is the dimension of the space. Displacements in a plane have dimension two, displacements in the room have dimension three, and the polynomials and the wave-equation solutions above have infinite dimension, which changes some technical details and none of the conceptual ones. My friend and I were not disagreeing about the walk. We had selected different bases, hers rotated relative to mine, and the transformation law we derived with so much trigonometry is nothing but the recipe for converting components with respect to one basis into components with respect to another. The physicist’s definition is the mathematician’s definition plus one nonnegotiable demand: nature never saw our axes, so no physical statement may depend on which basis we happened to pick. It is the same instinct that made us write the invariant volume $\sqrt{-g}\,d^4x$ in the previous post instead of the coordinate volume $d^4x$. The laws belong to the objects, not to the bookkeeping.

Once you see components as coordinates-with-respect-to-a-basis, a familiar tool from a completely different corner of physics turns out to be this same idea in disguise. Functions on an interval form a vector space, and the sines and cosines can serve as a basis for it. Writing a signal as a Fourier series is choosing that basis and reading off the components; the Fourier coefficients of a sound wave are, in the most literal sense, the components of a vector.13To make “component along a basis function” precise you need an inner product on the function space, and the natural one is an integral, $\langle f, g\rangle = \int f(x)\,g(x)\,dx$ over the interval. The sines and cosines are orthogonal with respect to it, which is what makes the coefficients easy to extract, one integral each. Infinite dimensions bring genuine subtleties about convergence that I am stepping over here, but the geometric picture, a function as a vector and its Fourier coefficients as components, is exactly right. Nobody would ask for the magnitude and direction of a violin note. Its vectorhood lies elsewhere.

Which raises the last question honestly: if direction is not essential, what exactly is the status of magnitude? The answer is that magnitude is optional equipment. To speak of lengths you must add a further gadget to the space, an inner product, the familiar dot product being the standard issue:

$$ u \cdot v = u_x v_x + u_y v_y. $$

Run it through the dictionary and watch the cancellation happen, this time in full. Her value of the product is

$$ u’\cdot v’ = (u_x\cos\theta + u_y\sin\theta)(v_x\cos\theta + v_y\sin\theta) + (u_y\cos\theta – u_x\sin\theta)(v_y\cos\theta – v_x\sin\theta). $$

Multiply out all four products. The terms carrying $\cos\theta\sin\theta$ arrive in matched pairs of opposite sign, $u_x v_y$ once with a plus and once with a minus, $u_y v_x$ the same, and annihilate. What survives is $u_x v_x$ and $u_y v_y$, each dressed with a factor of $\cos^2\theta + \sin^2\theta$:

$$ u’\cdot v’ = u_x v_x + u_y v_y = u\cdot v. $$

The dot product of two vectors comes out identical in every rotated frame, and the length computation we did earlier is the special case $u = v$. From the dot product you then recover everything the arrows promised, length as $|v| = \sqrt{v\cdot v}$ and the angle between two vectors through $u\cdot v = |u|\,|v|\cos\alpha$.14That last formula only defines an angle if the quantity $u\cdot v/(|u||v|)$ actually lies between $-1$ and $1$, and the guarantee is the Cauchy–Schwarz inequality, $(u\cdot v)^2 \le (u\cdot u)(v\cdot v)$. The proof is short enough to give in full. For any number $t$, the vector $u – t\,v$ has a non-negative dot product with itself, so $0 \le (u – t\,v)\cdot(u – t\,v) = u\cdot u – 2t\,(u\cdot v) + t^2\,(v\cdot v)$. The right side is a quadratic in $t$ that never dips below zero, so it cannot have two distinct real roots, so its discriminant obeys $4(u\cdot v)^2 – 4(u\cdot u)(v\cdot v) \le 0$, which is the inequality. Nothing about arrows was used, only the rules of the inner product, so the result holds in every inner product space, the function spaces of the previous footnote included. The definition of angle rests on a theorem, and now you have seen the theorem.

The inner product also settles a practical question we have been finessing: given a vector and a basis, how do you actually extract the components? Call a basis orthonormal if its vectors are mutually perpendicular and of unit length, $e_x\cdot e_x = e_y\cdot e_y = 1$ and $e_x\cdot e_y = 0$. Then dot the expansion $v = v_x\,e_x + v_y\,e_y$ with $e_x$ and let the orthonormality do the sorting:

$$ e_x\cdot v = v_x\,(e_x\cdot e_x) + v_y\,(e_x\cdot e_y) = v_x. $$

A component is a dot product with the corresponding basis vector, $v_x = e_x\cdot v$ and $v_y = e_y\cdot v$: to find how much of $v$ lies along a reference direction, project onto it. And this closes the loop left open with the violin note. Extracting a Fourier coefficient is this same projection, executed with the integral inner product in the space of functions,

$$ c_n = \langle f, e_n \rangle, $$

one coefficient per basis function, one integral each. The formula engineers use daily to decompose a signal is $v_x = e_x\cdot v$ wearing an integral sign.

A vector space with no inner product installed, meanwhile, is still a complete, functioning vector space; its members add, scale, form bases, transform. They simply have no lengths, because no one has defined any. Magnitude and direction, the two pillars of the high-school definition, turn out to be features of one optional attachment.

So, standing where my professor stood, here is how I would now answer his question, in layers. At the foundation, a vector is a member of a vector space: any object, arrow or polynomial or sound wave or quantum state, that can be added to its fellows and scaled by numbers under the usual rules. In physics we ask for one thing more: that the numbers describing the object in different frames be linked by the transformation law, because the object exists and our axes do not. And at the top, as a special case, sits the beloved picture, the arrow in two or three dimensions with an inner product installed, where magnitude and direction happen to capture everything. Running underneath all three layers is the mechanism this article has really been about: change the basis and the components change in the opposite way, so that their combination, the object itself, does not move. If one sentence is worth carrying out of the room, it is that one. The high-school definition is not wrong. It is a portrait of one member of the family, mistaken for the family itself.

I think that is what the smile meant. He knew the definition I gave him was the beginning of the story, and he knew the story was the same one mathematics always tells. The naturals ran out and gave us $-2$; the rationals ran out and gave us $\sqrt 2$; the reals ran out and gave us $i$; and one number per quantity ran out the moment somebody asked which way, and gave us this. Nor does the story stop here. There are quantities in physics that outgrow vectors in turn: the stress inside a loaded beam answers every direction you probe it with by handing back a force vector, and the machine doing the handing back is no longer a vector but a tensor. The metric $g_{\mu\nu}$ from the previous post is exactly such a machine.

The language runs out again. It always does. Vectors are enough to describe quantities with one directional character, but physics soon confronts objects that relate several directions at once. Stress is one example; the metric $$g_{\mu\nu}$$ from the previous posts are another. And those objects are tensors. And we have already met them.15Read more in-depth interpretation of tensors here: https://aronnomirdha.com/expository/990 What vectors reveal is principle underlying the entire ladder: components depend on the language we choose, while the object itself does not.

References and Footnotes

  • 1
    The ancient Greek discovery that the diagonal of a unit square cannot be expressed as a ratio of two integers was one of the earliest demonstrations that the rational numbers are incomplete. ↩︎
  • 2
    The parallelogram rule for composing forces was stated clearly by Simon Stevin in 1586 and appears as a corollary in Newton’s Principia, so the geometric content of vector addition is centuries older than the word. The word itself comes from William Rowan Hamilton in the 1840s, from the Latin vehere, to carry, as part of his quaternions; the modern, stripped-down vector analysis was carved out of the quaternion system by Gibbs and Heaviside toward the end of the nineteenth century. It is a pleasing echo of the professor’s story: Hamilton was searching for a three-dimensional analogue of the complex numbers, which had themselves been invented when the real numbers ran out. ↩︎
  • 3
    The two identities used here are $\cos(\varphi-\theta) = \cos\varphi\cos\theta + \sin\varphi\sin\theta$ and $\sin(\varphi-\theta) = \sin\varphi\cos\theta – \cos\varphi\sin\theta$. They are the whole trigonometric input of this article; everything else is substitution. ↩︎
  • 4
    A bookkeeping note: throughout, the walk is held fixed and the axes rotate, which is called a passive transformation. One can instead hold the axes fixed and rotate the object by $\theta$, an active transformation; the formulas are the same with $\theta$ replaced by $-\theta$. Both conventions are common, which is why signs in other books may differ from mine. ↩︎
  • 5
    I am borrowing a fact here: that a displacement in the plane determines its east part and north part uniquely. Geometrically this is clear enough. In the general setting it is exactly the uniqueness of representation in a basis, which I prove properly once bases are defined below, at which point this step stops being a loan. ↩︎
  • 6
    This opposition has standard names: components that transform inversely to the basis are called contravariant, and objects that transform along with the basis, as the gradient does, are called covariant. The raised and lowered indices of the previous post are precisely this bookkeeping, $x^\mu$ upstairs for contravariant components and $\partial_\mu$ downstairs for covariant objects, and the convention exists so that a paired upper and lower index always assembles a language-independent quantity. ↩︎
  • 7
    For rotations, $R$ is an orthogonal matrix: its rows are the new axis directions written in the old language, it satisfies $\sum_i R_{ij}R_{ik} = \delta_{jk}$, where $\delta_{jk}$ equals $1$ when $j=k$ and $0$ otherwise. This orthogonality is forced rather than assumed: demand that $\sum_i v_i’^2 = \sum_j v_j^2$ for every vector, substitute the transformation law to get $\sum_{j,k}\left(\sum_i R_{ij}R_{ik}\right)v_j v_k$, and the only way the result can equal $\sum_j v_j^2$ for all choices of the $v_j$ is for the bracketed sum to be $\delta_{jk}$. The two-dimensional cancellation we did by hand is the $n=2$ case. Requiring in addition that the frame is not mirror-reflected fixes $\det R = +1$. ↩︎
  • 8
    In modern language: rotations compose to give rotations, every rotation has an inverse, and doing nothing counts as a rotation, so the collection of all rotations forms a group, written $SO(2)$ in the plane and $SO(n)$ in $n$ dimensions. “Transforms like a vector” is, said fully, a statement about how an object responds to this entire group at once, and group theory is the systematic study of such consistency requirements. It becomes one of the load-bearing structures of modern physics. ↩︎
  • 9
    The failure is a finite-angle effect. Rotations through very small angles do commute, up to corrections of second order in the angles, and this is why angular velocity, which is built from infinitesimal rotations, is a legitimate vector while “the rotation itself” is not. There is a further subtlety I am deliberately postponing: quantities like angular velocity behave correctly under rotations but pick up an extra sign under mirror reflection, which is why they are called axial vectors or pseudovectors. That distinction deserves its own discussion. ↩︎
  • 10
    The rules, stated compactly: addition is commutative and associative, there is a zero element, every element has a negative, scaling distributes over both vector addition and number addition, scaling is compatible with multiplication of numbers, and scaling by $1$ does nothing. Eight axioms in all. The “numbers” doing the scaling can be the real numbers or the complex numbers; quantum mechanics is built on complex vector spaces, and everything in this article goes through in either case. ↩︎
  • 11
    Independence here means: if $c_0 + c_1 x + \cdots + c_n x^n$ is the zero vector of this space, which is the polynomial equal to zero for every value of $x$, then all the $c_i$ vanish. Set $x=0$ and $c_0$ falls. Differentiate once, set $x=0$ again, and $c_1$ falls. Keep differentiating; each coefficient drops in turn. So the monomials carry no redundancy, and the components below are the unique ones. ↩︎
  • 12
    A basis is far from unique, but a theorem guarantees that every basis of a given space contains the same number of elements, and that shared number is the dimension of the space. Displacements in a plane have dimension two, displacements in the room have dimension three, and the polynomials and the wave-equation solutions above have infinite dimension, which changes some technical details and none of the conceptual ones. ↩︎
  • 13
    To make “component along a basis function” precise you need an inner product on the function space, and the natural one is an integral, $\langle f, g\rangle = \int f(x)\,g(x)\,dx$ over the interval. The sines and cosines are orthogonal with respect to it, which is what makes the coefficients easy to extract, one integral each. Infinite dimensions bring genuine subtleties about convergence that I am stepping over here, but the geometric picture, a function as a vector and its Fourier coefficients as components, is exactly right. ↩︎
  • 14
    That last formula only defines an angle if the quantity $u\cdot v/(|u||v|)$ actually lies between $-1$ and $1$, and the guarantee is the Cauchy–Schwarz inequality, $(u\cdot v)^2 \le (u\cdot u)(v\cdot v)$. The proof is short enough to give in full. For any number $t$, the vector $u – t\,v$ has a non-negative dot product with itself, so $0 \le (u – t\,v)\cdot(u – t\,v) = u\cdot u – 2t\,(u\cdot v) + t^2\,(v\cdot v)$. The right side is a quadratic in $t$ that never dips below zero, so it cannot have two distinct real roots, so its discriminant obeys $4(u\cdot v)^2 – 4(u\cdot u)(v\cdot v) \le 0$, which is the inequality. Nothing about arrows was used, only the rules of the inner product, so the result holds in every inner product space, the function spaces of the previous footnote included. The definition of angle rests on a theorem, and now you have seen the theorem. ↩︎
  • 15
    Read more in-depth interpretation of tensors here: https://aronnomirdha.com/expository/990 ↩︎
Posted in Expository | Tagged , , , , | 1 Comment

Deriving the Euler–Lagrange Equation for Particles, Fields, and Curved Spacetime

A good way to understand the laws of motion is not to memorize a separate equation for each situation, but to notice there is really only one trick, and every equation of motion is what falls out when you apply it. That sounds like an exaggeration. It isn’t. I’m going to take the trick apart in front of you, use it to derive Newton’s law, then run the very same steps to get the equation for a field spread through space, and once more for a field living in curved spacetime. Three arenas that look nothing alike in a textbook, one move that never changes. Try something as we go: before I write the next line, guess it. If the trick really is this general, you should be able to see the physics coming before the algebra confirms it.

Stated before any symbols get in the way, the trick goes like this. You take every history the world could possibly have, every path a particle might follow and every way a field might wave, and to each one you attach a single number. You call that number the action.1The action is a functional: unlike an ordinary function, which eats a number and returns a number, it eats an entire history $q_r(t)$ and returns a single number $S$. Its units are energy multiplied by time. The name and the idea run back through Maupertuis, Euler, Lagrange, and Hamilton across the eighteenth and nineteenth centuries. Then you make one demand: the history nature actually chooses is the one for which the number does not change, to first order, when you wiggle the history a little. Carry out that demand honestly and it hands you the equation of motion. That’s the whole engine. What follows is just running it in three arenas and watching it do the same job every time.

Start with the simplest arena there is, a single particle moving along a path. To say where the particle is at each moment, give its coordinates $q_r(t)$, where the label $r$ runs over as many numbers as you need to pin it down.2The label $r$ collects what are called generalized coordinates. For one particle in three dimensions it runs over $1,2,3$; for $N$ particles it runs to $3N$; and it need not be a Cartesian position at all, since an angle or any other convenient parameter works just as well. The derivation never cares which you choose. The action is a running total of a quantity $L$, the Lagrangian, tallied all along the path from a starting time to a finishing time:

$$ S = \int_{t_1}^{t_2} L(q_r, \dot q_r, t)\, dt. $$

The Lagrangian is allowed to depend on where the particle is, on how fast it is going, and possibly on the time itself, and on nothing else. Hold onto that restriction a little harder than usual; it’s doing more work than it looks like. Ask yourself what would go wrong if $L$ were also allowed to depend on the acceleration $\ddot q_r$. Keep that question in your pocket. We’ll come back to it right after the equation falls out, and the answer isn’t a footnote you can skip.3If $L$ depended on $\ddot q_r$ as well, the variation would generically produce a fourth-order equation, $\frac{\partial L}{\partial q_r} – \frac{d}{dt}\frac{\partial L}{\partial \dot q_r} + \frac{d^2}{dt^2}\frac{\partial L}{\partial \ddot q_r} = 0$, since integrating the acceleration term by parts twice brings down two extra time derivatives. Generically, but not always: a degenerate dependence on $\ddot q_r$, or a piece of $L$ that is itself a total time derivative, can bring the order back down. Ostrogradski showed in 1850 that a genuinely non-degenerate higher-derivative theory carries a Hamiltonian unbounded below, energy extractable without limit, a real instability rather than a mathematical curiosity, which is why most fundamental Lagrangians avoid the non-degenerate case. General relativity is itself a caveat worth knowing: the Einstein–Hilbert density does contain second derivatives of the metric, but they arrange themselves into a special boundary-term structure that keeps the resulting field equations second order. And when we later move to fields, this same restriction usually reappears as $\mathcal L$ depending only on $\phi_A$ and its first spacetime derivatives. It is the safe, common choice, not a law of nature the way the Euler–Lagrange equation itself is.

Time for the actual move. Take the true path, whatever it happens to be, and nudge it aside by a small amount:

$$ q_r(t) \longmapsto q_r(t) + \varepsilon\, \eta_r(t). $$

The function $\eta_r(t)$ is the shape of the nudge, an arbitrary bend you are free to invent, and $\varepsilon$ is a small knob that controls how hard you push. There is exactly one string attached. The nudge has to vanish at the two ends,

$$ \eta_r(t_1) = 0, \qquad \eta_r(t_2) = 0, $$

which says in plain words that you may bend the path however you like in the middle, but the particle must start where it started and arrive where it arrived. That condition will come back later, much enlarged, and recognizing it when it does is half of the whole story.

q t t₁ t₂ εη(t) q(t) q(t) + εη(t)

Figure 1: The true path and a nearby varied path. Both are pinned at the two endpoints, so the nudge $\varepsilon\eta(t)$ is free to bulge in the middle but must close up at the ends. The action is stationary on the true path.

To say the action is stationary is to say that as you turn the knob $\varepsilon$ up through zero, the action does not move, to first order:

$$ \left.\frac{dS}{d\varepsilon}\right|_{\varepsilon=0} = 0. $$

This is the precise version of the loose phrase about wiggling the path.4Stationary is the honest word, not least. The true history is guaranteed only to be a critical point of the action, where the first-order change vanishes, and depending on the problem it may be a minimum, a maximum, or a saddle. The traditional name “principle of least action” is therefore a mild historical misnomer, a point Feynman was fond of making. What is always true is that the action does not change to first order, which is exactly the condition $dS/d\varepsilon|_{0}=0$. Differentiating is straightforward: pull the derivative inside the integral sign, apply the chain rule, and remember that when the path wiggles, its velocity wiggles too, at the rate $\dot\eta_r$:5Pulling the derivative $d/d\varepsilon$ inside the integral sign is itself a small theorem, not a free move. It requires $L$ and its first derivatives to be continuous in $\varepsilon$ near $\varepsilon=0$ and the path $q_r(t)$ to be differentiable on $[t_1,t_2]$; under those mild smoothness conditions, differentiation under the integral sign (a form of the Leibniz integral rule) is justified. Physics almost never worries about this, and it almost never needs to, but it is worth knowing the fine print exists.

$$ \left.\frac{dS}{d\varepsilon}\right|_{0} = \int_{t_1}^{t_2} \sum_r \left[ \frac{\partial L}{\partial q_r}\, \eta_r + \frac{\partial L}{\partial \dot q_r}\, \dot\eta_r \right] dt. $$

The whole difficulty of this subject lives inside that bracket, and so does the solution. The nudge $\eta_r$ appears bare in the first term but differentiated in the second, as $\dot\eta_r$. I’d like to factor it out and declare that whatever multiplies it must vanish, except I can’t: the nudge is wearing two different outfits. Integration by parts fixes this. It’s really the only idea in the whole business: it peels the derivative off the nudge and lays it onto everything else:

$$ \frac{\partial L}{\partial \dot q_r}\, \dot\eta_r = \frac{d}{dt}\!\left( \frac{\partial L}{\partial \dot q_r}\, \eta_r \right) – \frac{d}{dt}\!\left( \frac{\partial L}{\partial \dot q_r} \right)\eta_r. $$

Substitute this back. The first piece is a total time derivative, so when you integrate it from $t_1$ to $t_2$ it collapses to its values at the two endpoints, and the rest gathers into a single bracket multiplying the bare nudge:

$$ \delta S = \int_{t_1}^{t_2} \sum_r \left[ \frac{\partial L}{\partial q_r} – \frac{d}{dt}\!\left(\frac{\partial L}{\partial \dot q_r}\right) \right]\eta_r\, dt \;+\; \sum_r \left[ \frac{\partial L}{\partial \dot q_r}\, \eta_r \right]_{t_1}^{t_2}. $$

The endpoint term now removes itself. It’s the nudge evaluated at the two ends, and we agreed at the outset that the nudge is zero there. So it’s gone:

$$ \left[ \frac{\partial L}{\partial \dot q_r}\, \eta_r \right]_{t_1}^{t_2} = 0. $$

What is left is an integral of some bracket multiplied by the nudge, and it has to equal zero not for one clever choice of nudge but for every nudge you could ever draw. A fixed function that integrates to zero against absolutely every wiggle has nowhere to hide. It must be zero at every point.6This is the fundamental lemma of the calculus of variations, and it is worth stating carefully. If $f(t)$ is continuous on $[t_1,t_2]$ and $\int_{t_1}^{t_2} f(t)\,\eta(t)\,dt = 0$ for every smooth $\eta$ that vanishes at the two endpoints, then $f(t)=0$ everywhere on the interval. The argument is by contradiction: if $f$ were nonzero, say positive, at some interior point $t_0$, continuity guarantees $f$ stays positive on some small interval around $t_0$; choosing an $\eta$ that bulges upward only inside that interval and is zero elsewhere forces the integral to be strictly positive, contradicting the assumption that it is zero for every such $\eta$. Therefore the bracket vanishes on its own:

$$ \boxed{\; \frac{d}{dt}\!\left( \frac{\partial L}{\partial \dot q_r} \right) – \frac{\partial L}{\partial q_r} = 0. \;} $$

This is the Euler–Lagrange equation, and it came out of exactly three steps. You wiggled the path. You integrated by parts to move the derivative off the wiggle. You threw away the endpoint term because you had pinned the wiggle down. Photograph the shape of that equation and keep it in your wallet: a time derivative of the derivative of $L$ with respect to the velocity, minus the derivative of $L$ with respect to the position. In a moment we’ll meet its identical twin. As promised, only $\ddot q_r$ ever entered the story, through the single time derivative acting on $\partial L/\partial \dot q_r$. The acceleration you were tempted to add to $L$ at the very start never got the chance to appear twice over, which is exactly what keeps this equation second order and well behaved.

Before we go anywhere bigger, let me convince you the equation is not an empty formality by feeding it the most ordinary situation in physics. One particle, three dimensions, and the Lagrangian everyone meets first, kinetic energy minus potential energy:

$$ L = \tfrac{1}{2} m\left(\dot x_1^2 + \dot x_2^2 + \dot x_3^2\right) – V(x_1, x_2, x_3, t). $$

The derivative of $L$ with respect to a velocity is the momentum in that direction, and its time derivative is mass times acceleration:

$$ \frac{\partial L}{\partial \dot x_r} = m \dot x_r, \qquad \frac{d}{dt}\!\left(\frac{\partial L}{\partial \dot x_r}\right) = m \ddot x_r. $$

The derivative of $L$ with respect to a position is minus the slope of the potential in that direction:

$$ \frac{\partial L}{\partial x_r} = -\frac{\partial V}{\partial x_r}. $$

Put both into the box, rearrange, and out falls

$$ \boxed{\; m \ddot x_r = -\frac{\partial V}{\partial x_r} = F_r. \;} $$

That is $F = ma$, with the force revealed as minus the gradient of the potential.7If the force depends on velocity, as the magnetic force does, it cannot be written as minus the gradient of an ordinary potential $V(x)$. The remedy is a velocity-dependent potential $U(x,\dot x,t)$, but the generalized force it produces is $Q_r = \frac{d}{dt}\!\left(\frac{\partial U}{\partial \dot q_r}\right) – \frac{\partial U}{\partial q_r}$, not simply $-\partial U/\partial q_r$; the extra total-time-derivative piece is exactly what lets $U$ generate a velocity-dependent force at all. The full Lorentz force on a charge comes out of exactly this construction, with $L = \frac12 m\mathbf v^2 + q\mathbf A\cdot\mathbf v – q\Phi$. We didn’t assume Newton’s law anywhere; we wrote down a single scalar, kinetic minus potential, demanded its running total be stationary, and Newton’s law came out the far end with no further input. The principle of stationary action isn’t a decorative restatement of Newtonian mechanics sitting politely beside it. It contains Newtonian mechanics, and, as we’re about to see, a great deal that Newton never wrote down.

Here’s the real leap, so take it slowly. For the particle, the thing with a life of its own was a position that depends on time, $q_r(t)$. We now hand that starring role to a field: a number $\phi_A$ that has a value at every point of spacetime at once,

$$ \phi_A = \phi_A(x^\mu), \qquad x^0 = ct,\; x^1 = x,\; x^2 = y,\; x^3 = z. $$

The lower label $A$ is a roster, telling you which field you mean when there are several of them.8The label $A$ simply counts the independent components of whatever field you are studying. A single real scalar has one, a complex scalar has effectively two, the electromagnetic potential $A_\mu$ has four, and the metric of general relativity has ten. What follows never cares what the count is. The upper index $\mu$, running from $0$ to $3$, is a compass, telling you which of the four spacetime directions you are differentiating along. I will write $\partial_\mu$ for the derivative along direction $\mu$, so that the single symbol $\partial_\mu \phi_A$ gathers up how fast the field is changing in time and in each of the three directions of space, all in one stroke.

Only one thing about the action has to change, and for a simple reason. The particle’s action was a total along a one-dimensional path through time. A field is not at a point, it is everywhere, so its action must be a total over a whole slab of spacetime:

$$ S = \int_\Omega \mathcal L\left(\phi_A, \partial_\mu \phi_A, x^\mu\right) d^4x. $$

The script $\mathcal L$ is a Lagrangian density, meaning an amount of action packed into each little box of spacetime volume, and $d^4x = dx^0\,dx^1\,dx^2\,dx^3$ is the size of the box.9There is a factor of $c$ hiding in $x^0 = ct$, which makes $d^4x = c\,dt\,d^3x$. You can absorb that constant into the definition of $\mathcal L$, or you can set $x^0 = t$ and carry the factors of $c$ explicitly. Either way it is bookkeeping: a constant multiplying the whole action never changes which history makes the action stationary, so it cannot affect the equation of motion. Set this next to the particle line by line and the correspondence is almost embarrassing. Where the particle had $L$, the field has a density $\mathcal L$. Where the particle’s Lagrangian leaned on the velocity $\dot q_r$, the field’s density leans on the spacetime gradient $\partial_\mu \phi_A$. Where the particle summed $dt$ along a line, the field sums $d^4x$ over a region. It’s the same skeleton with one dimension traded for four. If you’ve really absorbed the particle derivation, you can already write down the field equation before I derive it. Try it.

Wiggle the field, and pin the wiggle not at two endpoints but on the entire boundary of the region:

$$ \phi \longmapsto \phi + \varepsilon\, \eta, \qquad \eta = 0 \text{ on } \partial\Omega. $$

I’ll carry a single field through the algebra to keep the symbols uncluttered; the case of many fields costs nothing, and I’ll collect that dividend at the end.10With several fields $\phi_A$, you introduce an independent nudge $\eta_A$ for each and run the same variation. Because the nudges are independent, the demand that the total variation vanish for all of them at once splits into one Euler–Lagrange equation per field. This is why the several-field case, which looks more general, is not any harder. The pinning condition is worth a second look. For the particle I froze the two endpoints of the path; for the field I freeze the entire boundary surface of the spacetime slab. But a boundary is a boundary. The two endpoints of the particle were nothing more than the boundary of a one-dimensional interval, and this is the same condition, promoted from the ends of a line to the skin of a four-dimensional region.

Demand stationarity, run the chain rule exactly as we did for the particle, and note that when the field wiggles its gradient wiggles at the rate $\partial_\mu \eta$:

$$ \delta S = \int_\Omega \left[ \frac{\partial \mathcal L}{\partial \phi}\, \eta + \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)}\, \partial_\mu \eta \right] d^4x. $$

There is the same obstacle we already know how to beat. The nudge is bare in the first term and differentiated in the second, so it cannot yet be factored out. Reach for the same tool. Integration by parts, now carried out in spacetime, lifts the derivative off the nudge:

$$ \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)}\, \partial_\mu \eta = \partial_\mu\!\left[ \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)}\, \eta \right] – \partial_\mu\!\left[ \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)} \right] \eta. $$

The first piece is a four-dimensional divergence, and integrating a divergence over a region converts it into a flux through the boundary of that region. That conversion is the divergence theorem:

$$ \int_\Omega \partial_\mu V^\mu\, d^4x = \int_{\partial\Omega} V^\mu\, d\Sigma_\mu. $$

It’s easy to read past this without noticing what’s actually going on. The endpoint term from the particle case and this new surface term aren’t just analogous; they’re the same theorem. When we integrated a total time derivative from $t_1$ to $t_2$ and collapsed it to its values at the two ends, we were using the fundamental theorem of calculus. And the fundamental theorem of calculus is nothing but the divergence theorem in a single dimension, with the two endpoints playing the part of the boundary.11In one dimension the divergence theorem is just the fundamental theorem of calculus: $\int_{t_1}^{t_2}\frac{dF}{dt}\,dt = F(t_2)-F(t_1)$. The boundary of the interval $[t_1,t_2]$ is the two-point set $\{t_1,t_2\}$, and the signs in $F(t_2)-F(t_1)$ are just the outward orientation of that boundary: $+1$ at $t_2$, where the outward direction agrees with increasing $t$, and $-1$ at $t_1$, where it opposes it. Both the one-dimensional and the four-dimensional statements are special cases of the general Stokes theorem, which says that integrating a derivative over a region equals integrating the original object over the region’s boundary. Raise the number of dimensions from one to four, and the boundary of an interval fattens into a boundary surface; the evaluation at two endpoints swells into a flux through that surface. One theorem, read once in one dimension and once in four. The endpoint term and the surface term aren’t cousins. They’re the same creature at two different sizes.

And the surface term dies for the same reason its one-dimensional ancestor did. We pinned the nudge to zero on the boundary, so the flux through the boundary is zero, and the surface term is gone. What survives is an integral of a bracket multiplied by an arbitrary nudge, which forces the bracket to vanish:

$$ \boxed{\; \partial_\mu\!\left( \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)} \right) – \frac{\partial \mathcal L}{\partial \phi} = 0. \;} $$

Lay this beside the particle’s equation and the resemblance is not a resemblance, it is a template with the words swapped. Where the particle had the time derivative $d/dt$, the field has the spacetime divergence $\partial_\mu$. Where the particle differentiated $L$ with respect to the velocity, the field differentiates $\mathcal L$ with respect to the gradient. Where the particle differentiated with respect to position, the field differentiates with respect to the field value. The equation did not change its shape at all. It only learned to count in four directions where before it counted in one.

Let’s feed this field equation its own simplest case, the way we fed Newton to the particle equation, and see what nature’s plainest field does. Take the density

$$ \mathcal L = \tfrac{1}{2}\, \partial_\mu \phi\, \partial^\mu \phi – \tfrac{1}{2}\, \mu^2 \phi^2, $$

where $\partial^\mu = \eta^{\mu\nu}\partial_\nu$ is the gradient with its index raised by the Minkowski metric, and the constant $\mu$ sets a scale we will read off in a moment.12Apologies for the notation: $\mu$ is doing double duty here, as a spacetime index in $\partial_\mu$ and as the mass constant. Both usages are completely standard, and context keeps them apart, but it is worth flagging so you are not caught off guard. The derivative of the density with respect to the field is

$$ \frac{\partial \mathcal L}{\partial \phi} = -\mu^2 \phi, $$

and the derivative with respect to the gradient comes out clean, because each of the two gradient factors in the first term contributes equally:

$$ \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)} = \partial^\mu \phi. $$

Drop both into the boxed field equation and you get

$$ \partial_\mu \partial^\mu \phi + \mu^2 \phi = 0, \qquad\text{or, more compactly,}\qquad \boxed{\; \Box \phi + \mu^2 \phi = 0. \;} $$

The symbol $\Box = \partial_\mu \partial^\mu$ is the wave operator, the four-dimensional relative of the Laplacian, and written out in ordinary time and space it reads

$$ \frac{1}{c^2}\frac{\partial^2 \phi}{\partial t^2} – \nabla^2 \phi + \frac{m^2 c^2}{\hbar^2}\, \phi = 0. $$

This is the Klein–Gordon equation, the simplest relativistic wave equation a single scalar field can obey, and the constant we called $\mu$ has turned out to be $\mu = mc/\hbar$, an inverse length fixed by the mass.13The constant works out to $\mu = mc/\hbar$, so its reciprocal $1/\mu = \hbar/mc$ is the reduced Compton wavelength of a particle of mass $m$, the natural length scale the mass sets. Throughout I use the signature $(+,-,-,-)$, so that $\partial_\mu\partial^\mu = \frac{1}{c^2}\partial_t^2 – \nabla^2$; with the opposite convention a sign flips, but no physics does. Set the mass to zero and the last term disappears, leaving

$$ \Box \phi = 0, $$

which is the plain wave equation. Calling this “the equation that carries light and sound” is only loosely right: light obeys Maxwell’s vector equations and sound is a pressure disturbance in matter, neither one a single scalar field, but individual components of both do satisfy an equation of exactly this mathematical form, which is part of why the wave equation feels so universal. So the field arena, handed its simplest Lagrangian, produces a wave, in parallel with the particle arena, which handed its simplest Lagrangian produced Newton.

What turned my own picture of these two arenas from “similar” into “the same” is a special case, though it takes a little bookkeeping to come out clean. Suppose the field is lazy and refuses to vary through space, depending on time alone, $\phi = \phi(t)$. Every spatial derivative is now zero, and integrating $\mathcal L$ over the now-irrelevant spatial directions just multiplies it by the coordinate volume $V$ of the region, so the effective, particle-like Lagrangian is $L_{\rm eff} = \int d^3x\, \mathcal L = V\mathcal L$; the constant $V$ divides straight out of the resulting equation of motion and never touches it. There is a second bit of bookkeeping too: with $x^0=ct$, the time part of the spacetime derivative is $\partial_0\phi = \frac{1}{c}\dot\phi$, not $\dot\phi$ itself, so the factors of $c$ only disappear cleanly if you set $x^0=t$ from the start or carry them through by hand. Small thing, easy to trip on. Do that, and the boxed field equation collapses into

$$ \frac{d}{dt}\!\left( \frac{\partial L_{\rm eff}}{\partial \dot\phi} \right) – \frac{\partial L_{\rm eff}}{\partial \phi} = 0, $$

which is the particle’s Euler–Lagrange equation in every structural respect, with $L_{\rm eff}$ standing in for $L$. So the field equation is not an analogy of the particle equation. It is a generalization that contains the particle equation exactly, as the special case in which nothing depends on where you are, once the volume factor and the coordinate convention have been tracked honestly rather than waved away. Peel away the three dimensions of space and the field remembers that it was the same law all along.

The last arena is curved spacetime: a metric $g_{\mu\nu}(x)$ now sets how distances and times are measured, and it changes from place to place. Two things change in the passage from flat to curved, one about how you measure volume and one about how you differentiate. Neither, it turns out, touches the move itself.

The first change is the volume of the little box. On flat spacetime the box had volume $d^4x$; on a curved manifold the correct, coordinate-independent volume is $\sqrt{-g}\, d^4x$, where $g$ is the determinant of the metric.14For a Lorentzian metric the determinant $g$ is negative, so $-g$ is positive and $\sqrt{-g}$ is real. The combination $\sqrt{-g}\,d^4x$ is the invariant four-volume, the same number in every coordinate system. In flat Minkowski coordinates the metric determinant is $-1$, so $\sqrt{-g}=1$ and the invariant volume reduces to the ordinary $d^4x$, which is why the flat-space derivation never had to mention it. The second change is that the plain derivative $\partial_\mu$ ought to be promoted to the covariant derivative $\nabla_\mu$, the derivative that accounts for the bending of the coordinates. For a scalar field these two happen to coincide, $\nabla_\mu \phi = \partial_\mu \phi$, which is a small mercy I am going to accept gratefully.15Two facts are doing quiet work here. First, a scalar field has no indices for the connection to act on, so $\nabla_\mu\phi = \partial_\mu\phi$; the first place the two derivatives part ways is a vector, where $\nabla_\mu V^\nu = \partial_\mu V^\nu + \Gamma^\nu{}_{\mu\lambda}V^\lambda$. Second, the divergence identity $\nabla_\mu V^\mu = \frac{1}{\sqrt{-g}}\partial_\mu(\sqrt{-g}\,V^\mu)$ used two lines below follows from the contracted Christoffel symbol $\Gamma^\lambda{}_{\mu\lambda} = \partial_\mu \ln\sqrt{-g}$. To keep the notation consistent from here on, the density should be understood as depending on the covariant derivative rather than the plain one,

$$ \boxed{\; \mathcal L = \mathcal L(\phi, \nabla_\mu \phi, x^\mu). \;} $$

For this scalar the two derivatives happen to coincide, which is why writing $\partial_\mu\phi$ in its place cost us nothing above, but it is $\nabla_\mu\phi$ that is the geometrically correct object, and it is the one that generalizes once the field stops being a scalar. The action reads

$$ S = \int_\Omega \mathcal L\, \sqrt{-g}\; d^4x. $$

Run the move a third time: wiggle the field, pin it on the boundary, demand stationarity, integrate by parts, this time with the covariant product rule. The one step with any right to go wrong is the boundary term. The curvature has to show up somewhere, after all, and if not there, it must be hiding elsewhere. It is. It sits in plain sight, inside a small identity:

$$ \nabla_\mu V^\mu = \frac{1}{\sqrt{-g}}\, \partial_\mu\!\left( \sqrt{-g}\, V^\mu \right). $$

The identity doesn’t make two copies of $\sqrt{-g}$ literally cancel; what it does is convert the covariant divergence $\nabla_\mu V^\mu$, sitting inside the honest curved-space measure $\sqrt{-g}\,d^4x$, into an ordinary coordinate divergence $\partial_\mu(\sqrt{-g}\,V^\mu)$ against the flat measure $d^4x$. That ordinary coordinate divergence is exactly the object the flat-space divergence theorem already knows how to turn into a boundary flux. So the boundary piece becomes an ordinary flux, which vanishes because the nudge vanishes on the boundary. Same death, same reason, third time in a row; the curvature never got a chance to spoil the argument because it was absorbed into the very identity that lets the divergence theorem apply at all. What survives is

$$ \boxed{\; \nabla_\mu\!\left( \frac{\partial \mathcal L}{\partial(\nabla_\mu \phi)} \right) – \frac{\partial \mathcal L}{\partial \phi} = 0, \;} $$

and, if you prefer to see the metric factor out in the open,

$$ \frac{1}{\sqrt{-g}}\, \partial_\mu\!\left[ \sqrt{-g}\, \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)} \right] – \frac{\partial \mathcal L}{\partial \phi} = 0. $$

Feed in the same free scalar density used in flat spacetime, now built from the curved metric instead of the flat one,

$$ \mathcal L = \tfrac{1}{2}\, g^{\mu\nu}\partial_\mu\phi\,\partial_\nu\phi – \tfrac{1}{2}\, \mu^2 \phi^2, $$

and the two derivatives you need are exactly the flat-space ones with $\eta^{\mu\nu}$ replaced by $g^{\mu\nu}$:

$$ \frac{\partial \mathcal L}{\partial \phi} = -\mu^2 \phi, \qquad \frac{\partial \mathcal L}{\partial(\partial_\mu \phi)} = g^{\mu\nu}\partial_\nu\phi. $$

Drop both into the boxed curved-space equation and out comes the curved-space Klein–Gordon equation:

$$ \boxed{\; \frac{1}{\sqrt{-g}}\, \partial_\mu\!\left( \sqrt{-g}\, g^{\mu\nu}\partial_\nu\phi \right) + \mu^2 \phi = 0. \;} $$

This is the third arena’s own famous equation, and it completes the parallel exactly: the particle gave Newton, the flat field gave Klein–Gordon, and the curved field gives this, the same Klein–Gordon equation with the flat metric $\eta^{\mu\nu}$ replaced by the curved $g^{\mu\nu}$ and a $\sqrt{-g}$ riding inside the derivative to keep the divergence covariant. Set $g_{\mu\nu}\to\eta_{\mu\nu}$, so that $\sqrt{-g}\to1$, and it collapses straight back to the flat equation we already derived, exactly as it must.

There is a subtlety in that last line. It looks as though the equation has forgotten that the metric depends on position, but it hasn’t forgotten at all. The factor $\sqrt{-g}$ depends only on where you are, never on the field or on the field’s gradient, so when you differentiate the density with respect to the field, the $\sqrt{-g}$ is not a variable to be differentiated but a bystander riding along inside the explicit coordinate dependence. It’s luggage the field carries, not a control the field can turn. That is precisely why it can sit inside the derivative in one line and outside it in the next without contradiction.

I leaned, twice now, on the field being a scalar, and I owe you the reason it mattered. When the field carries indices of its own, when it is a vector or a tensor rather than a plain number, the covariant derivative stops being equal to the plain derivative, the connection terms that measure the bending of the coordinates come out of hiding, and the bookkeeping grows real teeth.16For a vector or tensor field you must differentiate the density with respect to $\nabla_\mu\phi_A$ with its indices intact, and the resulting Euler–Lagrange operator carries extra connection terms. The Maxwell and Proca equations, and the linearized equations for a metric perturbation, all come out this way. The scalar spares us those terms without hiding the essential move. The scalar is the clean room in which you can watch the move work with nothing extra clinging to it. It is the right place to learn the pattern before the extra machinery arrives.

So what did we actually do? We wrote an action for a particle, demanded it be stationary, and Newton’s law came out. We wrote an action for a field, demanded it be stationary, and a wave came out. We wrote an action for a field in curved spacetime, demanded it be stationary, and the curved-space wave equation came out. Three arenas that share no furniture, and in every one of them the derivation was the same three motions: wiggle the history, integrate by parts to move the derivative off the wiggle, and discard the boundary term because the wiggle was pinned down there.

The many-field dividend I promised earlier is now free. If there are many fields $\phi_A$ instead of one, you introduce an independent nudge for each and run the same variation, and because the nudges are independent, each field simply hands you its own copy of the boxed equation. Nothing in the three steps ever asked how many fields were present. The move does not become harder when the world becomes more complicated. It only becomes longer to write down.

There’s a second dividend here, one I’ve been quietly setting up all along and haven’t cashed in yet. We discarded the boundary term every single time because we forced the nudge to vanish there. It is tempting to think that if you refuse to force it, say by choosing $\eta_r=\dot q_r$, the shift generated by letting time flow forward, the endpoint term $\left[\partial L/\partial\dot q_r\,\dot q_r\right]_{t_1}^{t_2}$ is itself the conserved energy. It isn’t. That guess is missing a piece, and the route to what actually is conserved goes through the Lagrangian’s own total time derivative along the true trajectory:

$$ \frac{dL}{dt} = \sum_r \frac{\partial L}{\partial q_r}\dot q_r + \sum_r \frac{\partial L}{\partial \dot q_r}\ddot q_r + \frac{\partial L}{\partial t}. $$

Now use the Euler–Lagrange equation itself, $\partial L/\partial q_r = \frac{d}{dt}(\partial L/\partial \dot q_r)$, to trade the first term for a total time derivative. The right-hand side collapses into a single derivative:

$$ \frac{d}{dt}\left( \sum_r \dot q_r\, \frac{\partial L}{\partial \dot q_r} – L \right) = -\frac{\partial L}{\partial t}. $$

So whenever $L$ carries no explicit time dependence, the quantity in parentheses is exactly conserved:

$$ \boxed{\; E = \sum_r \dot q_r\, \frac{\partial L}{\partial \dot q_r} – L. \;} $$

This is the energy, or in Hamiltonian language, $H$. It needed both the momentum-like piece $\dot q_r\,\partial L/\partial\dot q_r$ and the compensating $-L$; the boundary term alone, without subtracting $L$, is not a conserved quantity at all, which is exactly the trap the naive guess above falls into.17This is the shadow of Noether’s theorem falling across the same calculation we already did, though done correctly it requires varying time itself alongside the coordinates rather than treating $\eta_r=\dot q_r$ as an ordinary fixed-time nudge. In the field case the analogous argument, built from a spacetime translation, produces the canonical energy–momentum tensor $T^\mu{}_\nu = \frac{\partial \mathcal L}{\partial(\partial_\mu\phi_A)}\,\partial_\nu\phi_A – \delta^\mu{}_\nu\,\mathcal L$, with $\partial_\mu T^\mu{}_\nu=0$ the local conservation law for the field’s energy and momentum; it reduces to $\partial^\mu\phi\,\partial_\nu\phi – \delta^\mu{}_\nu\mathcal L$ only for the specific free-scalar density used earlier, since that is the one case where $\partial\mathcal L/\partial(\partial_\mu\phi)=\partial^\mu\phi$. Momentum and angular momentum come from spacetime symmetries in exactly this family. Electric charge conservation is a genuine Noether current too, but it comes from an internal symmetry, a global phase rotation of the field, not from a spacetime translation, so it is not itself a component of $T^\mu{}_\nu$. The lesson survives the correction: a piece of the calculation you learned to throw away, tracked instead of discarded, is where every conservation law in physics comes from. That’s not a small aside. It may be the more important half of this entire trick, and it’s the subject for the next post.

None of this is a coincidence of notation. Every experimentally successful fundamental classical field theory currently known admits an action formulation. That’s why a single trick can reach from a bead on a wire all the way to a field in curved spacetime. The principle of stationary action is one sentence, that nature makes a certain total stationary, and the equation of motion is simply what that sentence looks like once you’ve carried out the wiggle without cheating. Newton’s law, the Klein–Gordon equation, and a scalar field threading its way through curved spacetime aren’t three laws that happen to rhyme. They’re one demand, put to three arenas, and answered by one move in each.

And if you ever get bold and let the metric itself be the thing you wiggle, instead of a fixed stage you have set the field upon, the same move hands you Einstein’s field equations for gravity.18The gravitational case takes the Einstein–Hilbert action $S = \frac{1}{2\kappa}\int R\,\sqrt{-g}\,d^4x$, where $R$ is the Ricci scalar built from the metric, and varies with respect to the metric $g^{\mu\nu}$ itself rather than a field living on top of it. Here “more careful boundary analysis” is not a figure of speech: the variation of $R$ produces terms containing derivatives of $\delta g_{\mu\nu}$, not just $\delta g_{\mu\nu}$ itself, so pinning $\delta g_{\mu\nu}=0$ on the boundary is not by itself enough to kill the boundary term. A well-posed variational principle needs the Gibbons–Hawking–York term added to the action to cancel the offending piece. With that included, the same wiggle-and-discard procedure yields Einstein’s field equations $G_{\mu\nu} = \kappa T_{\mu\nu}$. But that’s a story for another day, and it’s the same trick once more. When you’ve truly seen this move, you haven’t learned a drawer full of separate equations. You’ve learned one thing, and then watched it grow up.

References and Footnotes

  • 1
    The action is a functional: unlike an ordinary function, which eats a number and returns a number, it eats an entire history $q_r(t)$ and returns a single number $S$. Its units are energy multiplied by time. The name and the idea run back through Maupertuis, Euler, Lagrange, and Hamilton across the eighteenth and nineteenth centuries. ↩︎
  • 2
    The label $r$ collects what are called generalized coordinates. For one particle in three dimensions it runs over $1,2,3$; for $N$ particles it runs to $3N$; and it need not be a Cartesian position at all, since an angle or any other convenient parameter works just as well. The derivation never cares which you choose. ↩︎
  • 3
    If $L$ depended on $\ddot q_r$ as well, the variation would generically produce a fourth-order equation, $\frac{\partial L}{\partial q_r} – \frac{d}{dt}\frac{\partial L}{\partial \dot q_r} + \frac{d^2}{dt^2}\frac{\partial L}{\partial \ddot q_r} = 0$, since integrating the acceleration term by parts twice brings down two extra time derivatives. Generically, but not always: a degenerate dependence on $\ddot q_r$, or a piece of $L$ that is itself a total time derivative, can bring the order back down. Ostrogradski showed in 1850 that a genuinely non-degenerate higher-derivative theory carries a Hamiltonian unbounded below, energy extractable without limit, a real instability rather than a mathematical curiosity, which is why most fundamental Lagrangians avoid the non-degenerate case. General relativity is itself a caveat worth knowing: the Einstein–Hilbert density does contain second derivatives of the metric, but they arrange themselves into a special boundary-term structure that keeps the resulting field equations second order. ↩︎
  • 4
    Stationary is the honest word, not least. The true history is guaranteed only to be a critical point of the action, where the first-order change vanishes, and depending on the problem it may be a minimum, a maximum, or a saddle. The traditional name “principle of least action” is therefore a mild historical misnomer, a point Feynman was fond of making. What is always true is that the action does not change to first order, which is exactly the condition $dS/d\varepsilon|_{0}=0$. ↩︎
  • 5
    Pulling the derivative $d/d\varepsilon$ inside the integral sign is itself a small theorem, not a free move. It requires $L$ and its first derivatives to be continuous in $\varepsilon$ near $\varepsilon=0$ and the path $q_r(t)$ to be differentiable on $[t_1,t_2]$; under those mild smoothness conditions, differentiation under the integral sign (a form of the Leibniz integral rule) is justified. Physics almost never worries about this, and it almost never needs to, but it is worth knowing the fine print exists. ↩︎
  • 6
    This is the fundamental lemma of the calculus of variations, and it is worth stating carefully. If $f(t)$ is continuous on $[t_1,t_2]$ and $\int_{t_1}^{t_2} f(t)\,\eta(t)\,dt = 0$ for every smooth $\eta$ that vanishes at the two endpoints, then $f(t)=0$ everywhere on the interval. The argument is by contradiction: if $f$ were nonzero, say positive, at some interior point $t_0$, continuity guarantees $f$ stays positive on some small interval around $t_0$; choosing an $\eta$ that bulges upward only inside that interval and is zero elsewhere forces the integral to be strictly positive, contradicting the assumption that it is zero for every such $\eta$. ↩︎
  • 7
    If the force depends on velocity, as the magnetic force does, it cannot be written as minus the gradient of an ordinary potential $V(x)$. The remedy is a velocity-dependent potential $U(x,\dot x,t)$, but the generalized force it produces is $Q_r = \frac{d}{dt}\!\left(\frac{\partial U}{\partial \dot q_r}\right) – \frac{\partial U}{\partial q_r}$, not simply $-\partial U/\partial q_r$; the extra total-time-derivative piece is exactly what lets $U$ generate a velocity-dependent force at all. The full Lorentz force on a charge comes out of exactly this construction, with $L = \frac12 m\mathbf v^2 + q\mathbf A\cdot\mathbf v – q\Phi$. ↩︎
  • 8
    The label $A$ simply counts the independent components of whatever field you are studying. A single real scalar has one, a complex scalar has effectively two, the electromagnetic potential $A_\mu$ has four, and the metric of general relativity has ten. What follows never cares what the count is. ↩︎
  • 9
    There is a factor of $c$ hiding in $x^0 = ct$, which makes $d^4x = c\,dt\,d^3x$. You can absorb that constant into the definition of $\mathcal L$, or you can set $x^0 = t$ and carry the factors of $c$ explicitly. Either way it is bookkeeping: a constant multiplying the whole action never changes which history makes the action stationary, so it cannot affect the equation of motion. ↩︎
  • 10
    With several fields $\phi_A$, you introduce an independent nudge $\eta_A$ for each and run the same variation. Because the nudges are independent, the demand that the total variation vanish for all of them at once splits into one Euler–Lagrange equation per field. This is why the several-field case, which looks more general, is not any harder. ↩︎
  • 11
    In one dimension the divergence theorem is just the fundamental theorem of calculus: $\int_{t_1}^{t_2}\frac{dF}{dt}\,dt = F(t_2)-F(t_1)$. The boundary of the interval $[t_1,t_2]$ is the two-point set $\{t_1,t_2\}$, and the signs in $F(t_2)-F(t_1)$ are just the outward orientation of that boundary: $+1$ at $t_2$, where the outward direction agrees with increasing $t$, and $-1$ at $t_1$, where it opposes it. Both the one-dimensional and the four-dimensional statements are special cases of the general Stokes theorem, which says that integrating a derivative over a region equals integrating the original object over the region’s boundary. ↩︎
  • 12
    Apologies for the notation: $\mu$ is doing double duty here, as a spacetime index in $\partial_\mu$ and as the mass constant. Both usages are completely standard, and context keeps them apart, but it is worth flagging so you are not caught off guard. ↩︎
  • 13
    The constant works out to $\mu = mc/\hbar$, so its reciprocal $1/\mu = \hbar/mc$ is the reduced Compton wavelength of a particle of mass $m$, the natural length scale the mass sets. Throughout I use the signature $(+,-,-,-)$, so that $\partial_\mu\partial^\mu = \frac{1}{c^2}\partial_t^2 – \nabla^2$; with the opposite convention a sign flips, but no physics does. ↩︎
  • 14
    For a Lorentzian metric the determinant $g$ is negative, so $-g$ is positive and $\sqrt{-g}$ is real. The combination $\sqrt{-g}\,d^4x$ is the invariant four-volume, the same number in every coordinate system. In flat Minkowski coordinates the metric determinant is $-1$, so $\sqrt{-g}=1$ and the invariant volume reduces to the ordinary $d^4x$, which is why the flat-space derivation never had to mention it. ↩︎
  • 15
    Two facts are doing quiet work here. First, a scalar field has no indices for the connection to act on, so $\nabla_\mu\phi = \partial_\mu\phi$; the first place the two derivatives part ways is a vector, where $\nabla_\mu V^\nu = \partial_\mu V^\nu + \Gamma^\nu{}_{\mu\lambda}V^\lambda$. Second, the divergence identity $\nabla_\mu V^\mu = \frac{1}{\sqrt{-g}}\partial_\mu(\sqrt{-g}\,V^\mu)$ used two lines below follows from the contracted Christoffel symbol $\Gamma^\lambda{}_{\mu\lambda} = \partial_\mu \ln\sqrt{-g}$. ↩︎
  • 16
    For a vector or tensor field you must differentiate the density with respect to $\nabla_\mu\phi_A$ with its indices intact, and the resulting Euler–Lagrange operator carries extra connection terms. The Maxwell and Proca equations, and the linearized equations for a metric perturbation, all come out this way. The scalar spares us those terms without hiding the essential move. ↩︎
  • 17
    This is the shadow of Noether’s theorem falling across the same calculation we already did, though done correctly it requires varying time itself alongside the coordinates rather than treating $\eta_r=\dot q_r$ as an ordinary fixed-time nudge. In the field case the analogous argument, built from a spacetime translation, produces the canonical energy–momentum tensor $T^\mu{}_\nu = \frac{\partial \mathcal L}{\partial(\partial_\mu\phi_A)}\,\partial_\nu\phi_A – \delta^\mu{}_\nu\,\mathcal L$, with $\partial_\mu T^\mu{}_\nu=0$ the local conservation law for the field’s energy and momentum; it reduces to $\partial^\mu\phi\,\partial_\nu\phi – \delta^\mu{}_\nu\mathcal L$ only for the specific free-scalar density used earlier, since that is the one case where $\partial\mathcal L/\partial(\partial_\mu\phi)=\partial^\mu\phi$. Momentum and angular momentum come from spacetime symmetries in exactly this family. Electric charge conservation is a genuine Noether current too, but it comes from an internal symmetry, a global phase rotation of the field, not from a spacetime translation, so it is not itself a component of $T^\mu{}_\nu$. ↩︎
  • 18
    The gravitational case takes the Einstein–Hilbert action $S = \frac{1}{2\kappa}\int R\,\sqrt{-g}\,d^4x$, where $R$ is the Ricci scalar built from the metric, and varies with respect to the metric $g^{\mu\nu}$ itself rather than a field living on top of it. Here “more careful boundary analysis” is not a figure of speech: the variation of $R$ produces terms containing derivatives of $\delta g_{\mu\nu}$, not just $\delta g_{\mu\nu}$ itself, so pinning $\delta g_{\mu\nu}=0$ on the boundary is not by itself enough to kill the boundary term. A well-posed variational principle needs the Gibbons–Hawking–York term added to the action to cancel the offending piece. With that included, the same wiggle-and-discard procedure yields Einstein’s field equations $G_{\mu\nu} = \kappa T_{\mu\nu}$. ↩︎
Posted in Expository | Tagged , , , , , , , , , , , | Leave a comment

So what are tensors, really?

I want to start with a confession. When I first heard the word “tensor,” I assumed it was one of those words that exists to make physicists sound clever. Something you learn in graduate school, surrounded by people who already know what it means, and everyone pretends they understood it the first time. It is not like that at all. A tensor is a completely natural object. It is what you are forced to invent when you ask a simple question about a crystal and discover that vectors are not enough to answer it.

Here is the question. You take a crystal, say a piece of calcite, the kind that makes double images when you look through it. You apply an electric field $\vec{E}$ to it. The field pushes the charges inside the crystal a little, and the crystal develops a dipole moment per unit volume, which we call the polarization $\vec{P}$. For a material like water or glass, which looks the same in every direction, the polarization is simply proportional to the field and points in the same direction: $\vec{P} = \alpha \vec{E}$, where $\alpha$ is a single number. One number, and you are done.

A crystal is different. Its atoms sit in a rigid lattice that has a definite structure, and that structure is not the same in every direction. Think of it mechanically. Imagine the atoms connected by springs, but the springs in one direction are stiff and the springs in a perpendicular direction are loose. If you push with a force at some angle, the charges move mostly along the loose direction and barely at all along the stiff direction. The displacement, which is the polarization, comes out pointing somewhere completely different from the force. The polarization is not parallel to the electric field. One number $\alpha$ cannot describe this, because one number cannot encode a relationship that depends on direction in this asymmetric way.

Isotropic vs anisotropic polarization
Left: an isotropic material. The polarization $\vec{P}$ always comes out parallel to the field $\vec{E}$, so one number $\alpha$ suffices. Right: an anisotropic crystal. The polarization is rotated away from the field by an angle $\Delta\theta$ that depends on the crystal structure. One number cannot describe this.

So what do you need? Let us work it out from scratch, without assuming anything. Suppose you apply a field purely in the $x$-direction, magnitude $E_x$. Where does the polarization go? There is no reason it has to go in the $x$-direction. It will have some $x$-component, some $y$-component, and some $z$-component. Since polarization is proportional to field strength (we are assuming small fields), each component of $\vec{P}$ is proportional to $E_x$, and we just have to give the proportionality constants names. Call them $\alpha_{xx}$, $\alpha_{yx}$, and $\alpha_{zx}$, where the first subscript tells you which component of $\vec{P}$ you are measuring and the second tells you which direction the field is in: $$P_x = \alpha_{xx}\, E_x, \qquad P_y = \alpha_{yx}\, E_x, \qquad P_z = \alpha_{zx}\, E_x.$$ Three numbers to describe what happens when you push in the $x$-direction. Now do the same for a field in the $y$-direction: $P_x = \alpha_{xy} E_y$, $P_y = \alpha_{yy} E_y$, $P_z = \alpha_{zy} E_y$. Three more numbers. And three more for a field in the $z$-direction. Nine coefficients in total.1The subscript ordering $\alpha_{ij}$ means: $i$ labels the component of $\vec{P}$ being produced, $j$ labels the direction of $\vec{E}$ that produces it. So $\alpha_{yx}$ is the coefficient that tells you how much $y$-polarization is produced by an $x$-field. Keep this straight from the beginning or the indices will confuse you forever.

Now suppose the field has all three components at once. Because polarization is linear in the field (doubling the field doubles the polarization), the contributions from each direction simply add. The $x$-component of $\vec{P}$ gets a contribution from the $x$-field (coefficient $\alpha_{xx}$), a contribution from the $y$-field (coefficient $\alpha_{xy}$), and a contribution from the $z$-field (coefficient $\alpha_{xz}$). Adding them up: $$\begin{aligned} P_x &= \alpha_{xx} E_x + \alpha_{xy} E_y + \alpha_{xz} E_z, \\ P_y &= \alpha_{yx} E_x + \alpha_{yy} E_y + \alpha_{yz} E_z, \\ P_z &= \alpha_{zx} E_x + \alpha_{zy} E_y + \alpha_{zz} E_z. \end{aligned}$$ These three equations can be written compactly as $$P_i = \sum_j \alpha_{ij} E_j$$ where $i$ and $j$ each run over $x$, $y$, $z$. The nine numbers $\alpha_{ij}$ are called the polarizability tensor. That is where tensors come from. Not from abstraction. From the forced recognition that describing direction-dependent proportionality requires not one number but nine.

The nine tensor coefficients as a table
Each cell in the table is one of the nine coefficients $\alpha_{ij}$. Read the table as: the coefficient in row $i$, column $j$ tells you how much of the $i$-th component of polarization is produced by the $j$-th component of the electric field. The shaded diagonal cells are where the field and polarization point in the same direction.

Look at that table for a moment before moving on. If you apply an $x$-field, you read down the first column: you get some $x$-polarization ($\alpha_{xx}$), some $y$-polarization ($\alpha_{yx}$), and some $z$-polarization ($\alpha_{zx}$). If you apply a $y$-field, you read down the second column. The whole table encodes the complete directional response of the crystal. This is what a tensor is: a table of coefficients that tells you, for each possible direction of the input, what the output is in each possible direction. Not a mysterious geometric object. A table with a physical meaning for each entry.

Now something wonderful happens. Of the nine numbers, not all are independent. Here is the argument. Think about the energy you spend polarizing the crystal. The work done per unit volume when you polarize it is $u_P = \frac{1}{2}\vec{E}\cdot\vec{P}$. Substitute $P_i = \sum_j \alpha_{ij} E_j$ and you get $$u_P = \frac{1}{2}\sum_{i,j} \alpha_{ij} E_i E_j.$$ Now imagine a thermodynamic cycle. Turn on an $x$-field, then turn on a $y$-field, then turn off the $x$-field, then turn off the $y$-field. The crystal is back where it started. The net work done must be zero. When you compute the work around this cycle carefully, the terms proportional to $\alpha_{xx}$ and $\alpha_{yy}$ cancel out cleanly. The net work is proportional to $\alpha_{xy} – \alpha_{yx}$. For this to vanish, you need $\alpha_{xy} = \alpha_{yx}$. The same argument applies to any pair of indices, so the tensor is symmetric: $\alpha_{ij} = \alpha_{ji}$.2This symmetry is a physical consequence of energy conservation, not a mathematical assumption. There exist tensors in physics that are not symmetric, for example the angular momentum flux tensor, and they genuinely need all nine components. Always ask which physical law enforces symmetry before counting independent components. Symmetry cuts nine components to six independent ones. That is not a small reduction. In higher-rank tensors, symmetry can reduce hundreds of components to just a handful, which is why physicists care about it so much.

Six is still more than one. But here is where the geometry becomes beautiful. The energy formula $u_P = \frac{1}{2}\sum_{i,j} \alpha_{ij} E_i E_j$ is a quadratic function of the field components. In two dimensions it would look like $\alpha_{xx} E_x^2 + 2\alpha_{xy} E_x E_y + \alpha_{yy} E_y^2 = 2u_0$. Ask: which electric field vectors $\vec{E}$ produce exactly the same energy density $u_0$? The set of all such vectors traces out a curve in the $(E_x, E_y)$ plane. Since the energy is always positive and finite for any nonzero field, that curve is always an ellipse, never a parabola or hyperbola. In three dimensions, the set of all $\vec{E}$ producing the same $u_0$ traces out an ellipsoid. The shape and orientation of this ellipsoid encodes everything about the tensor.

The energy ellipse of the polarization tensor
Left: when the off-diagonal components are nonzero, the energy ellipse is tilted with respect to the coordinate axes. Right: when you choose the principal axes as your coordinate system, the ellipse aligns with the axes and the off-diagonal terms vanish. The semi-axis lengths are $1/\sqrt{\alpha_{aa}}$ and $1/\sqrt{\alpha_{bb}}$, encoding the two principal values directly.

Every ellipsoid has three perpendicular axes called the principal axes, the directions along which it is longest, shortest, and intermediate. When you align your coordinate system with those axes, something remarkable happens to the tensor: all the off-diagonal components vanish. The tensor becomes diagonal: $$\alpha_{ij} = \begin{pmatrix} \alpha_{aa} & 0 & 0 \\ 0 & \alpha_{bb} & 0 \\ 0 & 0 & \alpha_{cc} \end{pmatrix}.$$ Along the principal axes, the polarization is parallel to the field, with one proportionality constant per axis. The crystal may be complicated in any arbitrary direction, but it has three special directions where it behaves like an isotropic material, each with its own effective $\alpha$. This is always possible for any symmetric tensor of rank two, in any number of dimensions. It is one of the deepest structural facts about this type of object, and it has nothing to do with the specific physics of crystals. It is purely geometric.

If all three principal values are equal, $\alpha_{aa} = \alpha_{bb} = \alpha_{cc} = \alpha$, the ellipsoid is a sphere and the material is isotropic. The tensor reduces to $\alpha_{ij} = \alpha \delta_{ij}$, where $\delta_{ij}$ is one if $i = j$ and zero otherwise. This is called the Kronecker delta, and it plays the role of the identity: $P_i = \alpha \sum_j \delta_{ij} E_j = \alpha E_i$, which is just $\vec{P} = \alpha \vec{E}$. The tensor framework contains the simple isotropic case as a special instance.3The Kronecker delta $\delta_{ij}$ has the remarkable property that it looks identical in every coordinate system. You can rotate your axes by any angle and its components stay the same: ones on the diagonal, zeros everywhere else. This is because it is proportional to the metric of flat Euclidean space, which has no preferred direction. The identity tensor is the one tensor that genuinely has no directional content.

The polarizability of a crystal is one example, but the same structure appears everywhere in physics. The conductivity of a crystal is a tensor: $j_i = \sum_j \sigma_{ij} E_j$ because the current density $\vec{j}$ and the electric field $\vec{E}$ are generally not parallel in a crystal. The moment of inertia is a tensor. You may have seen the scalar moment of inertia $I$ in a basic mechanics course, with $L = I\omega$. But this only works when you spin an object about one of its symmetry axes. For a general axis, the angular momentum $\vec{L}$ and the angular velocity $\vec{\omega}$ are not parallel, in exactly the same way that $\vec{P}$ and $\vec{E}$ are not parallel in an anisotropic crystal.4The inertia tensor has the explicit form $I_{ij} = \sum_\text{particles} m(r^2 \delta_{ij} – r_i r_j)$, where the sum is over all particles in the body and $r^2 = x^2+y^2+z^2$. The diagonal component $I_{xx} = \sum m(y^2+z^2)$ is exactly the moment of inertia about the $x$-axis that you saw in introductory mechanics. The off-diagonal components $I_{xy} = -\sum m\, xy$ are the products of inertia, and they vanish when your axes align with the principal axes of the body.

The inertia tensor: omega and L not parallel
Left: spinning a rod about its symmetry axis. The angular momentum $\vec{L}$ is parallel to $\vec{\omega}$, and a single number $I$ suffices. Right: spinning the same rod about an oblique axis. $\vec{L}$ is tilted away from $\vec{\omega}$ by an angle $\Delta\theta$. Describing this relationship for any axis requires the full inertia tensor $I_{ij}$.

Now for the question that most introductions skip. The nine components of $\alpha_{ij}$ were measured with respect to some particular set of coordinate axes. If you rotate your axes, the components change. Does the tensor change? No. The crystal is the same crystal. The physical relationship between $\vec{P}$ and $\vec{E}$ is unchanged. What changes is only the numerical values of the components, in a specific and predictable way determined by how you rotated the axes. The components are not the tensor. They are the tensor’s representation in a particular coordinate system, just as the components $(E_x, E_y, E_z)$ are not the electric field but its representation in a particular coordinate system.

Components change under rotation, tensor does not
The same physical vector $\vec{E}$ drawn in two coordinate systems rotated by 35 degrees. The arrow is identical. The numerical components $E_x = 3.2, E_y = 2.4$ in the original system become $E_{x’} \approx 3.9, E_{y’} \approx 0.1$ in the rotated system. The same thing happens to every index of a tensor when you rotate your axes.

The transformation law for a rank-2 tensor under a rotation is $$\alpha’_{\mu\nu} = \sum_{\rho,\sigma} R_{\mu\rho}\, R_{\nu\sigma}\, \alpha_{\rho\sigma},$$ where $R_{\mu\rho}$ is the rotation matrix. Each index picks up exactly one factor of the rotation matrix. A rank-1 tensor (a vector) picks up one factor: $v’_\mu = \sum_\rho R_{\mu\rho} v_\rho$. A rank-0 tensor (a scalar) picks up no factors: it is unchanged. This is the general pattern. The rank of a tensor tells you how many rotation-matrix factors appear in its transformation law, and therefore how many indices it carries.

This transformation law is not just bookkeeping. It is the guarantee that tensor equations are coordinate-independent. If $P_i = \sum_j \alpha_{ij} E_j$ holds in one coordinate system, it holds in every coordinate system, because both sides transform in exactly the same way. You write the physics once and it is automatically valid for any observer, any orientation of the laboratory, any choice of axes. This is why tensors are indispensable in general relativity, where there is no preferred coordinate system at all, and in continuum mechanics, where the material has no reason to know which way you pointed your $x$-axis.

General tensor vs diagonal tensor
Left: the tensor in an arbitrary coordinate system, with all six independent components nonzero. Right: the same tensor in the principal-axis coordinate system, with all off-diagonal terms zero. Only three numbers remain. The choice of coordinate system changes the representation, not the physical content.

Tensors can have more than two indices. The elastic constants of a crystal, which relate the stress tensor $S_{ij}$ (internal forces) to the strain tensor $T_{ij}$ (deformations), form a rank-4 tensor $\gamma_{ijkl}$ with $3^4 = 81$ components in principle. Symmetry reduces this to 21 independent constants for the least symmetric crystal. For a cubic crystal, 3. For an isotropic material, just 2. The rank-4 tensor must have the form $\gamma_{ijkl} = a\,\delta_{ij}\delta_{kl} + b\,(\delta_{ik}\delta_{jl} + \delta_{il}\delta_{jk})$ in the isotropic case, because those are the only rank-4 combinations of $\delta_{ij}$ that have the required symmetry and pick out no preferred direction. Two constants, $a$ and $b$, and you have described the full elastic behaviour of any isotropic solid. The same pattern of symmetry reducing an enormous number of apparent components to a small set of physical ones repeats at every rank.

The rank ladder: scalar, vector, tensor
The hierarchy of tensors by rank. A scalar (rank 0) needs one number. A vector (rank 1) needs three components and one index. A rank-2 tensor needs nine components and two indices. At each step up in rank, you add one more direction index and multiply the number of components by three.

So here is the summary, stated as plainly as possible. A tensor is what you get when one physical vector quantity depends linearly on another, and the proportionality is not the same in every direction. The components of the tensor tell you, for each possible input direction, what the output is in each possible output direction. That is why a rank-2 tensor in three dimensions has $3 \times 3 = 9$ components: three possible input directions times three possible output directions. The components change when you rotate your axes, but the physical relationship they encode does not. For a symmetric tensor, you can always find principal axes where the off-diagonal components vanish, reducing the description to three numbers. And a tensor equation, once you have written it correctly, is automatically true in every coordinate system without any further work.

None of this required linear algebra. It required only the idea of proportionality, the fact that polarization is linear in the electric field, the simple geometry of ellipses, and the physical demand that energy be conserved. The full formalism of linear algebra makes the same ideas faster to compute with, but the ideas themselves are prior to the formalism. A tensor is a physical object. The indices are just how we write it down.

References and Footnotes

  • 1
    The subscript ordering $\alpha_{ij}$ means: $i$ labels the component of $\vec{P}$ being produced, $j$ labels the direction of $\vec{E}$ that produces it. So $\alpha_{yx}$ is the coefficient that tells you how much $y$-polarization is produced by an $x$-field. Keep this straight from the beginning or the indices will confuse you forever. ↩︎
  • 2
    This symmetry is a physical consequence of energy conservation, not a mathematical assumption. There exist tensors in physics that are not symmetric, for example the angular momentum flux tensor, and they genuinely need all nine components. Always ask which physical law enforces symmetry before counting independent components. ↩︎
  • 3
    The Kronecker delta $\delta_{ij}$ has the remarkable property that it looks identical in every coordinate system. You can rotate your axes by any angle and its components stay the same: ones on the diagonal, zeros everywhere else. This is because it is proportional to the metric of flat Euclidean space, which has no preferred direction. The identity tensor is the one tensor that genuinely has no directional content. ↩︎
  • 4
    The inertia tensor has the explicit form $I_{ij} = \sum_\text{particles} m(r^2 \delta_{ij} – r_i r_j)$, where the sum is over all particles in the body and $r^2 = x^2+y^2+z^2$. The diagonal component $I_{xx} = \sum m(y^2+z^2)$ is exactly the moment of inertia about the $x$-axis that you saw in introductory mechanics. The off-diagonal components $I_{xy} = -\sum m\, xy$ are the products of inertia, and they vanish when your axes align with the principal axes of the body. ↩︎
Posted in Expository | Tagged , , , , | 1 Comment