Newtonian mechanics contains two masses with no obvious reason to be related. One appears in the law of gravitation, $$\mathbf{F}=-\frac{GMm_g}{r^{2}}\hat{\mathbf r},$$ where $m_g$ measures how strongly a body couples to a gravitational field. The other appears in the second law, $$m_i\mathbf a=\mathbf F,$$ where $m_i$ measures resistance to being accelerated by any force at all and has nothing to do with gravity in particular. Combining them gives $$\mathbf a=-\frac{GM}{r^{2}}\frac{m_g}{m_i}\hat{\mathbf r},$$ so the acceleration of a falling body depends on the ratio $m_g/m_i$. Nothing in Newtonian mechanics requires that ratio to be the same for different materials. Experimentally, though, it’s universal to very high precision, so we set $m_i=m_g$ and write $$\mathbf a=-\frac{GM}{r^{2}}\hat{\mathbf r}.$$
The question worth asking is why gravity should be the only interaction whose trajectories don’t care what the test body is made of. For electromagnetism, $$\mathbf a=\frac{q}{m}\mathbf E,$$ so the motion depends explicitly on properties of the particle. A gravitational acceleration apparently doesn’t, and the consequence is that a free-fall trajectory is fixed entirely by an initial position and an initial velocity, with no other information about the falling body required.
That’s a strange property for a force law. It’s a much more ordinary property for geometry: once a starting point and a starting direction are given on a sphere, the corresponding great circle is fixed without asking what is moving along it. This doesn’t prove that gravity is geometric, but it suggests that the universality of free fall may be telling us something about spacetime itself.
The useful next question is how much of gravity can be removed by changing the observer. If the geometric suggestion is pointing in the right direction, then some of what we call a gravitational field may turn out to depend on how spacetime is being described, while some part of it should survive every such change. The way to find out is to start with the simplest possible experiment: a falling object inside an accelerating laboratory.
The way to test it is to see how much of gravity can be removed by a change of description. Put a laboratory in deep space, far from any mass, and accelerate it with proper acceleration $a$ using a rocket. Someone inside releases a ball; no force acts on it, so relative to the distant stars it moves at constant velocity while the floor accelerates upward to meet it, and to the occupant the ball appears to accelerate toward the floor at $a$. Every ball behaves identically, since the ball is doing nothing and the floor is doing all the work. Now stand the same laboratory on a planet where the gravitational acceleration happens to equal $a$. A released ball accelerates toward the floor at $a$, again independently of mass or composition, this time because of gravity. The two situations have completely different descriptions, one with an inertial ball and an accelerating lab, the other with a static lab and a ball being pulled, and no experiment with falling bodies performed inside a windowless box distinguishes them. Universality of free fall is what guarantees this; if different balls fell differently, dropping two of them would settle the question at once.1This is a reconstruction of the logic rather than the history. Einstein worked at the problem for most of a decade with several wrong turns, and the question here is whether we, knowing modern physics, can reconstruct why this structure makes sense.
Running the argument in reverse suggests that the effect of a gravitational field can be removed by accelerating appropriately, which for a field pointing downward means letting the laboratory fall. Cut the cable on an elevator near the Earth’s surface, and the occupant and everything with them accelerates downward at the same rate, so relative to the elevator nothing accelerates. A released ball stays where it’s put, two released balls hover next to each other, and a ball given a gentle push drifts across the elevator in a straight line at constant speed. Mechanics experiments performed inside return the answers of special relativity with no gravitational term in them. It’s tempting to say the gravitational field has been cancelled, and that phrasing causes trouble later, so it’s worth fixing the language now. What has happened is that at a chosen event a freely falling observer can adopt a frame in which the non-gravitational laws take their special-relativistic form and the gravitational acceleration doesn’t appear. Whether anything else survives is a separate question, and the answer to it is the whole point of what follows.
The statement that this holds for all of physics rather than only for mechanics is the equivalence principle, and the version needed here says that in a sufficiently small region of spacetime around any event, local non-gravitational experiments give the same results as in an inertial frame of special relativity with no gravitational field present. The word local is carrying most of the weight in that sentence. Take the freely falling elevator and make it very large, say a thousand kilometres tall, still falling toward the Earth. Release two balls, one near the ceiling and one near the floor. The lower ball sits where $GM/r^{2}$ is larger, so it accelerates slightly more strongly than the elevator average and drifts toward the floor, while the upper ball drifts toward the ceiling. Release two balls side by side instead. Both accelerate toward the Earth’s centre, so their acceleration vectors aren’t quite parallel and the balls drift toward each other. Inside a large enough freely falling laboratory, something is left over.
Local therefore means small enough that this residue sits below whatever precision you care about, and the comparison is quantitative rather than rhetorical. Near the Earth’s surface the tidal gradient is $2GM_\oplus/R_\oplus^{3}\approx3.1\times10^{-6}\ \mathrm{s^{-2}}$, so across a laboratory two metres tall the relative acceleration between top and bottom is about $6\times10^{-6}\ \mathrm{m\,s^{-2}}$, or roughly six parts in $10^{7}$ of $g$. An experiment that can’t resolve that level sees a two-metre freely falling lab as an inertial frame; one that can does not, and needs a smaller lab. There’s no absolute answer to how small counts as local, only an answer relative to the curvature scale and the precision of the measurement. The region also has to be small in time, since an elevator falling for an hour has moved a long way through a field that changes along the way.
So what exactly is the residue? Consider a single freely falling particle first. Construct a small frame that falls along with it, and by the equivalence principle that frame is inertial in a neighbourhood of each event on the worldline. In fact one can do better than event by event: a single frame can be built that stays locally inertial all along the worldline, so an observer riding with the particle finds no gravitational acceleration anywhere in their own history.2Coordinates of this kind are called Fermi normal coordinates. They put the metric in Minkowski form and set the connection coefficients to zero at every point of the central worldline, with corrections appearing only at second order in the distance away from it. Both of those objects are defined later in this post. What curvature obstructs is not the maintenance of such a frame along the curve, but its extension to a finite region surrounding the curve. Nothing local along the way distinguishes a gravitational field from empty space, and the accurate statement is that curvature can’t be detected by measurements confined to an ideally pointlike freely falling frame. A real observer with a gradiometer, or an interferometer with arms of finite length, measures tidal gravity perfectly well; what’s needed is a comparison between points at some finite separation, and the cleanest version of that comparison uses two nearby worldlines.
Release two particles a short distance apart in the freely falling elevator and watch the vector between them. Each particle is individually weightless and each has a local inertial frame in which nothing happens, and yet the separation changes. Two particles released side by side at the same altitude both fall toward the Earth’s centre, and since that centre is a single point their acceleration vectors converge, so the horizontal separation shrinks. Two released one above the other fall along the same radius, the lower one accelerates more strongly, and the vertical separation grows. A small spherical cloud of freely falling particles is therefore stretched radially and squeezed transversely. What no change of coordinates can do is make this covariant relative acceleration disappear. One can even choose coordinates in which both particles keep fixed spatial labels, but then their changing physical separation is carried by the metric rather than by their coordinate positions. It would be too quick to say that all observers agree about the distance between them, since coordinate distance and simultaneity are not invariant and different observers assign different components to the separation vector; what’s invariant is the relative acceleration of the two worldlines described covariantly, and that cannot be transformed away. Making that description precise is the business of the rest of this post.
In Newtonian language the residue has a clean form. Write the field as the gradient of a potential, $\mathbf{g}=-\nabla\Phi$, so a particle obeys $\ddot{x}^{i}=-\partial_{i}\Phi$. Take two particles, one at $x^{i}$ and one at $x^{i}+\xi^{i}$ with $\xi$ small, subtract their equations of motion and expand the gradient to first order in $\xi$: $$\ddot{\xi}^{i}=-\partial_{i}\partial_{j}\Phi\,\xi^{j}.$$ The relative acceleration of two nearby freely falling particles is governed by the Hessian of the potential. Not the potential itself, which can be shifted by a constant with no physical consequence, and not its gradient, which is exactly what falling absorbs. With $\Phi=-GM/r$ and $\partial_{i}\partial_{j}\Phi=GM(\delta_{ij}/r^{3}-3x_{i}x_{j}/r^{5})$, the eigenvalue along the radial direction is $-2GM/r^{3}$ and in each transverse direction $+GM/r^{3}$, so the radial separation grows and the transverse ones shrink. The three eigenvalues sum to zero, so a small cloud keeps its volume to leading order while its shape distorts, which is just $\nabla^{2}\Phi=0$ holding in the vacuum outside the Earth.
That equation is the target, and it’s also the obstacle. Everything in it is written in a fixed background space with a universal time, and it treats the components of a separation as though they meant something on their own. But the lesson of the falling elevator is that the acceleration of a body depends on who’s describing it: the occupant and the observer on the ground disagree about whether the ball accelerates, and neither of them is making a mistake. If part of gravity is an artefact of the description, then a claim that some other part is not an artefact has to be made in a language where the choice of description can’t influence the answer. Otherwise we have no way of telling a residue from a bookkeeping accident. So the next thing we need isn’t more physics, it’s a way of writing physics in which relabelling events changes nothing.
Suppose Alice describes the falling elevator using coordinates $x^{\mu}$ and Bob uses different coordinates $x’^{\mu}$. Nothing physical has changed; they’ve relabelled events. A coordinate system is a perfectly real and useful thing, a rule for attaching four numbers to each event, and Alice can measure everything she needs with it. What has no invariant meaning is the labels themselves and the components taken with respect to them: the statement that a clock reads $t=5$, or that a vector has components $(3,4)$, is a statement about a labelling scheme, whereas the proper time along a worldline is something a clock actually measures and comes out the same whoever computes it. The laws we’re looking for therefore can’t depend on which labels were used, and this becomes a working requirement: we want statements whose truth doesn’t depend on the coordinate system.
To make that requirement usable we need to be more careful about what a vector is, because the whole difficulty ahead is hiding in a place where the usual picture of an arrow is too loose. Fix an event $P$. Take all smooth curves through $P$ and, for each one, its velocity at $P$. These velocities add and scale like arrows, and together they form a four-dimensional vector space attached to $P$ and to no other point, written $T_{P}M$ and called the tangent space at $P$. A vector at $P$ is an element of $T_{P}M$, and that’s all it is; it isn’t an object stretching from one place to another. Choosing coordinates gives a basis $\mathbf{e}_{\mu}$ of $T_{P}M$, namely the velocities of the four curves along which one coordinate varies and the rest are held fixed, and the components $V^{\mu}$ are the numbers appearing in $\mathbf{V}=V^{\mu}\mathbf{e}_{\mu}$. A covector at $P$ is a linear map from $T_{P}M$ to the real numbers, which in plain terms is a machine that eats one vector and hands back a single number, in a way that respects adding and scaling. The gradient of a function is the standard example, since it takes a velocity and returns the rate at which the function changes along that curve, and its components carry a lower index. A tensor at $P$ is a multilinear map that takes some vectors and some covectors and returns a number, one index per slot; multilinear just means linear in each slot separately, so doubling what you feed into one slot doubles the answer and the other slots don’t notice. That’s the whole content of index notation: an index is a slot, an upper index is a slot for a covector and a lower index a slot for a vector, and the components are the numbers the map returns when fed basis elements. Notice what hasn’t been defined. Nothing so far relates $T_{P}M$ to $T_{Q}M$ at a different event, and that omission is about to become the central problem.
The transformation law follows from this. Under a change of coordinates the basis of $T_{P}M$ changes, so the components change to compensate, giving $$V’^{\mu}=\frac{\partial x’^{\mu}}{\partial x^{\nu}}V^{\nu},\qquad \omega’_{\mu}=\frac{\partial x^{\nu}}{\partial x’^{\mu}}\omega_{\nu},$$ for a vector and a covector, and one such factor per index for anything with more slots, as in $T’^{\mu}{}_{\nu}=\frac{\partial x’^{\mu}}{\partial x^{\alpha}}\frac{\partial x^{\beta}}{\partial x’^{\nu}}T^{\alpha}{}_{\beta}$. This is what makes tensors worth the notation. If $A^{\mu}{}_{\nu}=B^{\mu}{}_{\nu}$ holds in Alice’s system, then multiplying both sides by the same transformation factors gives $A’^{\mu}{}_{\nu}=B’^{\mu}{}_{\nu}$ in Bob’s, so a tensor equation is true in every coordinate system or in none. In particular, if every component of a tensor vanishes for Alice then every component vanishes for Bob, since each of Bob’s components is a linear combination of Alice’s. A tensor equation is therefore a statement about spacetime rather than about labels, which is exactly the kind of statement we were told to look for.
A concrete example makes the danger clear, and it will come back later. Draw an arrow on a flat tabletop. In one set of Cartesian axes it has components $(3,4)$, and its length follows from $dl^{2}=dx^{2}+dy^{2}$ as $5$. Rotate the axes to lie along the arrow and the components become $(5,0)$. Two different pairs of numbers, one arrow, and a length that never moved. Now describe the same tabletop in polar coordinates. The line element reads $$dl^{2}=dr^{2}+r^{2}d\theta^{2},$$ which looks nothing like $dx^{2}+dy^{2}$: there’s a coefficient that varies from place to place, as though the geometry were doing something at large $r$ that it wasn’t doing near the origin. The table is still a plane. What’s changed is that the polar basis vectors are genuinely different objects at different points, unlike the Cartesian ones: $\mathbf{e}_{r}$ points away from the origin, so on opposite sides of the origin it points in opposite directions, and $\mathbf{e}_{\theta}$ gets longer as $r$ grows, because a step $d\theta$ carries you a distance $r\,d\theta$. Take the unit radial vector field as a test case. Its components with respect to the local polar basis are $(1,0)$ at every point of the plane, and yet the arrows at two different points plainly point in different directions. Three things have to be held apart from here on: the geometric vector, which lives in the tangent space at one event; its components, which are just numbers; and the basis, which is what turns the numbers back into the vector. Only the first of the three is coordinate-independent, and equal components in a position-dependent basis don’t mean equal vectors. Hold on to that line element too, because we’ll meet it again in a more dangerous disguise.
Some notation, then. An event is a point of spacetime labelled by four coordinates $x^{\mu}$ with $\mu$ running over $0,1,2,3$, where $x^{0}=ct$, and Greek indices run over all four values while Latin indices $i,j,k$ run over the three spatial ones.3Two smaller conventions, used without further comment below. An index appearing once raised and once lowered is summed over, so $A_{\mu}B^{\mu}$ means $\sum_{\mu=0}^{3}A_{\mu}B^{\mu}$. And I keep $c$ explicit throughout rather than setting it to one, so that every formula can be checked against a Newtonian limit. In special relativity the geometry is carried by the interval between nearby events, $$ds^{2}=\eta_{\mu\nu}\,dx^{\mu}dx^{\nu}=-c^{2}dt^{2}+dx^{2}+dy^{2}+dz^{2},$$ with $\eta_{\mu\nu}$ the Minkowski metric, a constant array with diagonal $(-1,+1,+1,+1)$. That sign pattern is the signature, and I use $(-,+,+,+)$ everywhere below, which matters because many signs flip under the other convention.
In the language just set up, the object defining that interval is a tensor at each event. The metric at $P$ is a symmetric bilinear form $$g_{P}:T_{P}M\times T_{P}M\longrightarrow\mathbb{R},$$ a map taking two vectors at $P$ and returning a number, linear in each argument, symmetric under exchanging them, and non-degenerate, meaning the only vector orthogonal to everything is the zero vector. Its components are its values on the basis, $g_{\mu\nu}=g_{P}(\mathbf{e}_{\mu},\mathbf{e}_{\nu})$, and symmetry leaves ten independent ones at each event, with matrix inverse $g^{\mu\nu}$ defined by $g^{\mu\alpha}g_{\alpha\nu}=\delta^{\mu}_{\nu}$. The interval itself is what that form returns when fed a displacement twice over, which is a number rather than a tensor. Special relativity is the case where coordinates exist making the components $\eta_{\mu\nu}$ everywhere at once. Dropping that requirement and letting them depend on position, $$ds^{2}=g_{\mu\nu}(x)\,dx^{\mu}dx^{\nu},$$ costs nothing in interpretation, because every physical use of the metric is local anyway. A clock carried along a timelike worldline with four-velocity $u$ accumulates proper time at a rate fixed by $g(u,u)=-c^{2}$; the length of a spatial separation is $\sqrt{g(\xi,\xi)}$ for a spacelike vector $\xi$ lying in an observer’s local rest space; two vectors are orthogonal when $g(A,B)=0$; and a direction is null, so a light ray can go that way, when $g(k,k)=0$, which is what fixes the light cones and with them the causal structure. Calling $g_{\mu\nu}$ a gravitational potential isn’t exactly wrong, but it undersells it: it’s the object that tells you what every clock, ruler and light signal in a small neighbourhood will do.
With the metric available, the equivalence principle turns into a statement about coordinates. Around any event $P$ we can choose coordinates in which $$g_{\mu\nu}(P)=\eta_{\mu\nu},\qquad \partial_{\alpha}g_{\mu\nu}(P)=0.$$ The first condition says that the geometry at $P$ is Minkowskian, which it always can be made to be, since a symmetric matrix can be brought to diagonal form with entries $\pm1$ and the signature fixes which. The second says the metric components are stationary at $P$, so to first order in the distance from $P$ the geometry looks flat. I want to resist saying that the first derivatives of the metric are the gravitational acceleration, because that isn’t quite what they are, and the precise version has to wait until we have the object they actually build. The obvious next question is whether we can go further and remove the second derivatives too. We’ll be in a position to answer that properly later, so I’ll leave it standing.
What we can’t do yet is compare anything at two different events, and this is where the tangent-space definition starts to bite. A vector at $P$ belongs to $T_{P}M$ and a vector at a neighbouring event $Q$ belongs to $T_{Q}M$, and these are two different vector spaces. Subtracting one from the other is not a defined operation. But subtraction is exactly what differentiation is: to say how a vector field changes near $P$ we have to take the vector at $Q$, compare it with the vector at $P$, and divide by the separation. The components, being numbers, can of course be subtracted, and that’s what $\partial_{\mu}V^{\nu}$ does, but the answer has no meaning on its own because it mixes a change in the vector with a change in the basis. The polar example already showed this: the radial field has constant components and is not a constant field.
The failure shows up algebraically as well, and it’s worth seeing both faces of it. If $\partial_{\nu}V^{\mu}$ were a tensor it would transform with two factors of the coordinate change. Differentiate the transformation law and see what happens. Using the chain rule for the outer derivative and then the product rule, $$\partial’_{\alpha}V’^{\mu}=\frac{\partial x^{\beta}}{\partial x’^{\alpha}}\partial_{\beta}\left(\frac{\partial x’^{\mu}}{\partial x^{\nu}}V^{\nu}\right)=\frac{\partial x^{\beta}}{\partial x’^{\alpha}}\frac{\partial x’^{\mu}}{\partial x^{\nu}}\,\partial_{\beta}V^{\nu}+\frac{\partial x^{\beta}}{\partial x’^{\alpha}}\frac{\partial^{2}x’^{\mu}}{\partial x^{\beta}\partial x^{\nu}}\,V^{\nu}.$$ The first term is the tensor transformation law we wanted. The second is an obstruction, and it contains second derivatives of the coordinate change, so it vanishes only for linear relabellings. A rotation of Cartesian axes is linear and does no damage, which is why nobody meets this problem in elementary vector calculus; a change to polar coordinates is not linear, and it does. The two faces are the same fact: the obstruction term is precisely the failure to have a canonical way of comparing $T_{P}M$ with $T_{Q}M$.
So the logic is forced. Differentiation requires comparing neighbouring vectors; no comparison exists by default; therefore one has to be supplied as extra structure. That structure is called a connection, and supplying it means saying how the basis vectors change from point to point, $$\nabla_{\nu}\mathbf{e}_{\mu}=\Gamma^{\rho}_{\ \nu\mu}\,\mathbf{e}_{\rho},$$ which declares what it means to carry a vector from one tangent space to the next without changing it. Given that rule, differentiating a field $\mathbf{V}=V^{\mu}\mathbf{e}_{\mu}$ is just the product rule, $$\nabla_{\nu}\mathbf{V}=(\partial_{\nu}V^{\rho})\mathbf{e}_{\rho}+V^{\mu}\nabla_{\nu}\mathbf{e}_{\mu}=\left(\partial_{\nu}V^{\rho}+\Gamma^{\rho}_{\ \nu\mu}V^{\mu}\right)\mathbf{e}_{\rho},$$ so that in components $$\nabla_{\mu}V^{\nu}=\partial_{\mu}V^{\nu}+\Gamma^{\nu}_{\ \mu\lambda}V^{\lambda}.$$ The first term asks how the components changed, the second corrects for the fact that the basis changed too, and the sum is ordinary differentiation with the change of basis subtracted off. For a lower index the correction comes with the opposite sign, $\nabla_{\mu}V_{\nu}=\partial_{\mu}V_{\nu}-\Gamma^{\lambda}_{\ \mu\nu}V_{\lambda}$, and a general tensor picks up one $\Gamma$ term per slot.
The coefficients themselves can’t be tensors, and the reason is structural rather than accidental. For $\nabla_{\mu}V^{\nu}$ to transform correctly, the $\Gamma$ term has to cancel the obstruction found above, so the $\Gamma$ must transform with an inhomogeneous piece of their own, $$\Gamma’^{\rho}_{\ \mu\nu}=\frac{\partial x’^{\rho}}{\partial x^{\lambda}}\frac{\partial x^{\alpha}}{\partial x’^{\mu}}\frac{\partial x^{\beta}}{\partial x’^{\nu}}\Gamma^{\lambda}_{\ \alpha\beta}+\frac{\partial x’^{\rho}}{\partial x^{\lambda}}\frac{\partial^{2}x^{\lambda}}{\partial x’^{\mu}\partial x’^{\nu}},$$ where the first term is what a tensor would do and the second is the repair. The second derivatives of the coordinate change appear here with exactly the structure needed to cancel the ones in the obstruction. So a connection is not a tensor, and an equation saying $\Gamma=0$ in some coordinates carries no invariant information at all. That will matter quite a lot in a moment.
Nothing so far says which connection to take, and in fact any set of coefficients with the right transformation behaviour defines a legitimate covariant derivative. General relativity makes two additional demands, and it’s worth stating them separately because they’re independent and neither is a theorem. The first is metric compatibility, $$\nabla_{\rho}g_{\mu\nu}=0,$$ which says that transporting two vectors along a curve preserves their inner product, so lengths and angles are carried unchanged. The second is vanishing torsion, $$\Gamma^{\rho}_{\ \mu\nu}=\Gamma^{\rho}_{\ \nu\mu},$$ symmetry in the two lower indices, which says that the failure of an infinitesimal parallelogram to close is not part of the structure. Metric compatibility doesn’t imply the symmetry, and theories keeping one without the other are perfectly consistent.4Einstein-Cartan theory is the standard example, where torsion is sourced by intrinsic spin and vanishes outside matter, so the vacuum predictions coincide with general relativity. These are choices that define standard general relativity, not mathematical necessities, and together they pin the connection down completely.
The derivation is short. Expanding metric compatibility with one $\Gamma$ per index gives $\partial_{\rho}g_{\mu\nu}=\Gamma_{\nu\rho\mu}+\Gamma_{\mu\rho\nu}$, where $\Gamma_{\alpha\beta\gamma}\equiv g_{\alpha\lambda}\Gamma^{\lambda}_{\ \beta\gamma}$. Write it three times with the indices permuted, $$\partial_{\mu}g_{\nu\rho}=\Gamma_{\rho\mu\nu}+\Gamma_{\nu\mu\rho},\qquad \partial_{\nu}g_{\rho\mu}=\Gamma_{\mu\nu\rho}+\Gamma_{\rho\nu\mu},\qquad \partial_{\rho}g_{\mu\nu}=\Gamma_{\nu\rho\mu}+\Gamma_{\mu\rho\nu},$$ then add the first two and subtract the third. Symmetry in the last two indices makes $\Gamma_{\nu\mu\rho}$ cancel against $\Gamma_{\nu\rho\mu}$ and $\Gamma_{\mu\nu\rho}$ against $\Gamma_{\mu\rho\nu}$, leaving $2\Gamma_{\rho\mu\nu}$. Raising the first index, $$\Gamma^{\rho}_{\ \mu\nu}=\tfrac12 g^{\rho\sigma}\left(\partial_{\mu}g_{\sigma\nu}+\partial_{\nu}g_{\sigma\mu}-\partial_{\sigma}g_{\mu\nu}\right),$$ the Christoffel symbols of the Levi-Civita connection. Every term is a first derivative of the metric, which is the fact we’ll use twice over in what follows.
It’s worth stopping to see what has and hasn’t been achieved. We set out to find the part of gravity that free fall can’t remove, and we haven’t found it. What we’ve built is the machinery for comparing directions at different events, which we needed because without it we couldn’t write down a single coordinate-independent statement about how anything varies from place to place. The connection is that machinery, and by construction it isn’t a tensor, so it can’t itself be the invariant we’re after. What it does give us is a way of saying what free motion means, and that’s the next thing to check.
A particle moves freely if its velocity is carried along unchanged by the transport the geometry provides, which is to say $$\nabla_{u}u=0,\qquad u^{\mu}=\frac{dx^{\mu}}{d\tau},\qquad g_{\mu\nu}u^{\mu}u^{\nu}=-c^{2}.$$ Writing it in components, $(\nabla_{u}u)^{\mu}=u^{\nu}\nabla_{\nu}u^{\mu}=u^{\nu}(\partial_{\nu}u^{\mu}+\Gamma^{\mu}_{\ \nu\lambda}u^{\lambda})$. The first piece simplifies because of what $u^{\nu}\partial_{\nu}$ means. The four-velocity is tangent to the worldline, so contracting it with a gradient is the directional derivative along that curve, and for any quantity defined along the worldline the chain rule gives $u^{\nu}\partial_{\nu}f=\frac{dx^{\nu}}{d\tau}\frac{\partial f}{\partial x^{\nu}}=\frac{df}{d\tau}$. Applied to $u^{\mu}$ this turns $u^{\nu}\partial_{\nu}u^{\mu}$ into $du^{\mu}/d\tau=d^{2}x^{\mu}/d\tau^{2}$, and the whole condition becomes $$\frac{d^{2}x^{\mu}}{d\tau^{2}}+\Gamma^{\mu}_{\ \alpha\beta}\frac{dx^{\alpha}}{d\tau}\frac{dx^{\beta}}{d\tau}=0,$$ the geodesic equation. A curve satisfying it parallel-transports its own tangent vector, which is the closest available statement of going straight, and when the metric components are constant every $\Gamma$ vanishes and it reduces to $d^{2}x^{\mu}/d\tau^{2}=0$.
The same equation follows from a completely different demand, and the agreement is worth noting because it isn’t obvious. In special relativity, among timelike paths joining two events the straight one has the largest proper time; that’s the twin paradox as a variational statement. Extremising $\tau=\frac{1}{c}\int\sqrt{-g_{\mu\nu}dx^{\mu}dx^{\nu}}$ is awkward because of the square root. It’s convenient to use an affine parameter $\lambda$ along the path and extremise instead $$S=\frac12\int g_{\mu\nu}\frac{dx^{\mu}}{d\lambda}\frac{dx^{\nu}}{d\lambda}\,d\lambda,$$ the same trick as using $\tfrac12 mv^{2}$ rather than a path length in ordinary mechanics. This gives the same curves, and for a timelike geodesic we may afterwards take $\lambda=\tau$, up to an affine rescaling. Run the Euler-Lagrange equations on that integrand and the derivatives of the metric assemble themselves, without being asked, into precisely the combination that defines $\Gamma^{\sigma}_{\ \mu\nu}$, and the geodesic equation comes back out.5The steps, for anyone checking: with $L=\tfrac12 g_{\mu\nu}\dot{x}^{\mu}\dot{x}^{\nu}$ and a dot meaning $d/d\lambda$, we have $\partial L/\partial\dot{x}^{\rho}=g_{\rho\nu}\dot{x}^{\nu}$, whose derivative along the curve is $\partial_{\mu}g_{\rho\nu}\dot{x}^{\mu}\dot{x}^{\nu}+g_{\rho\nu}\ddot{x}^{\nu}$, set against $\partial L/\partial x^{\rho}=\tfrac12\partial_{\rho}g_{\mu\nu}\dot{x}^{\mu}\dot{x}^{\nu}$. The first term is contracted with something symmetric in $\mu\nu$ and may be symmetrised, and multiplying through by $g^{\sigma\rho}$ produces the Christoffel combination. So a geodesic has two independent characterisations: it parallel-transports its own tangent, and it extremises proper time. That they agree is not automatic. The first notion needs only a connection and the second needs only a metric, and they coincide because we chose the connection to be metric-compatible and torsion-free; for a general connection the autoparallels and the extremal-proper-time curves are different families.6Timelike geodesics extremise proper time locally and are genuinely maximal for sufficiently nearby events, but over long paths families of geodesics can refocus at conjugate points, beyond which a geodesic is no longer the longest path between its endpoints. I say extremise rather than maximise throughout for that reason.
Now go back to the freely falling frame and put the pieces together. In those coordinates $g_{\mu\nu}(P)=\eta_{\mu\nu}$ and $\partial_{\alpha}g_{\mu\nu}(P)=0$, and every term in the Levi-Civita formula is a first derivative of the metric, so $$\Gamma^{\rho}_{\ \mu\nu}(P)=\tfrac12 g^{\rho\sigma}(0+0-0)=0.$$ This is the precise statement of what free fall accomplishes, and it’s worth spelling out the chain rather than compressing it. The first derivatives of the metric determine the connection; the connection coefficients are what appear in the geodesic equation; and there they supply the entire coordinate acceleration of a freely moving particle, since $d^{2}x^{\mu}/d\tau^{2}=-\Gamma^{\mu}_{\ \alpha\beta}\dot{x}^{\alpha}\dot{x}^{\beta}$. Kill the $\Gamma$ at $P$ and a free particle has no coordinate acceleration at $P$, which is the occupant of the elevator watching the ball hang in the air. The first derivatives of the metric aren’t themselves the gravitational field; they’re the ingredients of the object that shows up as acceleration once you ask how a free particle moves.
At this point everything appears to line up. The geodesic equation rearranges to acceleration equals minus a $\Gamma$ term, sitting in exactly the slot where $-\nabla\Phi$ sits in Newton’s law, and the connection coefficients are built from first derivatives of the metric, just as the Newtonian field is the first derivative of the potential. On that basis you’ll see the $\Gamma$ called the gravitational field, and it’s hard to see what else they could be.
So compute them for a space with no gravity in it at all. Take the flat tabletop from earlier in polar coordinates, $$dl^{2}=dr^{2}+r^{2}d\theta^{2},$$ so $g_{rr}=1$, $g_{\theta\theta}=r^{2}$, and the inverse components are $g^{rr}=1$, $g^{\theta\theta}=1/r^{2}$. The only nonvanishing derivative is $\partial_{r}g_{\theta\theta}=2r$, and feeding that into the Levi-Civita formula gives $$\Gamma^{r}_{\ \theta\theta}=-\tfrac12 g^{rr}\partial_{r}g_{\theta\theta}=-r,\qquad \Gamma^{\theta}_{\ r\theta}=\Gamma^{\theta}_{\ \theta r}=\tfrac12 g^{\theta\theta}\partial_{r}g_{\theta\theta}=\frac1r,$$ with every other component zero. The space is a plane. There’s no gravitational field anywhere, no matter anywhere, nothing curved about it, and the connection coefficients are not zero. The geodesic equation becomes $\ddot{r}-r\dot{\theta}^{2}=0$ together with an equation for $\theta$, whose solutions are straight lines written in polar form, and that $-r\dot{\theta}^{2}$ is the centrifugal term from first-year mechanics, appearing here as a connection coefficient rather than a force. Nonzero Christoffel symbols do not imply curvature. They record that the coordinate basis rotates and stretches from point to point, which is a fact about the grid, and the warning attached to this same line element earlier was exactly this.
The same thing happens in spacetime. An observer with constant proper acceleration in flat Minkowski spacetime can use coordinates adapted to their own motion, called Rindler coordinates, in which the metric components aren’t $\eta_{\mu\nu}$, the connection coefficients are nonzero, and objects released from rest appear to accelerate away. Locally that observer could reasonably call what they feel a gravitational field, while the spacetime is exactly flat. Put that beside the calculation two paragraphs above and both directions are now established. In flat spacetime we can have $\Gamma^{\rho}_{\ \mu\nu}\neq0$, and in curved spacetime we can have $\Gamma^{\rho}_{\ \mu\nu}(P)=0$ at any event we like. A quantity that can be made nonzero where there’s no curvature and zero where there is curvature cannot be the invariant measure of curvature, and this is not an argument about intuition; it follows from the inhomogeneous transformation law, since an object with a non-tensorial piece can always be shifted by a suitable relabelling. Whatever survives free fall isn’t the connection.
Which brings back the question left standing earlier: how far can a change of coordinates go? Counting the freedom gives a rough sense of where it runs out, and the count is supporting intuition rather than proof, so it’s worth doing quickly and not leaning on it too hard. A coordinate change near $P$ is specified by its Taylor coefficients: the first derivatives $\partial x’^{\mu}/\partial x^{\alpha}$ form a $4\times4$ array with $16$ entries; the second derivatives are symmetric in their two lower indices, giving $10$ pairs for each of $4$ values of $\mu$, so $40$; the third derivatives are symmetric in three lower indices, giving $20$ triples for each of $4$, so $80$. What we’d like removed is the metric and its derivatives: $g_{\mu\nu}$ has $10$ independent components, $\partial_{\alpha}g_{\mu\nu}$ has $40$, and $\partial_{\alpha}\partial_{\beta}g_{\mu\nu}$, symmetric in $\alpha\beta$ and in $\mu\nu$ separately, has $100$. Setting $g_{\mu\nu}=\eta_{\mu\nu}$ imposes $10$ conditions on $16$ parameters and leaves $6$ over, which are the dimension of the Lorentz group and the residual freedom to boost and rotate a local inertial frame that special relativity says should be there. Setting $\partial_{\alpha}g_{\mu\nu}=0$ imposes $40$ conditions with exactly $40$ parameters. That matching count is consistent with the existence of locally inertial coordinates, although the count itself is not a proof of it. Setting $\partial_{\alpha}\partial_{\beta}g_{\mu\nu}=0$ would impose $100$ conditions with only $80$ parameters, and it fails, short by $20$.
| What we try to set at $P$ | Conditions | Coordinate freedom | Result |
|---|---|---|---|
| $g_{\mu\nu}(P)=\eta_{\mu\nu}$ | 10 symmetric pair $(\mu\nu)$: $\frac{4(4+1)}{2}=10$ | 16 $\displaystyle \frac{\partial x’^{\mu}} {\partial x^{\alpha}}$ $4\times4=16$ | possible 6 parameters remain |
| $\partial_{\alpha}g_{\mu\nu}(P)=0$ | 40 $4$ choices for $\alpha$ $\times\,10$ symmetric $(\mu\nu)$ | 40 $\displaystyle \frac{\partial^2x’^{\mu}} {\partial x^{\alpha}\partial x^{\beta}}$ $4\times10=40$ | possible |
| $\partial_{\alpha}\partial_{\beta} g_{\mu\nu}(P)=0$ | 100 $10$ symmetric $(\alpha\beta)$ $\times\,10$ symmetric $(\mu\nu)$ | 80 $\displaystyle \frac{\partial^3x’^{\mu}} {\partial x^{\alpha} \partial x^{\beta} \partial x^{\gamma}}$ $4\times20=80$ | not all possible short by 20 |
Take that as a hint about where to look rather than as a result. The tally treats coordinate changes as though they acted on the metric derivatives one parameter at a time, which they don’t, and it proves nothing about what the leftover information is. Two things it does suggest. First, the coordinate freedom available to remove the metric’s first derivatives is just sufficient, while at second derivatives something is left that no relabelling reaches, which is the same order at which the Newtonian residue lived. Second, the size of the shortfall is consistent with the fact that the Riemann tensor in four dimensions has twenty independent components, which is a remark worth filing and not an argument for its existence or structure. The distinction to keep hold of is that the raw second derivatives $\partial_{\alpha}\partial_{\beta}g_{\mu\nu}$ are themselves coordinate-dependent, and individual components can be changed at will by a different choice of labels. What can’t be changed is a particular combination of them together with products of first derivatives, and finding that combination is a separate job the count can’t do for us.
One picture is worth clearing away before we start. The bowling ball on a stretched rubber sheet explains the marble’s motion by appealing to real gravity pulling it into the dip, so it uses gravity to explain gravity; it shows two spatial dimensions, so what it depicts is spatial curvature rather than the curvature of spacetime; and it hides the component that actually matters, since we’ll find in the weak-field limit that what makes a slowly moving body fall comes from $g_{00}$, which is a feature of the temporal part of the geometry that no curved spatial surface can display. We’ll work intrinsically instead, with measurements made inside the geometry and no reference to a higher-dimensional space for it to sit in.
The intrinsic notion we need is already in hand, because a connection is exactly a rule for carrying a vector from one tangent space to the next. Carry one along a curve, keeping it parallel by that rule at every step, and you have parallel transport. Now the question that matters: if two different curves run from $P$ to $Q$, do they deliver the same vector at $Q$? Equivalently, transport a vector around a closed loop and ask whether it returns to itself. In a simply connected flat region with the ordinary rule it does. On a sphere it doesn’t. Stand on the equator with an arrow pointing north, carry it a quarter of the way around the equator keeping it parallel to itself, then up a meridian to the pole, then back down the next meridian to where you started, and the arrow comes home rotated by a right angle. Nothing twisted it. The rotation is measurable from within the surface, with no reference to any space the sphere might sit in, and its size depends on the loop: for a sphere of radius $R$ the angle equals the enclosed area divided by $R^{2}$, so shrinking the loop shrinks the effect in proportion to the area. This path-dependence is holonomy. For an infinitesimal loop, its leading behaviour is controlled locally by the curvature tensor, which is the object we’re about to build, and that’s the route we’ll take: holonomy is built from the connection, which isn’t a tensor, and yet it’s stated as a comparison between two transports, which is a question with an observer-independent answer. It also explains the failure of the coordinate count in geometric terms: each small patch carries a good flat frame, and carrying that frame around a loop doesn’t return it to itself, so the patches can’t be assembled into one flat frame over a finite region.
To get a formula we make the loop infinitesimal. Transport around a small rectangle with sides along the $\mu$ and $\nu$ directions means going out one way and back the other, and comparing the two routes is the same as differentiating first in one direction and then the other and asking whether the order matters. So compute the commutator of two covariant derivatives. Write $W_{\nu}{}^{\rho}\equiv\nabla_{\nu}V^{\rho}$, which has one upper and one lower index, so differentiating it covariantly needs one $\Gamma$ for each: $$\nabla_{\mu}\nabla_{\nu}V^{\rho}=\partial_{\mu}W_{\nu}{}^{\rho}+\Gamma^{\rho}_{\ \mu\lambda}W_{\nu}{}^{\lambda}-\Gamma^{\lambda}_{\ \mu\nu}W_{\lambda}{}^{\rho}.$$ Substituting $W_{\nu}{}^{\rho}=\partial_{\nu}V^{\rho}+\Gamma^{\rho}_{\ \nu\lambda}V^{\lambda}$ and expanding, $$\nabla_{\mu}\nabla_{\nu}V^{\rho}=\partial_{\mu}\partial_{\nu}V^{\rho}+(\partial_{\mu}\Gamma^{\rho}_{\ \nu\lambda})V^{\lambda}+\Gamma^{\rho}_{\ \nu\lambda}\partial_{\mu}V^{\lambda}+\Gamma^{\rho}_{\ \mu\lambda}\partial_{\nu}V^{\lambda}+\Gamma^{\rho}_{\ \mu\lambda}\Gamma^{\lambda}_{\ \nu\sigma}V^{\sigma}-\Gamma^{\lambda}_{\ \mu\nu}\left(\partial_{\lambda}V^{\rho}+\Gamma^{\rho}_{\ \lambda\sigma}V^{\sigma}\right).$$ Now write the same thing with $\mu$ and $\nu$ exchanged and subtract, term by term. The second derivatives go first: $\partial_{\mu}\partial_{\nu}V^{\rho}-\partial_{\nu}\partial_{\mu}V^{\rho}=0$, since partial derivatives commute. The terms carrying a first derivative of $V$ go next: the pair $\Gamma^{\rho}_{\ \nu\lambda}\partial_{\mu}V^{\lambda}+\Gamma^{\rho}_{\ \mu\lambda}\partial_{\nu}V^{\lambda}$ is already symmetric under exchanging $\mu$ and $\nu$, so it cancels against itself. The last bracket cancels too, because $\Gamma^{\lambda}_{\ \mu\nu}$ is symmetric in its lower indices, which is where the torsion-free assumption earns its keep.7With torsion the antisymmetric part of $\Gamma^{\lambda}_{\ \mu\nu}$ survives here and contributes an extra term $-T^{\lambda}{}_{\mu\nu}\nabla_{\lambda}V^{\rho}$ to the commutator. Dropping it is a consequence of the choice made earlier, not an algebraic accident. Everything containing a derivative of $V$ has now gone, which is the important structural fact: what’s left acts on $V$ itself, with no derivatives, so it’s a linear map on vectors rather than a differential operator. What’s left is $$\left(\nabla_{\mu}\nabla_{\nu}-\nabla_{\nu}\nabla_{\mu}\right)V^{\rho}=\left(\partial_{\mu}\Gamma^{\rho}_{\ \nu\sigma}-\partial_{\nu}\Gamma^{\rho}_{\ \mu\sigma}+\Gamma^{\rho}_{\ \mu\lambda}\Gamma^{\lambda}_{\ \nu\sigma}-\Gamma^{\rho}_{\ \nu\lambda}\Gamma^{\lambda}_{\ \mu\sigma}\right)V^{\sigma},$$ and we give the bracket a name, $$R^{\rho}_{\ \sigma\mu\nu}=\partial_{\mu}\Gamma^{\rho}_{\ \nu\sigma}-\partial_{\nu}\Gamma^{\rho}_{\ \mu\sigma}+\Gamma^{\rho}_{\ \mu\lambda}\Gamma^{\lambda}_{\ \nu\sigma}-\Gamma^{\rho}_{\ \nu\lambda}\Gamma^{\lambda}_{\ \mu\sigma},$$ the Riemann curvature tensor. It’s worth being clear that this is not a formula someone wrote down and then found a use for. It’s the coefficient left behind when covariant derivatives fail to commute, and the expression is whatever the cancellation leaves.
The indices now say something definite. The left side is linear in $V$, so $\sigma$ labels the component of the vector fed in and $\rho$ the component of the vector that comes back out, exactly as for any linear map on the tangent space. The pair $\mu\nu$ is antisymmetric, since exchanging them changes the sign of the commutator, and an antisymmetric pair of directions is an oriented two-plane: it specifies the infinitesimal patch the transport was taken around and which way round it was traversed. So for each oriented two-plane at an event, $R^{\rho}_{\ \sigma\mu\nu}$ is a linear map taking a vector to the change it suffers when carried around that patch. Four indices is the least this can be done with, and it also explains why the object is a tensor even though the $\Gamma$ are not. The left side is manifestly a tensor: it’s built from covariant derivatives, each of which was constructed to transform correctly, and the inhomogeneous pieces of the $\Gamma$ cancel there rather than surviving. Since the left side is a tensor and $V^{\sigma}$ is arbitrary, the coefficient has to be one too. This is the identification we were promised, although the usual shorthand for it needs one qualification. People say the curvature is a combination of second derivatives of the metric, and it’s worth seeing why that’s only half the story. Each $\Gamma$ already carries one derivative of the metric, so the definition above is schematically $$R\sim\partial^{2}g+(\partial g)^{2},$$ with both kinds of term present. The quadratic piece isn’t decoration; it’s part of what makes the non-tensorial contributions cancel, and an object built from $\partial^{2}g$ alone wouldn’t be a tensor. The two terms separate only in a locally inertial frame at $P$, where $\Gamma(P)=0$ kills the quadratic part at that one event and leaves the curvature there expressed through second derivatives of the metric alone. With that understood, the Riemann tensor is invariant where the raw second derivatives weren’t, and being a tensor, nonzero in one coordinate system means nonzero in all of them.
With it in hand we can say how local a local inertial frame really is. In the locally inertial coordinates at $P$, usually called Riemann normal coordinates, the metric has the expansion $$g_{\mu\nu}(x)=\eta_{\mu\nu}-\tfrac13 R_{\mu\alpha\nu\beta}(P)\,x^{\alpha}x^{\beta}+O(x^{3}).$$ Coordinates of this kind can be built around any event of any spacetime; that’s a standard result which I’ll take on trust rather than prove, since what we need is not the construction but what the expansion says. Read it carefully. The constant term is the first condition free fall achieved and the absence of a linear term is the second, and the content of the expansion is that in such coordinates the first invariant information in the metric appears at quadratic order and is controlled by the curvature. That’s not the same as saying the leftover second derivatives are the curvature components; it’s saying that once the coordinate freedom has been spent, what remains at that order is governed by a tensor. Practically, a laboratory of size $L$ around $P$ departs from flat geometry by a fractional amount of order $\mathcal{R}L^{2}$, where $\mathcal{R}$ stands for the characteristic size of the relevant components of the Riemann tensor rather than the Ricci scalar or any other particular contraction, and that’s the number to hold against your experimental precision. It’s worth keeping this apart from the estimate made by hand much earlier, because the two measure different things. That one was mechanical, a tidal relative acceleration compared with $g$, and gave a few parts in $10^{7}$ across a two-metre laboratory. This one is a dimensionless departure of the metric from Minkowski form, and near the Earth the characteristic curvature is $\mathcal{R}\sim GM_\oplus/c^{2}R_\oplus^{3}\sim10^{-23}\ \mathrm{m^{-2}}$, so $\mathcal{R}L^{2}\sim10^{-22}$ for the same laboratory. The gap between the two numbers is a factor of order $gL/c^{2}$, which is what separates a statement about accelerations from a statement about geometry.
The polar example settles immediately, which is a useful check on the machinery. Substituting $\Gamma^{r}_{\ \theta\theta}=-r$ and $\Gamma^{\theta}_{\ r\theta}=1/r$ into the definition, the derivative terms cancel against the quadratic terms and every component vanishes, so $R^{\rho}_{\ \sigma\mu\nu}=0$ for a plane in any coordinates whatsoever. The connection was a fact about the grid; the curvature is a fact about the surface.
| Quantity | What it tells us | What it means |
|---|---|---|
| $\Gamma^{\rho}{}_{\mu\nu}\neq 0$ | The coordinate basis changes from point to point. | Not necessarily curvature. Christoffel symbols can be nonzero even in flat spacetime, as in polar coordinates. |
| $R^{\rho}{}_{\sigma\mu\nu}\neq 0$ | Parallel transport depends on the path. | Intrinsic curvature. No coordinate transformation can make the Riemann tensor vanish at that event. |
| $\displaystyle \frac{D^2\xi^{\mu}}{D\tau^2}\neq 0$ | Nearby freely falling worldlines accelerate relative to one another. | An observable tidal effect. For a chosen four-velocity and separation, curvature produces geodesic deviation. |
What remains is to connect this tensor to the thing we started from, the relative acceleration of two nearby falling bodies, and the derivation is short enough to do properly. Take a smooth one-parameter family of geodesics and label them by $s$, so that $x^{\mu}(\tau,s)$ is the point reached at proper time $\tau$ along the geodesic numbered $s$. The family sweeps out a two-dimensional surface in spacetime on which $\tau$ and $s$ serve as coordinates, and it carries two natural vector fields, $$u^{\mu}=\frac{\partial x^{\mu}}{\partial\tau},\qquad \xi^{\mu}=\frac{\partial x^{\mu}}{\partial s}.$$ The first is the four-velocity, tangent to each geodesic. The second points from a given geodesic to its neighbour at the same value of $\tau$, and it’s the mathematical version of the vector between the two balls in the elevator. Because $\tau$ and $s$ are coordinates on that surface, $u$ and $\xi$ are the basis fields belonging to those two directions, and mixed partial derivatives commute. Said without notation: move a little way along one geodesic and then step across to its neighbour, or step across first and move along afterwards, and you arrive at the same event either way. The object measuring the failure of two such routes to agree is called the Lie bracket $[u,\xi]$, and what we’ve just observed is that here it vanishes, $[u,\xi]=0$. For a torsion-free connection the bracket can be written with covariant derivatives, $[u,\xi]^{\mu}=u^{\nu}\nabla_{\nu}\xi^{\mu}-\xi^{\nu}\nabla_{\nu}u^{\mu}$, since the $\Gamma$ terms cancel by symmetry in the lower indices, and therefore $$\nabla_{u}\xi=\nabla_{\xi}u.$$ That’s the one nontrivial input, and it’s again the torsion-free choice doing work.
Now differentiate the separation twice along the worldline. Using the relation just obtained, $$\frac{D^{2}\xi^{\mu}}{D\tau^{2}}=\nabla_{u}\nabla_{u}\xi^{\mu}=\nabla_{u}\nabla_{\xi}u^{\mu}.$$ To evaluate the right side we swap the order of the two derivatives, which is precisely what the curvature tensor was built to account for, $$\nabla_{u}\nabla_{\xi}u^{\mu}=\nabla_{\xi}\nabla_{u}u^{\mu}+u^{\alpha}\xi^{\beta}\left(\nabla_{\alpha}\nabla_{\beta}-\nabla_{\beta}\nabla_{\alpha}\right)u^{\mu},$$ where the bracket term is the commutator acting on $u^{\mu}$ and there’s no extra contribution because $[u,\xi]=0$. The first term on the right vanishes, since each curve of the family is a geodesic and $\nabla_{u}u=0$. The second is the commutator identity, which gives $u^{\alpha}\xi^{\beta}R^{\mu}_{\ \nu\alpha\beta}u^{\nu}$. So $$\frac{D^{2}\xi^{\mu}}{D\tau^{2}}=R^{\mu}_{\ \nu\alpha\beta}\,u^{\nu}u^{\alpha}\xi^{\beta}.$$ One step is left, and it’s the one that’s usually skipped. The Riemann tensor is antisymmetric in its last two indices, which we know because the commutator changes sign when the two derivatives are exchanged, so $R^{\mu}_{\ \nu\alpha\beta}u^{\alpha}\xi^{\beta}=-R^{\mu}_{\ \nu\beta\alpha}u^{\alpha}\xi^{\beta}$. Relabelling the dummy indices, $\beta\to\rho$ and $\alpha\to\sigma$, this is $-R^{\mu}_{\ \nu\rho\sigma}\xi^{\rho}u^{\sigma}$, and the equation takes its standard form $$\frac{D^{2}\xi^{\mu}}{D\tau^{2}}=-R^{\mu}_{\ \nu\rho\sigma}\,u^{\nu}\xi^{\rho}u^{\sigma},$$ the equation of geodesic deviation, with signs fixed by the curvature convention chosen above.8Sign conventions for the Riemann tensor differ between textbooks, and the overall sign of this equation differs with them. If you compare against another source and the signs disagree, check its definition of $R^{\rho}_{\ \sigma\mu\nu}$ and its metric signature before concluding that one of the two is wrong. Every object in it is a tensor, so the statement holds in every coordinate system: the components of $\xi^{\mu}$ and of the relative acceleration are as observer-dependent as the components of any vector, and yet if the right side is nonzero for one observer it’s nonzero for all of them. That is the invariant we went looking for, and it took this long to state because nothing smaller would do.
Now compare it with the Newtonian tidal equation, and do it carefully, because the correspondence is more exact than a slogan. Fix an observer with four-velocity $u^{\mu}$. Then the geodesic deviation equation defines a map that takes the separation and returns the relative acceleration, $$\xi^{\rho}\;\longmapsto\;-R^{\mu}_{\ \nu\rho\sigma}u^{\nu}u^{\sigma}\xi^{\rho},$$ which is linear in $\xi$ and depends on $u$ but not on any property of the two falling bodies. This is the relativistic tidal operator measured along that worldline. Newtonian gravity has the same structure, $$\xi^{j}\;\longmapsto\;-\partial_{i}\partial_{j}\Phi\,\xi^{j},$$ a linear map from separation to relative acceleration built from second derivatives of the potential. The input and output roles match directly: $\rho$ and $j$ label the direction of separation, while $\mu$ and $i$ label the direction of the resulting relative acceleration. The difference is the pair of velocity contractions, which the relativistic version needs because the answer depends on how the pair is moving through spacetime, and which the Newtonian version could do without because it assumed a universal time in which everyone advances into the future at the same rate. So the right statement isn’t that the Riemann tensor is the tidal matrix. It’s that contracting the Riemann tensor twice with a given four-velocity produces the tidal operator for observers with that velocity, and the Newtonian Hessian is what that operator reduces to when the velocity is small and the field is weak. The dictionary, which we’ll now verify, is $R^{i}_{\ 0j0}\leftrightarrow\partial_{i}\partial_{j}\Phi/c^{2}$. Notice also that a single worldline appears nowhere on the right and couldn’t, since the left side needs two of them to be defined at all, which is the earlier remark about pointlike frames, now as an equation.
The correspondence can be made quantitative, and it’s worth stating the assumptions before starting, since all four are doing work. The field is weak, so $g_{\mu\nu}=\eta_{\mu\nu}+h_{\mu\nu}$ with $|h_{\mu\nu}|\ll1$ and we keep only first order in $h$. The field is static, so $\partial_{0}g_{\mu\nu}=0$. The test particle is slow, so $|dx^{i}/dt|\ll c$ and $d\tau\simeq dt$. And the signature is $(-,+,+,+)$, which fixes the signs throughout. In the geodesic equation the dominant term in $\Gamma^{\mu}_{\ \alpha\beta}\dot{x}^{\alpha}\dot{x}^{\beta}$ is the one with both indices zero, since $\dot{x}^{0}=c\,dt/d\tau\simeq c$ swamps the spatial components, leaving $d^{2}x^{i}/d\tau^{2}\simeq-c^{2}\Gamma^{i}_{\ 00}$. From the Levi-Civita formula, $\Gamma^{i}_{\ 00}=\tfrac12 g^{i\sigma}(2\partial_{0}g_{\sigma0}-\partial_{\sigma}g_{00})$, and staticity kills the first part while $g^{ij}\simeq\delta^{ij}$ to the order we’re working, so $$\Gamma^{i}_{\ 00}\simeq-\tfrac12\delta^{ij}\partial_{j}g_{00},\qquad \frac{d^{2}x^{i}}{dt^{2}}\simeq\frac{c^{2}}{2}\partial_{i}g_{00}.$$ Newton says $d^{2}x^{i}/dt^{2}=-\partial_{i}\Phi$, so the two agree if $\tfrac12 c^{2}\partial_{i}g_{00}=-\partial_{i}\Phi$, which integrates to $$g_{00}\simeq-\left(1+\frac{2\Phi}{c^{2}}\right)$$ with the constant of integration fixed by demanding $g_{00}\to-1$ far from any mass, where the geometry must reduce to Minkowski. Substituting back, $\Gamma^{i}_{\ 00}\simeq+\partial_{i}\Phi/c^{2}$, and $d^{2}x^{i}/dt^{2}\simeq-c^{2}\Gamma^{i}_{\ 00}=-\partial_{i}\Phi$, which is the check that the signs are consistent. The Newtonian potential isn’t a separate field sitting alongside the geometry. In this limit it’s the leading correction to the time-time component of the metric, which is also the honest replacement for the rubber sheet, since what makes an apple fall is a gradient in $g_{00}$ and no picture of a curved spatial surface can display that. And because $g_{00}$ is the component that fixes how much proper time a stationary clock accumulates, the potential turns up in exactly the part of the weak-field metric that controls gravitational time dilation: relative to infinity, the fractional clock-rate shift is $\Phi/c^{2}$. Since $\Phi<0$, a clock deeper in the well runs slower, by about $7\times10^{-10}$ at the Earth's surface, which the GPS system has to correct for.
The tidal correspondence follows from the same expansion. For a static weak field we can drop time derivatives and terms quadratic in $\Gamma$, since each $\Gamma$ is already first order in $h$, so the curvature reduces to $R^{i}_{\ 0j0}\simeq\partial_{j}\Gamma^{i}_{\ 00}$, and with $\Gamma^{i}_{\ 00}\simeq\partial_{i}\Phi/c^{2}$ from above this gives $$R^{i}_{\ 0j0}\simeq\frac{1}{c^{2}}\partial_{i}\partial_{j}\Phi,$$ which is the dictionary entry promised a moment ago. Feed it into the geodesic deviation equation with $u^{\mu}\simeq(c,0,0,0)$, so that only the $\nu=\sigma=0$ terms survive, and $$\frac{d^{2}\xi^{i}}{dt^{2}}\simeq-R^{i}_{\ 0j0}\,c^{2}\,\xi^{j}=-\partial_{i}\partial_{j}\Phi\,\xi^{j},$$ the Newtonian tidal equation we wrote down before any geometry had appeared. The two factors of $c$ from the four-velocity cancel the $1/c^{2}$ in the curvature, reproducing a finite Newtonian tidal acceleration in the nonrelativistic limit. For the Earth, with $\Phi=-GM/r$, this reproduces stretching along the radius at rate $2GM/r^{3}$ and squeezing in each transverse direction at $GM/r^{3}$, and the trace vanishes in vacuum, so a small falling cloud holds its volume to leading order while its shape distorts. The loop is closed: the quantity the Newtonian argument identified as the irremovable part of gravity is a set of components of the Riemann tensor, and the relativistic tidal operator reduces to the Newtonian Hessian in the limit where Newton’s assumptions hold.
So it’s now possible to say what gravity is in this framework, in terms that refer only to things one could measure. Spacetime carries a metric, which fixes proper times, distances, orthogonality and causal structure, and which determines a unique connection that’s compatible with it and free of torsion. That connection defines parallel transport, and free particles are those whose velocity is parallel-transported along their own worldline. At any event a freely falling observer can choose coordinates in which the metric components are Minkowskian and the connection coefficients vanish, so the non-gravitational laws take their special-relativistic form and a free particle shows no coordinate acceleration, and such a frame can be maintained all along that observer’s worldline. What can’t be done is extending it to a finite region, and the obstruction is the Riemann tensor, the particular combination of first and second derivatives of the metric that no coordinate change removes. Its observable signature is the relative acceleration of neighbouring freely falling bodies through the geodesic deviation equation, and in the weak-field limit the relevant components reduce to the Hessian of the Newtonian potential. It would be too strong to say that only the curvature is really there. The metric is the fundamental field of the theory, and it’s what everything else is built from; what the Riemann tensor provides is the coordinate-independent local gravitational content, the part of the description that survives every relabelling and every change of observer.
That’s as far as the equivalence principle takes us, and it’s worth being exact about where it stops. It told us to look for local inertial frames, and the failure of those frames to fit together forced a metric, then a connection, then covariant differentiation, then curvature, each one because the mathematics we had couldn’t express what the physics required. What we have at the end is a complete account of how free motion works in a given geometry, and a way of detecting the invariant curvature of that geometry by watching neighbouring free-fall trajectories. What we don’t have is any reason why spacetime should carry one geometry rather than another. Nothing in the argument so far mentions the Earth. The metric is still an unknown function, and we’ve never explained why a particular distribution of matter produces a particular metric, which means we can recognise gravity but can’t yet predict it.
That’s as far as the equivalence principle takes us, and it’s worth being exact about where it stops. It told us to look for local inertial frames, and the failure of those frames to fit together forced a metric, then a connection, then covariant differentiation, then curvature, each one because the mathematics we had couldn’t express what the physics required. What we have at the end is a complete account of how free motion works in a given geometry, and a way of detecting the invariant curvature of that geometry by watching neighbouring free-fall trajectories. What we don’t have is any reason why spacetime should carry one geometry rather than another. Nothing in the argument so far mentions the Earth. The metric is still an unknown function, and we’ve never explained why a particular distribution of matter produces a particular metric, which means we can recognise gravity but can’t yet predict it. Answering that needs an equation with matter on one side and geometry on the other: the stress-energy tensor $T_{\mu\nu}$ to describe the matter, and on the other side something built from the curvature, which turns out to be constrained rather tightly. Contracting the Riemann tensor gives the Ricci tensor $R_{\mu\nu}=R^{\alpha}{}_{\mu\alpha\nu}$ and contracting again gives the Ricci scalar $R=g^{\mu\nu}R_{\mu\nu}$, and the combination that matters,
$$
G_{\mu\nu}=R_{\mu\nu}-\frac{1}{2}R\,g_{\mu\nu},
$$
is singled out because the Bianchi identity forces $\nabla^{\mu}G_{\mu\nu}=0$ identically, matching the conservation law $\nabla^{\mu}T_{\mu\nu}=0$ that any sensible matter has to satisfy. Setting the two proportional,
$$
G_{\mu\nu}=\frac{8\pi G}{c^{4}}T_{\mu\nu},
$$
is the Einstein field equation, and that’s where this goes next.
References and Footnotes
- 1This is a reconstruction of the logic rather than the history. Einstein worked at the problem for most of a decade with several wrong turns, and the question here is whether we, knowing modern physics, can reconstruct why this structure makes sense. ↩︎
- 2Coordinates of this kind are called Fermi normal coordinates. They put the metric in Minkowski form and set the connection coefficients to zero at every point of the central worldline, with corrections appearing only at second order in the distance away from it. Both of those objects are defined later in this post. ↩︎
- 3Two smaller conventions, used without further comment below. An index appearing once raised and once lowered is summed over, so $A_{\mu}B^{\mu}$ means $\sum_{\mu=0}^{3}A_{\mu}B^{\mu}$. And I keep $c$ explicit throughout rather than setting it to one, so that every formula can be checked against a Newtonian limit. ↩︎
- 4Einstein-Cartan theory is the standard example, where torsion is sourced by intrinsic spin and vanishes outside matter, so the vacuum predictions coincide with general relativity. ↩︎
- 5The steps, for anyone checking: with $L=\tfrac12 g_{\mu\nu}\dot{x}^{\mu}\dot{x}^{\nu}$ and a dot meaning $d/d\lambda$, we have $\partial L/\partial\dot{x}^{\rho}=g_{\rho\nu}\dot{x}^{\nu}$, whose derivative along the curve is $\partial_{\mu}g_{\rho\nu}\dot{x}^{\mu}\dot{x}^{\nu}+g_{\rho\nu}\ddot{x}^{\nu}$, set against $\partial L/\partial x^{\rho}=\tfrac12\partial_{\rho}g_{\mu\nu}\dot{x}^{\mu}\dot{x}^{\nu}$. The first term is contracted with something symmetric in $\mu\nu$ and may be symmetrised, and multiplying through by $g^{\sigma\rho}$ produces the Christoffel combination. ↩︎
- 6Timelike geodesics extremise proper time locally and are genuinely maximal for sufficiently nearby events, but over long paths families of geodesics can refocus at conjugate points, beyond which a geodesic is no longer the longest path between its endpoints. I say extremise rather than maximise throughout for that reason. ↩︎
- 7With torsion the antisymmetric part of $\Gamma^{\lambda}_{\ \mu\nu}$ survives here and contributes an extra term $-T^{\lambda}{}_{\mu\nu}\nabla_{\lambda}V^{\rho}$ to the commutator. Dropping it is a consequence of the choice made earlier, not an algebraic accident. ↩︎
- 8Sign conventions for the Riemann tensor differ between textbooks, and the overall sign of this equation differs with them. If you compare against another source and the signs disagree, check its definition of $R^{\rho}_{\ \sigma\mu\nu}$ and its metric signature before concluding that one of the two is wrong. ↩︎