General relativity has been tested against the solar system, against binary pulsars, and against a black hole’s shadow, matching every one of those measurements. A new catalog of gravitational wave detections, GWTC-5.0, extends that record of tests to 168 independent mergers of black holes and neutron stars, using seven distinct analyses. None of the seven require new mathematics. Each is built from tools already used earlier in this series: inner products, projections, and a Taylor series.
Start with what the detectors actually hand you. A gravitational wave detector does not measure a black hole. It measures a single number as a function of time, the strain, and that number is modeled as a signal buried in noise,
$$ d(t) = h(t) + n(t). $$
The noise is assumed stationary and Gaussian, a strong assumption, but one close enough to true that the whole edifice below stands on it.1Stationary means the noise’s statistical properties do not change over the short duration of a signal; Gaussian means each frequency bin’s noise amplitude follows a normal distribution with a known variance, the power spectral density $S_n(f)$. Neither is exactly true. Real detector noise carries glitches, non-stationary drifts, and calibration uncertainty, all of which the analyses below have to account for separately, and several of the events discussed here needed exactly that kind of special handling. Given this model, comparing a candidate waveform $h$ to the data $d$ is not a vague notion of “does it look right.” It is an inner product, the very same kind of object the second post in this series built from scratch, just fitted with frequency-dependent weights so that noisy frequencies count for less:
$$ \langle a, b \rangle \equiv 4\,\mathrm{Re}!\int_0^\infty \frac{\tilde a^*(f)\,\tilde b(f)}{S_n(f)}\, df. $$
Under this Gaussian noise model, the log-likelihood ratio between “a template of shape $h$, scaled by an unknown amplitude $A$, is present” and “there is only noise” equals $\ln\Lambda(A) = A\langle d,h\rangle – \tfrac12 A^2\langle h,h\rangle$.2Write the noise as $n = d – Ah$ and use $\ln p(n) \propto -\tfrac12\langle n,n\rangle$ for stationary Gaussian noise. Expanding $\langle d-Ah,\,d-Ah\rangle = \langle d,d\rangle – 2A\langle d,h\rangle + A^2\langle h,h\rangle$ and dropping the term with no $A$ in it, since it is the same for every candidate amplitude and cancels in the ratio, leaves exactly this expression. Extremizing over $A$, the same stationarity condition that produced every equation of motion earlier in this series, gives
$$ \frac{d(\ln\Lambda)}{dA} = \langle d,h\rangle – A\langle h,h\rangle = 0 \quad\Longrightarrow\quad A_{\rm ML} = \frac{\langle d,h\rangle}{\langle h,h\rangle}, $$
and substituting back gives the maximized log-likelihood ratio, $\ln\Lambda_{\rm max} = \langle d,h\rangle^2/(2\langle h,h\rangle) \equiv \rho^2/2$, where
$$ \rho = \langle d, \hat h \rangle, \qquad \hat h \equiv \frac{h}{\sqrt{\langle h,h\rangle}}. $$
So the matched-filter signal-to-noise ratio is not a definition dropped in from outside; it is the statistic that falls out of maximizing the detection likelihood over the one free parameter a fixed template shape leaves open, its amplitude, and it is exactly the formula $v_x = e_x \cdot v$ from the vectors post, wearing a frequency integral instead of a sum. Every SNR quoted anywhere in this catalog, and there will be a very large one later, is this same projection.
The first of the seven tests is the most conservative one you could run, asking almost nothing of the theory. Take the best-fit general-relativistic template $h_{\rm maxL}(t)$, the one that maximizes the likelihood of the data, and subtract it out:
$$ r(t) = d(t) – h_{\rm maxL}(t). $$
If general relativity is complete, meaning the template captures every astrophysical feature present in the signal, then whatever is left in $r(t)$ should be indistinguishable from ordinary detector noise. The residual is examined with an algorithm called BayesWave, which reconstructs any coherent, multi-detector power in $r(t)$ using a flexible basis of wavelets rather than any particular waveform model, so it carries no bias toward any specific kind of deviation. This gives a posterior distribution over the residual’s own network SNR, and the analysis reports its 90th percentile, $\mathrm{SNR}_{90}$: a 90% credible statement that whatever signal-like power remains after subtraction is no larger than this number.
A single number on its own does not tell you whether that leftover power is real or just what noise does anyway, so the collaboration builds a null distribution to compare against. They run the identical BayesWave analysis on two hundred stretches of data known to contain no signal, and count how many of those noise-only stretches produce a residual SNR at least as large as the one observed. The fraction is the p-value,
$$
p = P!\left(\mathrm{SNR{90}^{\rm noise} \ge \mathrm{SNR}{90}\right).
$$
A p-value near 1 means ordinary noise regularly produces residuals this loud, so there is nothing here to explain. A p-value near 0 means noise almost never does that, and the leftover power demands an explanation the GR template failed to provide.
Because only 200 background trials were run, the p-value itself is uncertain, and deriving that uncertainty is more informative than simply quoting it. If $p$ is the true, unknown probability that a noise-only trial exceeds the threshold, and $n$ out of $N$ background trials did so, the likelihood of that outcome is the binomial,
$$ \mathcal L(\hat p) = \binom{N}{n} p^n (1-p)^{N-n}. $$
Put a flat prior on $p$ over $[0,1]$, since before running any trials every value is equally plausible. Bayes’ theorem then says the posterior is proportional to the likelihood alone, and the binomial coefficient is just a constant that drops out under normalization:
$$ P(p \mid N, n) \;\propto\; p^n (1-p)^{N-n}. $$
To normalize this, you need $\int_0^1 p^n(1-p)^{N-n}\,dp$, which is the definition of the Beta function, $B(n+1, N-n+1)$. Dividing by it gives the normalized posterior,
$$ \boxed{\; P(p \mid N, n) = \mathrm{Beta}(n+1,\, N-n+1). \;} $$
That is not an approximation borrowed from a statistics textbook. It is what falls straight out of Bayes’ theorem once you notice that a uniform prior is exactly what turns the beta distribution into the correct posterior for a binomial count.3The beta distribution is the conjugate prior for the binomial likelihood precisely because a flat prior is the special case $\mathrm{Beta}(1,1)$, and multiplying $\mathrm{Beta}(a,b)$ by a binomial likelihood always lands you back on another beta distribution, here $\mathrm{Beta}(1+n,\,1+N-n)$. This is the same conjugacy trick used throughout Bayesian statistics to keep posteriors in closed form.
The residual test also reports a fitting factor,
$$
\mathrm{FF}{90} = \frac{\mathrm{SNR}{\rm GR}}{\sqrt{\mathrm{SNR}{\rm GR}^2 + \mathrm{SNR}{90}^2}},
$$
and that formula’s shape follows from a simple geometric picture, not from an arbitrary definition. If the true signal’s total power split cleanly into a piece the GR template captured and an orthogonal leftover piece, the two would combine the way perpendicular legs combine into the hypotenuse of a right triangle, and the denominator above is exactly that hypotenuse. The fitting factor is then the cosine of the angle between the true signal and its GR-template component, the same $u\cdot v = |u||v|\cos\alpha$ from the vectors post. That right-triangle picture rests on the same fact the next test relies on directly: if $h_{\rm maxL}$ truly minimizes $\langle d-h,\,d-h\rangle$ over its template family, the residual is orthogonal to the fitted template to leading order, the same orthogonality a least-squares fit always produces. This Pythagorean picture is a simplification: $\mathrm{SNR}{\rm GR}$ and $\mathrm{SNR}{90}$ are both point estimates from different analyses rather than exactly orthogonal legs of one triangle, and the paper itself notes two events, both with unusually low SNR, where $\mathrm{SNR}{90}$ comes out larger than $\mathrm{SNR}{\rm GR}$, which a strict right triangle could never allow. The geometric picture is the right intuition; it is not an exact theorem here.
Across all 168 events the residuals came back consistent with noise, and the p-values were consistent with the flat distribution you would expect if nothing were being missed. That is the least demanding of the seven tests, and the right one to run first: before asking whether gravity generates waves correctly or whether black holes ring the way Kerr says they should, you check that nothing coherent was left on the table at all.
The second test asks a much sharper geometric question. General relativity predicts gravitational waves come in exactly two independent polarizations, usually called plus and cross. Many alternative theories of gravity predict more: a scalar breathing mode, vector modes, and so on. The question is how you test for extra polarizations without ever assuming what their waveform looks like in time, since a theory-specific search would only catch the one alternative theory you happened to guess.
Here is the setup. With a network of $N$ detectors, collect the strain data into a vector $\tilde{\mathbf d}(f) \in \mathbb C^N$, one entry per detector. This is related to the polarization content, an $M$-dimensional vector $\tilde{\mathbf h}(f) \in \mathbb C^M$ holding whichever modes you have decided to test for, by
$$ \tilde{\mathbf d}(f) = \mathbf F(\alpha,\delta,\psi,t_{\rm event})\, \tilde{\mathbf h}(f) + \tilde{\mathbf n}(f). $$
$\mathbf F$ is the beam pattern matrix, an $N\times M$ array whose entries depend only on the sky location, polarization angle, and detector geometry, never on the source’s astrophysics or the waveform’s time dependence. Its columns are $M$ vectors living in the $N$-dimensional space of detector outputs; if they are linearly independent, which they generically are, they span an $M$-dimensional subspace of that space, the same subspace idea from two posts ago, now doing physical work. Whatever gravitational wave signal actually arrives, however it was generated, however its waveform evolves in time, its imprint on the detector network is forced by geometry alone to lie inside $\mathrm{span}(\mathbf F)$. This is the whole test in one sentence: geometry, not waveform physics, tells you where the signal must live.
So project the data onto the orthogonal complement of that subspace and see what is left. To find that projector, look first for the point inside $\mathrm{span}(\mathbf F)$ closest to the data: the $\hat{\mathbf h}$ minimizing the squared distance $(\tilde{\mathbf d}-\mathbf F\hat{\mathbf h})^\dagger(\tilde{\mathbf d}-\mathbf F\hat{\mathbf h})$. Treating $\hat{\mathbf h}$ and $\hat{\mathbf h}^\dagger$ as independent variables, the standard trick for extremizing a real quantity built from complex vectors, and setting the derivative with respect to $\hat{\mathbf h}^\dagger$ to zero gives the normal equations,
$$ \mathbf F^\dagger!\left(\tilde{\mathbf d} – \mathbf F\hat{\mathbf h}\right) = \mathbf 0 \quad\Longrightarrow\quad \hat{\mathbf h} = (\mathbf F^\dagger\mathbf F)^{-1}\mathbf F^\dagger\tilde{\mathbf d}. $$
The closest point in the subspace is therefore $\mathbf F\hat{\mathbf h} = \mathbf F(\mathbf F^\dagger\mathbf F)^{-1}\mathbf F^\dagger\,\tilde{\mathbf d}$, which means $\mathbf F(\mathbf F^\dagger\mathbf F)^{-1}\mathbf F^\dagger$ is itself the projector onto $\mathrm{span}(\mathbf F)$; subtracting it from the identity gives the projector onto the orthogonal complement,
$$ \mathbf P = \mathbf I_N – \mathbf F\,(\mathbf F^\dagger \mathbf F)^{-1}\,\mathbf F^\dagger. $$
That formula is not asserted; it is the unique minimizer of a squared distance, produced by the same move, extremizing a quantity and solving whatever equation falls out, that has generated every result in this series so far. It is still necessary to check, rather than simply trust, that the object this minimization produced actually behaves like a projector. It should satisfy $\mathbf P^2 = \mathbf P$, since projecting twice ought to do nothing the first projection did not already do:
$$ \mathbf P^2 = \mathbf I – 2\mathbf F(\mathbf F^\dagger\mathbf F)^{-1}\mathbf F^\dagger + \mathbf F(\mathbf F^\dagger\mathbf F)^{-1}\underbrace{\mathbf F^\dagger\mathbf F(\mathbf F^\dagger\mathbf F)^{-1}}_{=\,\mathbf I_M}\mathbf F^\dagger = \mathbf I – \mathbf F(\mathbf F^\dagger\mathbf F)^{-1}\mathbf F^\dagger = \mathbf P. $$
And it should annihilate anything already inside $\mathrm{span}(\mathbf F)$:
$$ \mathbf P \mathbf F = \mathbf F – \mathbf F(\mathbf F^\dagger\mathbf F)^{-1}\mathbf F^\dagger \mathbf F = \mathbf F – \mathbf F = \mathbf 0. $$
Apply $\mathbf P$ to the actual data. If the true signal really does have exactly the polarization content you assumed, $\tilde{\mathbf h}$ lying entirely within the $M$ modes of $\mathbf F$, then
$$ \mathbf P\,\tilde{\mathbf d} = \mathbf P\mathbf F\tilde{\mathbf h} + \mathbf P\tilde{\mathbf n} = \mathbf 0 + \mathbf P\tilde{\mathbf n} = \mathbf P\tilde{\mathbf n}, $$
pure noise, regardless of what $\tilde{\mathbf h}(f)$ looked like as a function of frequency. This quantity is called the null stream, and if the real polarization content includes anything outside $\mathrm{span}(\mathbf F)$, an extra scalar mode, say, the projection $\mathbf P\tilde{\mathbf d}$ will retain a coherent signal component, because that piece of the true signal was never inside the subspace $\mathbf P$ was built to erase.
A practical constraint falls straight out of the linear algebra: the network needs more detectors than assumed polarization modes, $N > M$, or there is no orthogonal complement left to inspect. This is exactly why a two-detector observation can only probe a single effective polarization mode, and why several of the events in this catalog, observed by only two of the three instruments, get a correspondingly weaker version of this test. The paper turns the null stream comparison into a Bayesian one, computing evidence for competing polarization hypotheses, tensor only, scalar only, various mixtures, by sampling how well each geometry accounts for the data. Combining many events is then just multiplication of independent evidences, or equivalently, addition of their logarithms,4If events are statistically independent, their joint likelihood is the product of individual likelihoods, so the joint Bayesian evidence for a hypothesis is the product of per-event evidences, and the combined Bayes factor between two hypotheses is the product of per-event Bayes factors. Taking $\log_{10}$ turns that product into the sum used throughout the paper, $\log_{10}B^X_T\big|_{\rm combined} = \sum_i \log_{10}B^X_{T,i}$. the same additivity that will reappear when the ringdown results get combined later. Across the full catalog, purely scalar and purely vector polarizations are ruled out overwhelmingly, by roughly nineteen and six orders of magnitude against the tensor hypothesis, respectively. Mixed tensor-plus-extra hypotheses came back close to neutral, with one, tensor-vector, landing mildly on the non-GR side of zero; the collaboration traced this to noise systematics and an approximation used for one particular high-spin, unequal-mass event rather than to any genuine departure, and rated the result inconclusive rather than a discovery. Non-Gaussian noise can produce excess power that this projection cannot distinguish from a real signal, which is the stated reason for the caution.
The third and fourth tests move from geometry to dynamics: not which polarizations arrive, but whether the waves are generated the way general relativity’s field equations say they should be. During the long inspiral phase, while the two orbiting bodies are still much farther apart than their own size, their motion can be treated perturbatively in the small parameter $v/c$, the orbital speed over the speed of light, an approach called the post-Newtonian expansion. The gravitational wave phase accumulated by the time the signal reaches a given frequency takes the form
$$ \Psi(f) = 2\pi f t_c – \phi_c – \frac{\pi}{4} + \frac{3}{128\,\eta\, v^5}\sum_{n=0}^{7}\left(\psi_n + \psi_{nl}\log v\right) v^n, $$
where $\eta = m_1 m_2/(m_1+m_2)^2$ is the symmetric mass ratio,5The symmetric mass ratio is bounded, $\eta \le 1/4$, with equality only for equal masses. This falls out of a single line of algebra: $(m_1-m_2)^2 \ge 0$ expands to $m_1^2+m_2^2 \ge 2m_1m_2$, add $2m_1m_2$ to both sides to get $(m_1+m_2)^2 \ge 4m_1m_2$, and divide through. So $\eta$ is a compact way of encoding how lopsided a binary is, small $\eta$ meaning very unequal masses, which is exactly what makes an asymmetric system like GW241011 valuable: it explores a part of parameter space most binaries do not reach. and $v$ deserves an actual derivation rather than a bare proportionality. At leading, Newtonian order, a circular orbit of total mass $M$ and separation $r$ obeys Kepler’s third law, $\omega_{\rm orb}^2 = GM/r^3$, which gives the orbital speed $v = \omega_{\rm orb}\,r = \omega_{\rm orb}(GM/\omega_{\rm orb}^2)^{1/3} = (GM\,\omega_{\rm orb})^{1/3}$. The dominant radiation is quadrupolar, emitted at twice the orbital frequency, so the observed gravitational wave frequency is $f = \omega_{\rm orb}/\pi$, or $\omega_{\rm orb}=\pi f$. Substituting,
$$ v = (GM\pi f)^{1/3}, \qquad \text{or, restoring the speed of light,} \qquad \frac{v}{c} = \left(\frac{\pi GMf}{c^3}\right)^{1/3}, $$
which is precisely the small parameter appearing in $\Psi(f)$ above: post-Newtonian theory is built by starting from this Newtonian relation and adding relativistic corrections order by order in $v/c$. The term $(n/2)$PN by convention labels a correction of relative order $(v/c)^n$ beyond this leading term, and for fixed masses and spins, general relativity fixes every coefficient $\psi_n$ uniquely. There is no freedom left to fit.
That absence of freedom is exactly what the test exploits. TIGER and the Flexible Theory-Independent method each let one coefficient at a time float by a controlled fractional amount, $\psi_i \to (1+\delta\hat\varphi_i)\psi_i^{\rm GR,NS} + \psi_i^{\rm GR,S}$, and ask whether the data prefer $\delta\hat\varphi_i = 0$. Since a single binary’s inspiral covers many observed cycles while depending on only a handful of physical numbers, namely mass, mass ratio, and spins, this is an overdetermined system: measuring the phase evolution in detail checks far more than it takes to specify the source, and each PN order is an independent opportunity for a discrepancy to appear. One coefficient is missing from GR’s list entirely, for a specific reason. The leading term at $-1$PN would correspond to dipole radiation, and general relativity forbids it outright, for the same reason electromagnetic dipole radiation has no counterpart in gravity: dipole radiation requires an accelerating dipole moment, and a system’s mass dipole moment is just its center of mass times its total mass, whose second time derivative vanishes by conservation of momentum.6This is the gravitational-wave cousin of the fact that isolated systems cannot radiate monopole waves either, since total mass-energy is conserved and a constant monopole moment cannot accelerate. The leading radiating multipole in GR is therefore the mass quadrupole. Alternative theories that endow compact objects with an extra gravitational scalar charge, unequal on the two bodies, generically reintroduce a nonzero dipole moment and predict exactly this $-1$PN term, which is why testing for it is such a clean discriminator and why the best bound on it in this catalog still comes from the long, close binary neutron star inspiral GW170817 rather than any of the new black hole mergers. Across the whole catalog, every PN deformation parameter came back consistent with zero, with the bounds tightened by factors of 1.2 to 2.6 over the previous catalog. The collaboration also mapped these agnostic bounds onto a handful of concrete modified-gravity theories, scalar-tensor gravity, Einstein-dilaton-Gauss-Bonnet gravity, dynamical Chern-Simons gravity, each of which predicts a specific pattern of PN deviations. The mappings are carefully labeled as illustrative rather than as genuine theory-specific constraints, since a real modified theory would generically shift several PN coefficients at once and also touch the merger and ringdown, neither of which this particular mapping accounts for.
The last three tests turn to what happens after the two objects merge, and this part of the analysis has a direct analogue in atomic spectroscopy, made precise below. Immediately after merger, the newly formed black hole is not yet a quiet, stationary object. It is a perturbed one, ringing, and it sheds that perturbation by radiating a superposition of discrete, exponentially damped sinusoids called quasinormal modes,7Each term combines into a single complex frequency, $\bar\omega_{\ell m n} \equiv 2\pi f_{\ell m n} + i/\tau_{\ell m n}$, so that $e^{-t/\tau_{\ell m n}}e^{2\pi i f_{\ell m n}t} = e^{i\bar\omega_{\ell m n}t}$ under this sign convention. The frequencies come out complex, rather than the purely real spectrum you would get from, say, a vibrating string with fixed ends, because the linear perturbation equations governing the remnant are solved subject to purely outgoing radiation at infinity and purely ingoing radiation at the horizon. Those boundary conditions make the eigenvalue problem non-self-adjoint, and a non-self-adjoint eigenvalue problem is not obligated to return real eigenvalues; the imaginary part of each one is exactly that mode’s decay rate.
$$ h(t) = \sum_{\ell,m,n} A_{\ell m n}\, S_{\ell m n}(\iota,\varphi,\chi_f)\, e^{-t/\tau_{\ell m n}}\, e^{2\pi i f_{\ell m n} t}. $$
The whole test rests on one remarkable fact. A theorem in general relativity, built from work by Israel, Carter, Hawking, and Robinson through the 1960s and 70s, says that a stationary, axisymmetric, vacuum black hole with a well-behaved horizon must be a member of the Kerr family, fixed by exactly two numbers: its mass $M_f$ and its dimensionless spin $\chi_f$.8This is often loosely called the no-hair theorem; more precisely it is a uniqueness theorem for stationary electrovacuum solutions, which in the absence of charge, expected for astrophysical black holes since any net charge is neutralized almost instantly by surrounding plasma, leaves mass and spin as the only two parameters. No third number is available for the remnant to encode any extra information in. Since the entire quasinormal mode spectrum, every $f_{\ell m n}$ and $\tau_{\ell m n}$ for every choice of indices, is a fixed mathematical function of just those two numbers, measuring one mode gives you $M_f$ and $\chi_f$, and measuring any additional mode is then a zero-parameter prediction: general relativity has already spent its only two adjustable numbers and has nothing left to fit the second mode with. This is black hole spectroscopy, and it is the direct gravitational analogue of atomic spectroscopy: a hydrogen atom’s entire spectral line series is fixed by a single number, its nuclear charge, through the Rydberg formula, and observing a second spectral line beyond the one used to calibrate that charge is a genuine, over-constrained test of the same underlying quantum theory. Here the “nuclear charge” is replaced by a pair of numbers, and the “spectral lines” are damped sinusoids instead of photon energies, but the logic of the test is identical.
The deviation parametrization is built to stay physical no matter what the data prefer, writing $f_{\ell m n} = f^{\rm Kerr}{\ell m n}\,e^{\delta f{\ell m n}}$ and $\gamma_{\ell m n} = \gamma^{\rm Kerr}{\ell m n}\,e^{\delta\gamma{\ell m n}}$ for the damping rate $\gamma=1/\tau$, an exponential form chosen so that no matter how large a deviation the data prefer, the reconstructed frequency and damping rate stay positive, a small piece of engineering that keeps a wildly off posterior from returning nonsense. For one event in this catalog, GW240621, the collaboration found genuine, if not fully secure, evidence for a second mode, the first overtone of the dominant quadrupole, appearing in the earliest post-merger data. Scanning across different analysis start times showed the dominant mode’s inferred amplitude departing from its expected exponential decay at early times, in a way that a second damped sinusoid resolved cleanly; a self-consistent measurement of the remnant’s mass and spin only emerged once that overtone was included. But when the collaboration simulated a pure general-relativity signal with the same parameters and injected it into similar detector noise, the simulations could not reliably reproduce the same tightly constrained, high-frequency feature seen in the real data, so the result is reported as suggestive rather than a confirmed detection of Kerr’s overtone.
One event in this catalog has a network matched-filter signal-to-noise ratio far above the rest: GW250114 was observed with an SNR of 76.9. Because a matched-filter SNR is exactly the projection $\rho = \langle d,\hat h\rangle$ from the first paragraph of this post, a number that large means the data align with the general-relativistic template about as cleanly as noise will ever allow. Its ringdown alone constrains the dominant mode’s fractional frequency and damping-time deviations to $\delta\hat f_{220} = 0.02 \pm 0.02$ and $\delta\hat\tau_{220} = -0.01^{+0.10}_{-0.09}$, the tightest single-event ringdown bound in the whole catalog, and it dominates the combined result enough that removing it from the hierarchical combination visibly loosens the joint bound on the damping time. Combined across all qualifying events, the ringdown analyses improve the dominant mode’s frequency and damping-time constraints by factors of 1.95 and 1.48 over the previous catalog. The combined posterior places the general-relativity prediction at the edge, though not outside, the 90% credible region, a mild tension the collaboration attributes to a combination of finite catalog size, correlated systematics, and possibly unmodeled selection effects rather than to physics beyond general relativity. That finite-catalog uncertainty is quantified by bootstrapping over the event set rather than by reporting a single sharp number.
The seven tests probe different aspects of the same theory. The residual test checks that no coherent signal is left unaccounted for after subtracting the best-fit template. The polarization test checks that the data are consistent with exactly the two tensor polarizations general relativity predicts, using a projection onto a subspace. The generation tests check that the inspiral phase matches the post-Newtonian coefficients general relativity predicts, order by order, including the absence of a dipole term. The ringdown tests check that the postmerger spectrum matches the two-parameter Kerr prediction. Each of these tests uses the inner product, projection, and expansion tools built earlier in this series.
Across all seven tests and all 168 events, the data are statistically consistent with general relativity: no residual test found unexplained coherent power, no polarization test found significant evidence for extra polarizations, no generation test found a nonzero PN deformation at the 90% credible level, and the ringdown tests found the Kerr prediction within, or at the edge of, the combined 90% credible region. This is a statement about consistency with a specific, finite set of measurements, not a proof of the theory. Each additional independent event tightens the same set of bounds without closing them entirely.
References and Footnotes
- 1Stationary means the noise’s statistical properties do not change over the short duration of a signal; Gaussian means each frequency bin’s noise amplitude follows a normal distribution with a known variance, the power spectral density $S_n(f)$. Neither is exactly true. Real detector noise carries glitches, non-stationary drifts, and calibration uncertainty, all of which the analyses below have to account for separately, and several of the events discussed here needed exactly that kind of special handling.
- 2Write the noise as $n = d – Ah$ and use $\ln p(n) \propto -\tfrac12\langle n,n\rangle$ for stationary Gaussian noise. Expanding $\langle d-Ah,\,d-Ah\rangle = \langle d,d\rangle – 2A\langle d,h\rangle + A^2\langle h,h\rangle$ and dropping the term with no $A$ in it, since it is the same for every candidate amplitude and cancels in the ratio, leaves exactly this expression.
- 3The beta distribution is the conjugate prior for the binomial likelihood precisely because a flat prior is the special case $\mathrm{Beta}(1,1)$, and multiplying $\mathrm{Beta}(a,b)$ by a binomial likelihood always lands you back on another beta distribution, here $\mathrm{Beta}(1+n,\,1+N-n)$. This is the same conjugacy trick used throughout Bayesian statistics to keep posteriors in closed form.
- 4If events are statistically independent, their joint likelihood is the product of individual likelihoods, so the joint Bayesian evidence for a hypothesis is the product of per-event evidences, and the combined Bayes factor between two hypotheses is the product of per-event Bayes factors. Taking $\log_{10}$ turns that product into the sum used throughout the paper, $\log_{10}B^X_T\big|_{\rm combined} = \sum_i \log_{10}B^X_{T,i}$.
- 5The symmetric mass ratio is bounded, $\eta \le 1/4$, with equality only for equal masses. This falls out of a single line of algebra: $(m_1-m_2)^2 \ge 0$ expands to $m_1^2+m_2^2 \ge 2m_1m_2$, add $2m_1m_2$ to both sides to get $(m_1+m_2)^2 \ge 4m_1m_2$, and divide through. So $\eta$ is a compact way of encoding how lopsided a binary is, small $\eta$ meaning very unequal masses, which is exactly what makes an asymmetric system like GW241011 valuable: it explores a part of parameter space most binaries do not reach.
- 6This is the gravitational-wave cousin of the fact that isolated systems cannot radiate monopole waves either, since total mass-energy is conserved and a constant monopole moment cannot accelerate. The leading radiating multipole in GR is therefore the mass quadrupole. Alternative theories that endow compact objects with an extra gravitational scalar charge, unequal on the two bodies, generically reintroduce a nonzero dipole moment and predict exactly this $-1$PN term, which is why testing for it is such a clean discriminator and why the best bound on it in this catalog still comes from the long, close binary neutron star inspiral GW170817 rather than any of the new black hole mergers.
- 7Each term combines into a single complex frequency, $\bar\omega_{\ell m n} \equiv 2\pi f_{\ell m n} + i/\tau_{\ell m n}$, so that $e^{-t/\tau_{\ell m n}}e^{2\pi i f_{\ell m n}t} = e^{i\bar\omega_{\ell m n}t}$ under this sign convention. The frequencies come out complex, rather than the purely real spectrum you would get from, say, a vibrating string with fixed ends, because the linear perturbation equations governing the remnant are solved subject to purely outgoing radiation at infinity and purely ingoing radiation at the horizon. Those boundary conditions make the eigenvalue problem non-self-adjoint, and a non-self-adjoint eigenvalue problem is not obligated to return real eigenvalues; the imaginary part of each one is exactly that mode’s decay rate.
- 8This is often loosely called the no-hair theorem; more precisely it is a uniqueness theorem for stationary electrovacuum solutions, which in the absence of charge, expected for astrophysical black holes since any net charge is neutralized almost instantly by surrounding plasma, leaves mass and spin as the only two parameters. No third number is available for the remnant to encode any extra information in.








