Planet Musings

August 13, 2026

Terence TaoA digestion of the proof of Sendov’s conjecture

This post concerns the following conjecture of Sendov, as well as its strengthening by Phelps–Rodriguez:

Conjecture 1 (Sendov’s conjecture) Let {n \geq 2}, and let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. Then for every zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|\zeta-a| \leq 1}.

Conjecture 2 (Phelps–Rodriguez conjecture) Let {n \geq 2}, and let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. Then for every zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|\zeta-a| < 1}, unless {a} is on the unit circle and {p} is a scalar multiple of {z^n - a^n}.

By applying a rotation around the origin, we can normalize {a} to be a real number with {0 \leq a \leq 1}.

From the work of Rubinstein, both conjectures were already established in the {a=1} case, so one can restrict to the {0 \leq a < 1} case. Both of these conjectures then follow from

Conjecture 3 (Sendov’s conjecture in interior) Let {n \geq 2}. Let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. Then if {0 \leq a < 1} is a zero of {p}, there exists a critical point {\zeta} of {p} with {|\zeta-a| < 1}.

All three of these conjectures were established for {n \leq 8} (in a sequence of papers culminating in this paper of Brown and Xiang) and for sufficiently large {n} (in a paper of myself, which in turn built upon several partial results in this setting). This left the case of intermediate {n} to be settled. My arguments used some qualitative ingredients (most notably analytic continuation) and as such did not easily lend themselves to quantifying the threshold of {n} above which the argument was valid.

Recently, Lech Mazur was able to use an AI tool to resolve Sendov’s conjecture for all {n \geq 2}, with the proof verified in Lean. However, the AI-generated proof was not human-digested to be in the form of a publication-ready preprint; and it has taken me several days (with heavy AI assistance) to perform such a digestion, to place the proof in proper context with previous literature and to simplify and streamline the argument to highlight the main ideas. (Note: the above chat log only represents a portion of the digestion work: the rest was performed with pen and paper, or using some further AI agents.) The same arguments also give a new proof of Rubinstein’s theorem, which I also give below the fold.

One consequence of this digestion is that the argument in fact demonstrates Conjecture 3, and thus resolves both the Sendov conjecture and the Phelps–Rodriguez conjecture in full generality.

The proof ends up being remarkably elementary. No complex analysis is used other than the fundamental theorem of algebra (and very basic facts about Möbius transformations); and the deepest inequality used as input is the Maclaurin inequality (and we only need a special case of that inequality which can be derived from the arithmetic mean-harmonic mean inequality and an induction argument).

Using an AI agent, I have been able to formalize the entire argument in Lean, extended to {n \geq 2} by some minor modifications to the proof. This formalization is more streamlined than the original formalization (it has about 15,000 lines of code, compared with around 90,000 for the original proof).

We now prove Conjecture 3. The {n \leq 4} cases have long been known but need to be treated separately; a short proof using the machinery developed here is provided at the end of the post. Suppose now that we have a counterexample for some {n \geq 5}, thus one can find a degree {n} polynomial {p} with zeroes

\displaystyle  a, z_1, \dots,z_{n-1}

for some {0 \leq a < 1} and {z_j}, {j=1,\dots,n-1} in the closed unit disk, whose critical points all lie a distance at least {1} from {a}. We use {O()} notation here in the non-asymptotic sense, thus {X = O(Y)} means that {|X| \leq C Y} for some absolute constant {C} (independent of {n}). We will also use the notation {O_{\leq}(Y)} to denote a quantity that is bounded in magnitude by {Y}.

To capture the fact that the critical points lie at a distance at least {1} from {a}, we write these critical points as

\displaystyle  a - \frac{1}{q_1}, \dots, a - \frac{1}{q_{n-1}}

for some (non-zero) {q_1,\dots,q_{n-1}} in the closed unit disk.

Example 4 If {p(z) = z^n - 1} and {a=1}, then {z_1,\dots,z_{n-1}} are the non-trivial {n^{th}} roots of unity, while the {q_1,\dots,q_{n-1}} are all equal to {1}. Strictly speaking this is not actually a counterexample to Conjecture 3, because {a} is not strictly less than one; nevertheless this is an important motivating near-counterexample for the arguments below.

Example 5 A generalization of the previous example was studied in Section 4 of my paper. Here one took

\displaystyle  p(z) = (z + \frac{c_2}{n})^{n-m} P(z) - (a + \frac{c_2}{n})^{n-m} P(a)

where {n} was an asymptotic parameter going to infinity,

\displaystyle  P(z) = (z-\lambda_1) \dots (z - \lambda_m)

was a low-degree polynomial for some {m=O(1)},

\displaystyle  a = 1 - \frac{c_1}{n},

and {c_1,c_2 > 0} were constants. This polynomial has a zero at {a}, {n-m-1} critical points at {-\frac{c_2}{n}}, and {m} additional critical points near {\lambda_1,\dots,\lambda_m}. If all the critical points were at distance at least one from {a}, one would have

\displaystyle  c_2 \geq c_1

and

\displaystyle  |1-\lambda_j| \geq 1 - o(1)

while if all the zeroes were in the unit disk, the calculations in my paper showed that

\displaystyle  c_2 - c_1 - c_2 \cos \theta + \sum_{j=1}^m \log |\frac{1-\lambda_j}{e^{i\theta}-\lambda_j}| \leq o(1) \ \ \ \ \ (1)

Here {o(1)} denotes a quantity that goes to zero as {n \rightarrow \infty}. If one ignores the {o(1)} errors, one can show that these conditions are only simultaneously feasible if {c_1=c_2} and all the {\lambda_j} vanish, but the argument was somewhat subtle (I had to proceed by inspecting the second Fourier coefficient of (1)). This illustrates the fact that the regime {a = 1 - O(1/n)} is particularly delicate.

We now have two sets of points in the closed unit disk: {z_1,\dots,z_{n-1}} and {q_1,\dots,q_{n-1}}. They “communicate” with each other through the polynomial {p} and its first derivative {p'}, both of which can be expressed in terms of either set of points (as well as {n} and {a}). Indeed, if we normalize {p} to be monic, then we can factor {p} in terms of the zeroes as

\displaystyle  p(z) = (z-a) \prod_{j=1}^{n-1} (z-z_j) \ \ \ \ \ (2)

and thus upon differentiating

\displaystyle  p'(z) = \left(\prod_{j=1}^{n-1} (z-z_j)\right) \left(1 + (z-a) \sum_{j=1}^{n-1} \frac{1}{z-z_j}\right). \ \ \ \ \ (3)

Here and in the sequel we adopt the convention of removing singularities when dealing with expressions that involve multiplication by both {\frac{1}{z-z_j}} and {z-z_j}, by cancelling such terms first in the event that {z=z_j}.

In a similar vein, {p'} can be factored

\displaystyle  p'(z) = n \prod_{j=1}^{n-1} \left(z - a + \frac{1}{q_j}\right) \ \ \ \ \ (4)

and thus on integrating (and using {p(a)=0})

\displaystyle  p(z) = (z-a) \int_0^1 n \prod_{j=1}^{n-1} \left(t(z - a) + \frac{1}{q_j}\right)\ dt. \ \ \ \ \ (5)

It is convenient to rule out the easy case {a=0} right away. In this case we see from (3), (4) that

\displaystyle  p'(0) = \prod_{j=1}^{n-1} (-z_j) = n \prod_{j=1}^{n-1} \frac{1}{q_j}

which is absurd since the first product has magnitude at most one, and the second product has magnitude at least one. Thus we can assume henceforth that {a>0}.

By inspecting {p} or {p'} at various natural locations, we can thus obtain a number of identities relating the {z_j} to the {q_j}. We record the ones that we actually need here:

Lemma 6 Let {F} denote the function

\displaystyle  F(t) := \prod_{j=1}^{n-1} (1 - atq_j). \ \ \ \ \ (6)

  • (i) (Centroid identity) We have

    \displaystyle  \frac{1}{n} \left(a + \sum_{j=1}^{n-1} z_j\right) = \frac{1}{n-1} \sum_{j=1}^{n-1} \left(a - \frac{1}{q_j}\right). \ \ \ \ \ (7)

    That is to say, the centroid of the zeroes equals the centroid of the critical values.
  • (ii) (Polar identity) We have

    \displaystyle  \prod_{j=1}^{n-1} \frac{1-az_j}{a-z_j} = \int_0^1 \prod_{j=1}^{n-1} (t (1-a^2) q_j + a)\ dt. \ \ \ \ \ (8)

  • (iii) (First origin identity) We have

    \displaystyle  (-1)^{n-1} \prod_{j=1}^{n-1} z_j = \frac{n}{\prod_{j=1}^{n-1} q_j} \int_0^1 F(t)\ dt. \ \ \ \ \ (9)

  • (iv) (Second origin identity) We have

    \displaystyle  (-1)^{n-1} \prod_{j=1}^{n-1} z_j \left( 1 + a \sum_{j=1}^{n-1} \frac{1}{z_j} \right) = \frac{n}{\prod_{j=1}^{n-1} q_j} F(1). \ \ \ \ \ (10)

    (Again, we are using the convention of removing singularities to deal with the case where some of the {z_j} vanish.)

Proof: For (i), we inspect the behavior of {p'(z)} as {z \rightarrow \infty}. From (2) we have

\displaystyle  p(z) = z^n - \left( a + \sum_{j=1}^{n-1} z_j \right) z^{n-1} + O(z^{n-2})

and thus on differentiating term by term

\displaystyle  p'(z) = n z^{n-1} - (n-1) \left( a + \sum_{j=1}^{n-1} z_j \right) z^{n-2} + O(z^{n-3}).

Meanwhile, from (4) we have

\displaystyle  p'(z) = n z^{n-1} - n \left(\sum_{j=1}^{n-1} \left(a - \frac{1}{q_j}\right)\right) z^{n-2} + O(z^{n-3}).

Comparing coefficients, we obtain the claim.

For (ii), we consider the expression {p(1/a) / p'(a)}. On the one hand, from (2), (3) one has

\displaystyle  \frac{p(1/a)}{p'(a)} = \frac{(1/a - a) \prod_{j=1}^{n-1} (1/a - z_j)}{\prod_{j=1}^{n-1} (a-z_j)}.

(Note from hypothesis that {a} cannot be a critical point, so the denominator is non-zero.) On the other hand, from (4), (5) one has

\displaystyle  \frac{p(1/a)}{p'(a)} = \frac{(1/a - a) \int_0^1 n \prod_{j=1}^{n-1} \left(t\left(\frac{1}{a} - a\right) + \frac{1}{q_j}\right)\ dt}{n \prod_{j=1}^{n-1} \frac{1}{q_j}}.

Equating the two identities, we obtain (ii) after some algebra.

For (iii), we evaluate {p(0)}. From (2) we have

\displaystyle  p(0) = -a (-1)^{n-1} \prod_{j=1}^{n-1} z_j

while from (5) we have

\displaystyle  p(0) = -a \int_0^1 n \prod_{j=1}^{n-1} \left(-at + \frac{1}{q_j}\right)\ dt.

Equating the two identities, we obtain (iii) after some algebra using (6).

For (iv), we similarly evaluate {p'(0)}. From (3) we have

\displaystyle  p'(0) = (-1)^{n-1} \left(\prod_{j=1}^{n-1} z_j\right) \left(1 + a \sum_{j=1}^{n-1} \frac{1}{z_j}\right)

while from (4) one has

\displaystyle  p'(0) = n \prod_{j=1}^{n-1} \left(- a + \frac{1}{q_j}\right).

Equating the two identities, we obtain (iv) after some algebra using (6). \Box

Remarkably, the polynomial {p} will play no further role in the argument: the identities in (i)-(iv), together with the hypotheses that {0 \leq a < 1} and {z_j, q_j} lie in the closed unit disk, will be sufficient by themselves to obtain a contradiction.

Example 7 Continuing the example in Example 4, in (i) both sides vanish. In (ii), both sides are equal to one. For (iii) and (iv), we have {F(t) = (1-t)^{n-1}}, with both sides of (iii) equal to one, and both sides of (iv) equal to zero.

Remark 8 The centroid identity is extremely classical, going back to this 1948 paper of Popoviciu. The comparison of the polynomial at a location {a} and at the polar inversion {1/a} of that location across the closed unit disk is a familiar trick in the literature; see, e.g., Lemma 5 and Theorem 8 of Dégot. The specific form of the polar identity is implicit in the first part of Section 5 of Mazur’s AI-generated proof, while the origin identities are extracted from equation (6.3) of that proof. The first origin identity is also very close to Theorem 6 of Dégot, while the second origin identity is similar to some identities appearing in the proof of Lemma 6 of Dégot, as well as the work of Mir–Nazir–Wani and (in the {a=1} case) Rubinstein. The work of Meir–Sharma and Mir–Nazir–Wani also contain several further identities relating the {z_j} to the {q_j}; see in particular Lemma 15 below. Variants of (5) also appear in Proposition 10 of Miller.

Remark 9 The first origin identity (9) is already strong enough to handle asymptotically all examples of the form in Example 5, except in the endpoint case where {c_1, c_2} vanish and the {\lambda_j} are all {0}. Indeed, as the {z_j, q_j} are in the closed unit disk, (9) implies that

\displaystyle  n \left|\int_0^1 F(t)\ dt\right| \leq 1.

On the other hand, routine calculations (omitted here) show that

\displaystyle  n \int_0^1 F(t)\ dt = 1 + \frac{c_2 + \sum_{j=1}^{m} \frac{-\lambda_j}{1-\lambda_j}}{n} + O\left(\frac{1}{n^2}\right)

leading asymptotically to the constraint

\displaystyle  c_2 + \mathrm{Re} \sum_{j=1}^{m} \frac{-\lambda_j}{1-\lambda_j} \leq 0.

But all terms here are non-negative (since {|1-\lambda_j| \geq 1}), so this forces a contradiction unless {c_2} (and hence also {c_1}) and the {\lambda_j} all vanish.

As mentioned in Example 5, the most delicate regime occurs when {a = 1 - O(1/n)}. It is convenient to introduce the normalized version

\displaystyle  \alpha := \frac{n-1}{2} (1-a^2), \ \ \ \ \ (11)

of {a}, thus {0 < \alpha < \frac{n-1}{2}}, and the case {a = 1 - O(1/n)} corresponds to {\alpha = O(1)}. Informally, {\alpha} measures how close {a} is to {1} (at the scale of {O(1/n)}).

A key role in the argument will be played by the mean

\displaystyle  x + iy := \frac{1}{n-1} \sum_{j=1}^{n-1} q_j \ \ \ \ \ (12)

of the {q_j}, particularly the real part {x}. As the {q_j} all lie in the unit disk, the mean {x+iy} does also, so that

\displaystyle  -1 \leq x \leq 1

and

\displaystyle  |y| \leq \sqrt{1-x^2}. \ \ \ \ \ (13)

On the other hand, in the example in Example 4, {x} is equal to the extremal value of {1}, and {y=0}. In Example 5, we have {x = 1 - O(1/n)} (and {y = O(1/n)}).

It will be convenient to work with the quadratic polynomial

\displaystyle  \beta(t) := 1 - 2atx + a^2 t^2 = 1 - x^2 + (x - at)^2 \ \ \ \ \ (14)

with a particular emphasis on the value at {t=1}:

\displaystyle  \begin{array}{rl} \beta(1) &= 1 - 2ax + a^2 \\ &= (1-a)^2 + 2a(1-x) = 1 - x^2 + (x-a)^2. \end{array} \ \ \ \ \ (15)

One should primarily think of {\beta(1)} as a measure of how close {x} is to {1}. Clearly we have

\displaystyle  \beta(t) > 0

for all {0 \leq t \leq 1} (note that {at} is strictly less than {1}).

The arguments will revolve around the relationship between {\alpha} and {\beta(1)}. Specifically, we will establish the following two inequalities below the fold. The first inequality, which we call the “polar inequality”, comes in three forms:

Proposition 10 (Polar inequality)

It will be the inequality (18) that we use in practice, but it will be derived from (17), which in turn is a consequence of (16), which will follow from the polar identity (8) together with the fact that the {z_j} and {q_j} lie in the unit disk. The bound (18) is only slightly weaker than (17); see the (Gemini-generated) image below.

I was not able to find an exact duplicate of the above polar inequalities in past literature, but the paper of Dégot contains several similar inequalities. The inequality (16) was extracted from (5.1) of Mazur’s AI-generated proof; the subsequent bounds (17), (18) arose from my attempts to simplify the arguments after that point.

The second inequality, which is more difficult, also will come in several forms:

Proposition 11 (Origin inequality) Let {n \geq 5}.

Part (i) (which was extracted with some effort from Section 6 of the original AI-generated argument) will be deduced from the first and second origin identities (9), (10), as well as the centroid identity (7). Part (ii) will follow from (i) and the polar inequality (18), while part (iii) is an elementary consequence of (i).

As it turns out, the last three terms in (21) are asymptotically negligible as {n \rightarrow \infty}. Dropping those terms gives a competing feasibility region for {\alpha} and {\beta(1)} which is disjoint from the one coming from the polar inequality (17) (or (18)):

This already suggests that one can use this approach to recover my previous result on Sendov’s conjecture holding for all sufficiently large {n}. In fact, even with the three error terms in (21) added, there is enough room between the two inequalities (18), (21) to obtain a contradiction for all {n \geq 5} (using the additional bound {\alpha \leq 17} to control these errors), although showing this for medium-sized {n} (such as {5 \leq n \leq 200}) requires a certain amount of computer assistance.

For fixed {\alpha}, the right-hand side of (21) is monotone increasing in {\beta(1)} (or equivalently, monotone decreasing in {x}). In view of (18), we can thus replace {\beta(1)} by {\frac{\alpha}{3+\alpha}} in this inequality, so that {ax = 1 - \frac{\alpha}{n-1} - \frac{\beta(1)}{2}} is replaced by

\displaystyle c(\alpha) := 1 - \frac{\alpha}{n-1} - \frac{\alpha}{2(3+\alpha)},

and {\beta(t) = 1 - 2axt + a^2t^2} replaced by {1 - 2c(\alpha) t + a^2 t^2}. The inequality (21) then becomes an inequality involving only {\alpha} and {n}:

\displaystyle  \begin{array}{rl} 1 \leq & \frac{1}{6} + \frac{1}{4(3+\alpha)} + \frac{1}{2(n-1)} + \frac{1}{4(n-1)(3+\alpha)} \\ & + \frac{a^4 n (n-1) (n-2)}{4(3+\alpha)} \int_0^1 t^3 (1 - 2c(\alpha) t + a^2 t^2)^{\frac{n-4}{2}}\ dt. \end{array} \ \ \ \ \ (22)

We also note that the bounds {0 \leq c(\alpha) \leq xa} force the constraint

\displaystyle  c(\alpha)^2 \leq a^2 = 1 - \frac{2\alpha}{n-1}.

This prevents {\alpha} from getting too close to the upper limit {\frac{n-1}{2}} (or {a} getting too close to zero).

We can now eliminate all large degrees, e.g., {n > 200}, as follows. The quadratic {1 - 2c(\alpha) t + a^2 t^2} attains its minimum at {t =c(\alpha)/a^2}. For {t \leq c(\alpha)/a^2} we have

\displaystyle  1 - 2c(\alpha) t + a^2 t^2 \leq 1 - c(\alpha) t \leq \exp( - c(\alpha) t)

while for {c(\alpha) / a^2 < t \leq 1} (if this region is non-vacuous) we can bound the quadratic by its value {\frac{\alpha}{3+\alpha}} at {t=1}. Thus

\displaystyle  \begin{array}{rl} & \int_0^1 t^3 (1 - 2c(\alpha) t + a^2 t^2)^{\frac{n-4}{2}}\ dt \\ \leq & \int_0^\infty t^3 \exp\left( - \frac{n-4}{2} c(\alpha) t\right)\ dt + \int_0^1 t^3 \left(\frac{\alpha}{3+\alpha}\right)^{\frac{n-4}{2}}\ dt. \end{array}

Evaluating these expressions, we arrive at

\displaystyle  \begin{array}{rl} 1 \leq & \frac{1}{6} + \frac{1}{4(3+\alpha)} + \frac{1}{2(n-1)} + \frac{1}{4(n-1)(3+\alpha)} \\ & + \frac{24 a^4 n (n-1) (n-2)}{(3+\alpha) (n-4)^4 c(\alpha)^4} + \frac{a^4 n (n-1) (n-2)}{16(3+\alpha)} \left(\frac{\alpha}{3+\alpha}\right)^{\frac{n-4}{2}}. \end{array}

Since {\alpha \leq 17}, we have {\frac{\alpha}{3+\alpha} \leq \frac{17}{20}}. Next, we claim that {(3+\alpha) c(\alpha)^4 \geq 1}. As {c(\alpha)} is monotone increasing in {n}, it suffices to do this when {n=201}. Here one can directly compute that

\displaystyle  \frac{d}{d\alpha} ((3+\alpha) c(\alpha)^4) = - c(\alpha)^3 \frac{5(3+\alpha)^2 - 103(3+\alpha) + 900}{200(3+\alpha)} < 0

since the discriminant {-7391} of the numerator is negative, we conclude that

\displaystyle  (3+\alpha) c(\alpha)^4 \geq (3+17) c(17)^4 \geq 1.1529\dots > 1

as desired.

Dropping some {\alpha} and {a} terms, we conclude that

\displaystyle  \begin{array}{rl} 1 \leq & \frac{1}{6} + \frac{1}{12} + \frac{7}{12(n-1)} + \frac{24 n (n-1) (n-2)}{(n-4)^4} \\ & + \frac{n (n-1) (n-2)}{48} \left(\frac{17}{20}\right)^{\frac{n-4}{2}}. \end{array}

Every term on the right-hand side can be seen to be decreasing in {n} for {n \geq 201}. Thus the right-hand side can be bounded by

\displaystyle  \begin{array}{rl} & \frac{1}{6} + \frac{1}{12} + \frac{7}{12 \times 200} + \frac{24 \cdot 201 \cdot 200 \cdot 199}{197^4} \\ & + \frac{201 \cdot 200 \cdot 199}{48} \left(\frac{17}{20}\right)^{\frac{201-4}{2}} \leq 0.399, \end{array}

giving the desired contradiction.

The remaining range to handle is when

\displaystyle  5 \leq n \leq 200; \quad 0 \leq \alpha \leq 17; \quad c(\alpha)^2 \leq 1 - \frac{2\alpha}{n-1}.

It turns out that (22) remains infeasible in this range. This can be illustrated numerically without much difficulty: see this applet. For instance, in the most delicate case {n=53}, the right-hand side of (22) only gets as large as {0.853} (and in particular stays below {1}) throughout the range {0 \leq \alpha \leq 17}:

I have also verified this bound in Lean.

— 1. The polar inequality —

We begin with a proof of Proposition 10.

As is well known, the Möbius transform {z \mapsto \frac{a-z}{1-az}} maps the closed unit disk to itself. In particular, we have

\displaystyle  \left|\frac{1-az_j}{a-z_j}\right| \geq 1

for all of the zeroes {z_j}. Inserting this into the polar identity (8) and using the triangle inequality, we conclude the lower bound

\displaystyle  \int_0^1 \prod_{j=1}^{n-1} |t (1-a^2) q_j + a|\ dt \geq 1. \ \ \ \ \ (23)

We now convert this bound to a bound involving the quantity {x} in (12). From the arithmetic mean-geometric mean inequality we have

\displaystyle  \prod_{j=1}^{n-1} |t (1-a^2) q_j + a| \leq \left(\frac{1}{n-1} \sum_{j=1}^{n-1} |t (1-a^2) q_j + a|^2\right)^{\frac{n-1}{2}} \ \ \ \ \ (24)

and from (12) we have

\displaystyle  \begin{array}{rl} \sum_{j=1}^{n-1} |t (1-a^2) q_j + a|^2 = & (n-1) a^2 + 2 (n-1) a t (1-a^2) x \\ & + t^2 (1-a^2)^2 \sum_{j=1}^{n-1} |q_j|^2. \end{array}

Since {|q_j|^2 \leq 1}, we thus have

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} |t (1-a^2) q_j + a|^2 \leq a^2 + 2 a t x (1-a^2) + t^2 (1-a^2)^2 \ \ \ \ \ (25)

giving the raw polar inequality (16).

Bounding {t^2} by {t} and using the quantities {\alpha,\beta(1)} from (11), (15), we observe that

\displaystyle  a^2 + 2 a t x (1-a^2) + t^2 (1-a^2)^2 \leq 1 + \frac{2}{n-1} \alpha (-1 + (2-\beta(1)) t).

Using the basic inequality {1 + y \leq \exp(y)}, we thus have

\displaystyle  (a^2 + 2 a t x (1-a^2) + t^2 (1-a^2)^2)^{\frac{n-1}{2}} \leq \exp ( \alpha (-1 + (2-\beta(1)) t)),

with strict inequality for {t < 1}. From (16) we conclude (17). This also implies {\beta(1) < 1}, since otherwise the integrand is always bounded by {1}, which is absurd.

On evaluating the integral in (17), we obtain

\displaystyle 1 < \frac{e^{\alpha (1 - \beta(1))}- e^{-\alpha}}{\alpha (2 - \beta(1))}

and thus

\displaystyle  e^{\alpha (1 - \beta(1))} > \alpha (2 - \beta(1)) + e^{-\alpha} \geq \alpha,

so on taking logarithms we obtain

\displaystyle  1 - \beta(1) > \frac{\log \alpha}{\alpha}.

It remains to establish the bound

\displaystyle  \beta(1) < \frac{\alpha}{3+\alpha}. \ \ \ \ \ (26)

Here we use an AI-generated argument. One can directly calculate

\displaystyle  \int_0^1 \exp ( \alpha (-1 + (2-\beta(1)) t))\ dt = e^{-u} \frac{\sinh h}{h}

where {u = \alpha \beta(1)/2} and {h = \alpha (2-\beta(1))/2}. If we can show that

\displaystyle  \log \frac{\sinh h}{h} \leq \sqrt{h^2+9} - 3 \ \ \ \ \ (27)

for all {h > 0}, then taking logarithms in (17) yields

\displaystyle  u < \log \frac{\sinh h}{h} \leq \sqrt{h^2+9} - 3

from which (26) will follow by routine algebra.

Both sides of (27) vanish at {h=0}. Taking derivatives, it suffices to show that

\displaystyle  \coth h - \frac{1}{h} \leq \frac{h}{\sqrt{h^2+9}}

which rearranges to

\displaystyle  h^4 \sinh^2 h - (h^2+9) (h \cosh h - \sinh h)^2 \geq 0.

To expand the left-hand side, we use the double angle formulae {\sinh^2 h = \frac{\cosh 2h - 1}{2}} and

\displaystyle  (h \cosh h - \sinh h)^2 = \frac{(h^2+1) \cosh 2h + (h^2-1)}{2} - h \sinh 2h

to rewrite it as

\displaystyle  \frac{h^4 (\cosh 2h - 1)}{2} - (h^2+9) \left( \frac{(h^2+1) \cosh 2h + (h^2-1)}{2} - h \sinh 2h \right).

Collecting the coefficient of {h^{2k}} for {k \geq 3} and extracting a common factor of {\frac{2^{2k-3}}{(2k)!}}, one is left with

\displaystyle  \begin{array}{rl} & N (N-1) (N-2) - 10 N (N-1) + 36 N - 36 \\ &= (N-1) (N^2 - 12N + 36) = (N-1) (N-6)^2 \end{array} \ \ \ \ \ (28)

where {N := 2k}. (The remaining coefficients, which also receive contributions from the polynomial terms, all vanish.) Thus the left-hand side has the Taylor expansion

\displaystyle  \sum_{k=4}^\infty \frac{2^{2k-3} (2k-1) (2k-6)^2}{(2k)!} h^{2k},

in which every coefficient is non-negative, giving the claim.

Remark 12 As the image in the introduction suggests, the bound (18) is only slightly weaker than (17). For small {\alpha}, one can perform Taylor approximation on the latter bound to obtain

\displaystyle  \beta(1) < \frac{\alpha}{3} - \frac{\alpha^2}{9} + \frac{19 \alpha^3}{540} - \frac{17 \alpha^4}{1620} + O(\alpha^5)

while the former bound is

\displaystyle  \beta(1) < \frac{\alpha}{3} - \frac{\alpha^2}{9} + \frac{\alpha^3}{27} - \frac{\alpha^4}{81} + O(\alpha^5).

Note that {\frac{19}{540} = 0.0351\dots} is slightly smaller than {\frac{1}{27} = 0.0370\dots}.

Relating to this, the constant {9} in (27) cannot be improved.

— 2. The origin inequality —

Now we turn to the proof of Proposition 11, which is more difficult and revolves around an analysis of the function {F} defined in (6). We begin with a heuristic analysis. Inserting the approximation {1 - \varepsilon \approx \exp(-\varepsilon)} for small {\varepsilon} into (6) and using (12), we are led to the approximation

\displaystyle  F(t) \approx \exp( - (n-1) a t (x+iy) ), \ \ \ \ \ (29)

at least when {t} is small (which turns out to be the dominant regime in applications). This suggests a relation

\displaystyle  \int_0^1 F(t)\ dt \approx \frac{1 - F(1)}{(n-1) a (x+iy)} \ \ \ \ \ (30)

between the two expressions involving {F} in the origin identities in Lemma 6. Substituting in this approximation, we obtain some (slightly complicated) approximation for the sum {\sum_{j=1}^{n-1} \frac{1}{z_j}} in terms of {a}, {n}, {x}, {y}, and the product {J := \prod_{j=1}^{n-1} z_j q_j}.

As {z_j, q_j} lie in the closed unit disk, the product {J} does also. However, past experience with the Sendov conjecture has taught us that the worst cases tend to be when {z_j, q_j} lie very close to the boundary of the disk, so that {|J|} is close to one. For instance, in Example 4 all the {z_j} and {q_j} lie on the unit circle, and {J = (-1)^{n+1}}. See Remark 3 of Dégot or Theorem 1.10(ii) of my own paper for other places where this heuristic is noted. To simplify the discussion, let us assume for now that {|J|} is exactly one, so that {z_j, q_j} all lie on the unit circle. This leads in particular to the inversion identities

\displaystyle  \frac{1}{z_j} = \overline{z_j}; \quad \frac{1}{q_j} = \overline{q_j}. \ \ \ \ \ (31)

The centroid identity in Lemma 6(i) relates the sum of the {z_j} with the sum of the {1/q_j}. Using (31), this gives a similar identity relating the sum of the {1/z_j} with the sum of the {q_j}. The latter sum is of course just {(n-1) (x+iy)}. This combines well with the previous approximation, thus giving an approximate identity relating {x}, {y} to {J}, {a}, and {n}. As it turns out, the roles of {y} and {J} are minor and can be quickly eliminated for the purposes of obtaining useful bounds, leading eventually to the relation in Proposition 11.

We turn to the details. To make the approximation (30) more precise, we note that {F(0)=1}, and hence by the fundamental theorem of calculus

\displaystyle  1 = F(1) - \int_0^1 F'(t)\ dt.

The heuristic (29) predicts that {F'(t) \approx -(n-1) a(x+iy) F(t)}, which would give (30). If we actually differentiate (6) carefully, we obtain the exact identity

\displaystyle  F'(t) = - (n-1) a(x+iy) F(t) - \sum_{j=1}^{n-1} a^2 t q_j^2 \prod_{k \neq j} (1 - atq_k).

Bounding {|q_j^2| \leq 1}, we write this

\displaystyle  F'(t) = - (n-1) a(x+iy) F(t) + O_{\leq} \left( a^2 t \sum_{j=1}^{n-1} \prod_{k \neq j} |1 - atq_k| \right).

When faced with a similar expression in (24), we used the arithmetic mean-geometric mean inequality. Here, the analogous tool is Maclaurin’s inequality, which gives

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} \prod_{k \neq j} |1 - atq_k| \leq \left(\frac{1}{n-1} \sum_{j=1}^{n-1} |1 - at q_j|\right)^{n-2},

and hence by Cauchy–Schwarz

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} \prod_{k \neq j} |1 - atq_k| \leq \left(\frac{1}{n-1} \sum_{j=1}^{n-1} |1 - at q_j|^2\right)^{\frac{n-2}{2}}.

Repeating the calculations used to show (25), we have

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} |1 - at q_j|^2 \leq 1 - 2 a x t + a^2 t^2 = \beta(t)

and so we obtain the bound

\displaystyle  F'(t) = - (n-1) a(x+iy) F(t) + O_{\leq} \left( a^2 t (n-1) \beta(t)^{\frac{n-2}{2}} \right).

Integrating this, we obtain a rigorous analogue of (30),

\displaystyle  \begin{array}{rl} 1 - F(1) = & (n-1) a(x+iy) \int_0^1 F(t)\ dt \\ & + O_{\leq} \left( \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt \right) \end{array}

and thus by the triangle inequality

\displaystyle  \begin{array}{rl} 1 \leq & \left|F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt\right| \\ & + \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt. \end{array} \ \ \ \ \ (32)

From the first and second origin identities (9), (10) we have

\displaystyle  \begin{array}{rl} & F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt \\ &= \frac{(-1)^{n-1} J}{n} \left(1 + a \sum_{j=1}^{n-1} \frac{1}{z_j} + (n-1) a(x+iy) \right). \end{array} \ \ \ \ \ (33)

The next step is thus to estimate {\sum_{j=1}^{n-1} \frac{1}{z_j}}. When {|J|=1}, then all the {z_j, q_j} were on the unit circle and we could use (31) (and the centroid identity) to proceed. Now, we are no longer assuming {|J|} to equal {1}, but we can still adapt the previous arguments with a loss proportional to {1 - |J|^2}. The key lemma is

Lemma 13 (Defect lemma) Let {w_1,\dots,w_N} be some points in the closed unit disk. Then

\displaystyle  \prod_{j=1}^N |w_j| \times \sum_{j=1}^N \left|\frac{1}{w_j} - \overline{w_j}\right| \leq 1 - \prod_{j=1}^N |w_j|^2.

Proof: By a limiting argument we may assume that none of the {w_j} vanish. If we write {|w_j| = e^{-a_j}} for some {a_j \geq 0}, then we can calculate that

\displaystyle  \left|\frac{1}{w_j} - \overline{w_j}\right| = 2 \sinh a_j

and

\displaystyle  1 - \prod_{j=1}^N |w_j|^2 = \prod_{j=1}^N |w_j| \times 2 \sinh \sum_{j=1}^N a_j.

Thus the desired inequality reduces to the superadditivity property

\displaystyle  \sinh \sum_{j=1}^N a_j \geq \sum_{j=1}^N \sinh a_j.

But from the sinh addition formula {\sinh(a+b) = \sinh a \cosh b + \sinh b \cosh a} we have

\displaystyle  \sinh(a+b) \geq \sinh a + \sinh b

for all non-negative {a,b} (this also follows from the convex nature of {\sinh} together with {\sinh 0 = 0}), and the claim follows by induction. \Box

We remark that the lemma can also be proven by direct induction, without an appeal to hyperbolic trigonometry.

From taking complex conjugates of the centroid identity (7) and performing some algebra, we have

\displaystyle  J \sum_{j=1}^{n-1} \left(\frac{n-1}{n} \overline{z_j} + \frac{1}{\overline{q_j}}\right) = a J \frac{(n-1)^2}{n}.

Using the defect lemma (applied to the points {z_1,\dots,z_{n-1},\overline{q_1},\dots,\overline{q_{n-1}}}) and the triangle inequality we conclude that

\displaystyle  J \sum_{j=1}^{n-1} \left(\frac{n-1}{n} \frac{1}{z_j} + q_j\right) = a J \frac{(n-1)^2}{n} + O_{\leq}\left( 1 - |J|^2 \right),

where as before we are removing singularities when some of the {z_j} vanish. Applying (12) and some algebraic manipulation, we arrive at

\displaystyle  J \sum_{j=1}^{n-1} \frac{1}{z_j} = a J (n-1) - n J (x+iy) + O_{\leq}\left( \frac{n}{n-1} (1 - |J|^2) \right)

Substituting this back into (33), we conclude that

\displaystyle  \begin{array}{rl} & \left|F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt\right| \\ &= \left| \frac{|J|}{n} + \frac{a^2 |J|(n-1)}{n} - a|J| (x+iy) + \frac{a (n-1) (x+iy) |J|}{n} \right| \\ & \quad + O_{\leq}\left( \frac{a}{n-1} (1-|J|^2) \right) \end{array}

and hence after some algebra and the triangle inequality

\displaystyle  \begin{array}{rl} & \left|F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt\right| \\ &\leq |J| \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| + \frac{a}{n-1} (1-|J|^2). \end{array}

Inserting this into (32), we obtain

\displaystyle  \begin{array}{rl} 1 \leq & |J| \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| + \frac{a}{n-1} (1-|J|^2) \\ & + \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt. \end{array} \ \ \ \ \ (34)

We can simplify (34) by reducing to the {|J|=1} case. Indeed, we shall show that

\displaystyle  \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| \geq 2 \frac{a}{n-1} \ \ \ \ \ (35)

which implies that the right-hand side of (34) is non-decreasing in {|J|} in the range {0 \leq |J| \leq 1}. Thus we may replace {|J|} by {1} in (34) to conclude that

\displaystyle  1 \leq \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| + \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt. \ \ \ \ \ (36)

Let us now verify (35). Using {|x+iy| \leq 1} and the triangle inequality, we can lower bound

\displaystyle  \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| \geq \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a}{n}.

Inserting this into (35) and clearing denominators, we reduce after some algebra to

\displaystyle  (n-1)^2 a^2 - (3n-1) a + n-1 > 0.

But as a quadratic polynomial in {a}, the left-hand side has discriminant {(3n-1)^2 - 4(n-1)^3}, which one can check to be negative for sufficiently large {n} (in fact {n \geq 5} suffices), giving the claim (35).

Next we eliminate the role of the imaginary term {iy}. Observe for any complex number {s+it} with positive real part that

\displaystyle  |s+it| \leq s + \frac{t^2}{2s}

as can be seen by squaring both sides. The expression {\frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n}} has real part

\displaystyle  \frac{a^2 (n-1)}{n} + \frac{1-ax}{n}

which lies between {\frac{a^2 (n-1)}{n}} and {1} (in particular, it is positive), and imaginary part of magnitude at most

\displaystyle  \frac{a}{n} (1 - x^2)^{1/2}

by (13). We conclude that

\displaystyle  \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| \leq \frac{a^2 (n-1)}{n} + \frac{1-ax}{n} + \frac{1-x^2}{2n(n-1)}.

The right-hand side can be rearranged using the quantity {\alpha} from (11) as

\displaystyle  1 - \frac{2\alpha}{n} - \frac{ax}{n} + \frac{1-x^2}{2n(n-1)},

so the bound (36) gives (19).

— 2.1. Upper bound on {\alpha}

Now we can prove (20). Suppose for contradiction that {\alpha > 17}; since {\alpha \leq \frac{n-1}{2}}, this implies that {n \geq 36}. Crudely discarding the {ax} term in (19) and bounding {\frac{1-x^2}{2(n-1)}} by {\frac{1}{2(n-1)}}, we have

\displaystyle  2 \alpha \leq \frac{1}{2(n-1)} + a^2 n(n-1) \int_0^1 t \beta(t)^{\frac{n-2}{2}}\ dt.

The quadratic polynomial {\beta(t)} equals {1} at {t=0} and attains its minimum at {t = x/a} with value {1-x^2}. By convexity, we thus have

\displaystyle  \beta(t) \leq 1 - axt \leq \exp(-axt)

for {0 \leq t \leq x/a} and

\displaystyle  \beta(t) \leq \beta(1)

for {x/a \leq t \leq 1} (this latter statement is vacuous if {x/a > 1}). Since {\int_0^1 t\ dt = \frac{1}{2}}, we can therefore crudely bound

\displaystyle  \int_0^1 t (1 - 2 a x t + a^2 t^2)^{\frac{n-2}{2}}\ dt \leq \int_0^\infty t \exp\left( - \frac{n-2}{2} axt\right)\ dt + \frac{1}{2} \beta(1)^{\frac{n-2}{2}}

\displaystyle  = \frac{4}{(n-2)^2 a^2 x^2} + \frac{1}{2} \beta(1)^{\frac{n-2}{2}},

and hence

\displaystyle  2 \alpha \leq \frac{1}{2(n-1)} + \frac{4n(n-1)}{(n-2)^2 x^2} + \frac{a^2 n(n-1)}{2} \beta(1)^{\frac{n-2}{2}}.

From (15) we have {x^2 \geq 1-\beta(1)}, thus by (18) one has

\displaystyle  \frac{1}{x^2} \leq \frac{\alpha}{\log \alpha}.

From another application of (18) one has

\displaystyle  \beta(1) \leq \exp( - (1 - \beta(1)) ) \leq \alpha^{-1/\alpha}.

We conclude that

\displaystyle  2 \alpha \leq \frac{1}{2(n-1)} + \frac{4n(n-1)}{(n-2)^2 \log \alpha} \alpha + \frac{a^2 n(n-1)}{2} \alpha^{-\frac{n-2}{2\alpha}}.

It is now convenient to introduce the quantity {u := \frac{a^2}{1-a^2}}, thus {0 < u < \infty} with

\displaystyle  a^2 = \frac{u}{1+u}

and

\displaystyle  \frac{n-2}{2 \alpha} = 1 + u - \frac{1}{2\alpha}.

Inserting these bounds and dividing by {\alpha}, we conclude

\displaystyle  2 \leq \frac{1}{2(n-1) \alpha} + \frac{4n(n-1)}{(n-2)^2 \log \alpha} + \frac{u}{2(1+u) \alpha^2} n(n-1) \alpha^{-u} \alpha^{\frac{1}{2\alpha}}.

Since {\alpha = \frac{n-1}{2(1+u)}}, we obtain

\displaystyle  2 \leq \frac{1}{2(n-1) \alpha} + \frac{4n(n-1)}{(n-2)^2 \log \alpha} + \frac{2n}{n-1} u (1+u) \alpha^{-u} \alpha^{\frac{1}{2\alpha}}.

Since {n \geq 36} and {\alpha \geq 17}, we conclude that

\displaystyle  2 \leq \frac{1}{2 \times 35 \times 17} + \frac{4 \times 36 \times 35}{(34)^2 \log 17} + \frac{2 \times 36}{35} u (1+u) 17^{-u} 17^{\frac{1}{2 \times 17}}.

Routine calculus shows that {u(1+u) 17^{-u}} has a maximum of at most {0.1825}, and that the right-hand side here is at most {1.948}, giving the required contradiction. This proves (20).

— 2.2. A simplified estimate —

Now we show (21). Note from (11) that

\displaystyle  1-a^2 = \frac{2\alpha}{n-1} \ \ \ \ \ (37)

while from (15) we have

\displaystyle  1 - x^2 \leq \beta(1) \ \ \ \ \ (38)

and hence also

\displaystyle  \frac{1}{x^2} \leq 1 + \frac{\beta(1)}{1-\beta(1)}. \ \ \ \ \ (39)

From (15) we have

\displaystyle  \beta(t) = (1 - axt)^2 + a^2 t^2 (1-x^2).

By the mean value theorem (noting that {1-axt} is non-negative) we thus have

\displaystyle  \beta(t)^{\frac{n-2}{2}} \leq \left((1 - axt)^2\right)^{\frac{n-2}{2}} + a^2 t^2 (1-x^2) \frac{n-2}{2} \beta(t)^{\frac{n-4}{2}}.

From the standard beta function identity

\displaystyle  \int_0^{1/ax} t (1 - axt)^{n-2}\ dt = \frac{1}{a^2 x^2 n(n-1)}

(and the fact that {1/ax \geq 1}) we can thus replace (19) by

\displaystyle  \begin{array}{rl} 2\alpha + ax \leq & \frac{1-x^2}{2(n-1)} + \frac{1}{x^2} \\ & + \frac{a^4 n (n-1) (n-2) (1-x^2)}{2} \int_0^1 t^3 \beta(t)^{\frac{n-4}{2}}\ dt. \end{array}

From (15) we have

\displaystyle ax = 1 - \frac{\beta(1)}{2} - \frac{1-a^2}{2}.

Thus by (37), (38), (39)

\displaystyle  \begin{array}{rl} 2\alpha \leq & \frac{\beta(1)}{1-\beta(1)} + \frac{\beta(1)}{2} + \frac{\alpha}{n-1} + \frac{\beta(1)}{2(n-1)} \\ & + \frac{a^4 n (n-1) (n-2) \beta(1)}{2} \int_0^1 t^3 \beta(t)^{\frac{n-4}{2}}\ dt. \end{array}

Dividing by the positive quantity {2\alpha} gives the claim.

— 3. Rubinstein’s theorem —

We now adapt the arguments to give a proof of Rubinstein’s theorem that the Phelps–Rodriguez conjecture holds in the {a=1} case, i.e.,

Theorem 14 (Rubinstein’s theorem) Let {n \geq 2}, and let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. If {p(1)=0}, then there exists a critical point {\zeta} of {p} with {|\zeta-1| < 1}, unless {p} is a scalar multiple of {z^n - 1}.

Taking contrapositives, we may assume that the critical points {\zeta_j} are of the form {1 - 1/q_j} for some {q_j} in the closed unit disk, and normalize {p} to be monic; our task is to show that {p(z) = z^n-1}.

The polar identity (8), based on calculating {p(1/a) / p'(a)} degenerates to a triviality when {a=1}, but we have the following usable substitute, valid for any choice of {a}, first observed in equation (3.2) of Meir–Sharma:

Lemma 15 (Meir–Sharma identity) If {p(a)=0} and the critical points are of the form {a - 1/q_j} then all the zeroes {z_j} are not equal to {a}, and

\displaystyle  \sum_{j=1}^{n-1} q_j = 2 \sum_{j=1}^{n-1} \frac{1}{a-z_j}.

Proof: By hypothesis, {a} is not a critical point of {p}, so {p'(a) \neq 0} and {z_j \neq a} for all {j}. Instead of computing {p(1/a)/p'(a)}, we instead consider the expression {p''(a) / p'(a)}. On the one hand, from (4) we have

\displaystyle  p'(a) = n \prod_{j=1}^{n-1} \frac{1}{q_j}

while from differentiating (4) we have

\displaystyle  p''(a) = n \prod_{j=1}^{n-1} \frac{1}{q_j} \times \sum_{j=1}^{n-1} q_j.

Meanwhile, from (3) we have

\displaystyle  p'(a) = \prod_{j=1}^{n-1} (a-z_j)

and from differentiating (3) we have

\displaystyle  p''(a) = \left(\prod_{j=1}^{n-1} (a-z_j)\right) \times \left( \sum_{j=1}^{n-1} \frac{1}{a-z_j} + \sum_{j=1}^{n-1} \frac{1}{a-z_j} \right).

Using these identities to compute {p''(a) / p'(a)} in two different ways gives the claim. \Box

Now take {a=1}. Since {z_j} lie in the closed unit disk, {\frac{1}{1-z_j}} has real part at least {1/2}, while {\mathrm{Re} q_j} is at most {1}. Thus, the only way that the above identity can hold is if {\mathrm{Re} q_j = 1} for all {j}, hence {q_j = 1} for all {j}. Thus all critical points are at the origin, which forces {p(z) = z^n - c} for some {c}. Since {p(1)=0}, we conclude that {c=1}, giving the claim.

— 4. The {n \leq 4} cases —

We now prove the {n \leq 4} cases of Conjecture 3. The starting point is (23). Using the triangle inequality and {|q_j| \leq 1}, this implies that

\displaystyle  1 \leq \int_0^1 (a+(1-a^2) t)^{n-1}\ dt.

(This also follows from (16) and {x \leq 1}.) From Hölder’s inequality and {n \leq 4} we conclude that

\displaystyle  1 \leq \int_0^1 (a+(1-a^2) t)^3\ dt.

The right-hand side can be computed to equal

\displaystyle  1 - \frac{(a-1)^2}{4} (a^2 (1-a)^2 + 3(1-a^2) + 2a),

which is obviously less than {1} for {0 \leq a < 1}, giving the contradiction.

Remark 16 The same argument also works for {n=5}, but breaks down for higher {n}.

— 5. Further directions —

The Sendov and Phelps–Rodriguez conjectures are now resolved, but several related conjectures remain open. The following strengthening of Sendov’s conjecture, by Borcea, is open for any {q < \infty}:

Conjecture 17 (Borcea conjecture) Let {1 \leq q < \infty} and {n \geq 2}, and let {p} be a degree {n} polynomial with zeroes {z_1,\dots,z_n} satisfying {\frac{1}{n} \sum_{j=1}^n |z_j|^q \leq 1}. Then for every zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|\zeta - a| \leq 1}.

Sendov’s conjecture is the limiting case {q=\infty} of this conjecture. There has been relatively little progress on this conjecture: the cases {(q,n) = (1,3), (2,4)} were established by Khavinson, Pereira, Putinar, Saff, and Shimorin, and in this previous paper we reported the negative result that AlphaEvolve failed to find a counterexample to the conjecture. The proof methods here do not seem to extend easily; all the identities relating zeroes and critical points continue to hold, but now that the {z_j} are only constrained to the unit disk in an averaged moment sense, all of the inequalities developed above now fail.

Another strengthening of Sendov’s conjecture that remains open is Schmeisser’s conjecture:

Conjecture 18 (Schmeisser’s conjecture) Let {n \geq 2}, and let {p} be a degree {n} polynomial with all zeroes in the closed unit disk. Then for any {a} in the convex hull of the zeroes of {p}, there exists a critical point {\zeta} of {p} with {|\zeta - a| \leq 1}.

Schmeisser proved several special cases of this conjecture, and AlphaEvolve again failed to find a counterexample, but there has not been much further progress. Here, the {z_j} are now back in the closed unit disk, but we no longer have {p(a)=0}, again rendering most of the previous identities invalid. But perhaps some modification of the arguments here can make some progress on this conjecture.

A common generalization of the Borcea and Schmeisser conjectures was proposed in Conjecture 2.4 of this paper of Zhang.

Another well known variant of Sendov’s conjecture is Smale’s problem:

Conjecture 19 (Smale’s problem) Let {n \geq 2}, and let {p} be a degree {n} polynomial. Then for any zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|p(\zeta)| \leq (1-\frac{1}{n}) |\zeta-a| |p'(a)|}.

The constant {1-\frac{1}{n}} is best possible, as can be seen by the example {p(z) = z^n-z} and {a=0}. Using the Koebe one-quarter theorem, Smale proved this conjecture with {1 - \frac{1}{n}} replaced by {4}. Some slight improvements of this bound have been obtained over the years; for instance for {n \geq 8}, the improved bound of {4 - \frac{2.263}{\sqrt{n}}} was obtained by Crane. Again, AlphaEvolve failed to find a counterexample to this conjecture. This problem does not seem to have a direct relationship with Sendov’s conjecture, and there is no useful normalization of the zeroes and critical points that is confined to the unit disk. Nevertheless there may be some hope of making progress on this conjecture, perhaps working first in the asymptotic regime {n \rightarrow \infty}.

Needless to say, I did try some desultory attempts to use AI tools to attack these questions, but without much notable success.

One potential way forward is to find further proofs of Sendov’s conjecture that utilize other techniques that might be more broadly applicable to this larger family of problems. The proof here is remarkable in that the zeroes and critical points are treated almost as independent mathematical objects, communicating with each other only very narrowly through four identities in which one only inspects the underlying polynomial (and its derivative) at a small number of points. It could be that an approach focusing on more global features of the polynomial may lead to new proofs of Sendov’s conjecture, and perhaps also of its generalizations.

John BaezThree Generations in E7

It’s long been a mystery why there are 3 generations of quarks and leptons: three sets of particles, apparently identical except for how they interact with the Higgs boson. It would be nice if there were some good physical explanation. Nobody knows one. Barring that, it would be nice if some beautiful mathematical structure made this pattern seem natural. That’s what my new paper is about.

It’s my third paper about exceptional algebraic structures and the Standard Model. When you classify famous gadgets in algebra, beautiful gadgets with fancy names like ‘simple Lie algebras’ and ‘Euclidean Jordan algebras’ and ‘positive hermitian Jordan pairs’, you tend to get infinite series of them—together with a few exceptions that can be built using the octonions. This is a bit spooky, so I’ve been interested in this for a long time.

A few physicists have hoped that these exceptions are good for something. For example, maybe the quirky features of our best theory of particle physics, the Standard Model, aren’t accidental. Perhaps they fall out naturally from some exceptional algebraic structure.

It’s a long shot, but we’ve been stuck on figuring out new fundamental laws of particle physics for so long—roughly since the early 1980s—that it’s worth a try.

In 2018, Michel Dubois-Violette and Ivan Todorov noticed that the gauge group of the Standard Model falls out as symmetries of the so-called ‘exceptional Jordan algebra’ together with some ordinary Jordan algebras sitting inside it. I tried to clarify that here, with a huge amount of help from an excellent young mathematician:

• John Baez and Paul Schwahn, The Standard Model gauge group from the exceptional Jordan algebra. (Blog article here.)

It’s very nice, because the Jordan algebras in question arise naturally when you try to axiomatize the foundations of quantum physics. It would be so cool if something about quantum physics made the Standard Model seem mathematically natural!

But really this result only concerns the gauge bosons in the Standard Model: the photon, gluons, and the W and Z bosons. It says nothing about the fermions—that is, the quarks and leptons. And it seems quite hard to get those into the picture.

In 2020, Latham Boyle tried to solve this problem by tensoring the exceptional Jordan algebra with the complex numbers. This made one generation of fermions appear quite naturally! But the connection to the foundations of quantum physics seemed lost: tensoring the exceptional Jordan algebra with the complex numbers seems at first like it might be just a formal trick.

This spring, Latham and his student Endre Bokor and I showed the connection to quantum physics is not lost:

• John Baez, Endre Bokor and Latham Boyle, Jordan pair quantum theory and the Standard Model. (Blog article here.)

The idea is to work, not with Jordan algebras, but with more general things called Jordan pairs, which have been studied by mathematicians since at least 1975. We showed that you can still do quantum physics with Jordan pairs. And we showed that there’s an ‘exceptional’ Jordan pair that naturally contains the Standard Model gauge group and one generation of fermions!

This Jordan pair is built from the bioctonions: the octonions tensored with the complex numbers. And it’s closely related to an exceptional Lie algebra called \mathfrak{e}_6.

This is nice because the work of Dubois-Violette and Todorov used a smaller exceptional Lie algebra called \mathfrak{f}_4. Going up to \mathfrak{e}_6 gives the room to include one generation of fermions.

There’s an even larger exceptional Lie algebra you can use to build a Jordan pair: it’s called \mathfrak{e}_7. Bokor, Boyle and I tried using this to get three generations of fermions. There are things that make this tempting: not just the fact that \mathfrak{e}_7 is bigger, but the fact that the Jordan pair you get from it has a kind of three-fold symmetry. But we couldn’t get it to work.

Around this time I got very interested in some work that someone had sent me in October 2025. My inbox is packed with new theories of physics. Since the rise of large language models the inflow has increased: I get about two emails a day from someone telling me they’ve made a revolutionary discovery in physics. Practically none of these theories appeal to me. But this paper, and this thesis, were different:

• Benjamin Nasmith, An exceptional combinatorial sequence and Standard Model particles, 2020.

• Benjamin Nasmith, Tight Projective 5-Designs and Exceptional Structures, Ph.D. thesis, Royal Military College of Canada, 2023.

He claimed to fit three generations of fermions into the exceptional Lie algebra \mathfrak{e}_7.

When I started seriously trying to understand this paper, I wound up translating it into a language I’m more comfortable with, and expanding on the ideas a bit. So I wrote this:

• John Baez, Three generations in \mathfrak{e}_7.

Here’s the basic idea.

The idea

There is a standard way to fit the Lie algebra of the Standard Model gauge group, which I call \mathfrak{g}_{\text{SM}}, into the Lie algebra \mathfrak{e}_7. You can construct a Lie algebra L that fits between them:

\mathfrak{g}_{\text{SM}} \subset L  \subset \mathfrak{e}_7

As a vector space we have

\mathfrak{e}_7 \; \cong \; L \oplus V

for some vector space V of dimension 3 \times 32.

Moreover, the Lie algebra \mathfrak{g}_{\text{SM}} acts on V, via the \mathfrak{e}_7 Lie bracket, precisely as it does on three generations of Standard Model fermions and their antiparticles, including right-handed neutrino and its antiparticle—but ignoring spin!

There is, in fact, a very interesting three-fold symmetry built into \mathfrak{e}_7, which is revealed when we put the Standard Model Lie algebra \mathfrak{g}_{\text{SM}} into it. It permutes the three generations.

Like Nasmith, I am not proposing a theory of physics. I’m only observing a fascinating mathematical pattern that might (or might not) be of some use in physics.

There are lots of things this pattern does not include: basically, everything I didn’t already mention. It does not include the spin of the fermions and gauge bosons. It does not include the Higgs boson, though in some sense it comes close (see the paper). It does not include a Lagrangian, so it doesn’t say anything at all about particle masses or more generally how particles are coupled to the Higgs boson.

I could say a lot more about this… most importantly, where does this Lie algebra L come from! The details are very interesting. There’s also the curious role of the right-handed neutrinos. But I’ve already spent weeks explaining all these things in my paper, so I won’t do it here. Instead let me say a bit about how I wrote the paper.

Writing the paper

I’ve been wanting to keep up with how AI is transforming math. About a year ago a friend gave me a subscription to Claude Pro. I wanted to test it out, despite my many misgivings, including how large language models are contributing to global warming and income inequality. Given the amazing things that people have recently done in math using large language models, I didn’t think that never trying them out would put me in the best position to make good decisions about the future.

So, I wrote this paper with help from Claude Opus 4.8.

I started by giving it Nasmith’s paper and asking a long series of questions about that paper over several days. The results were very interesting and helpful. Eventually I asked it to summarize and expand on our conversation. It quickly spat out a 10-page paper.

This paper was written in a breezy, pleasant style—but also quite hard to understand in detail, since it mixed Nasmith’s terminology with the Lie algebra terminology I prefer, and the proofs skipped over some steps.

It took me about three weeks of hard work to fully understand and re-express all the ideas a way that I like. For a while I felt dumb and frustrated, because when I asked Claude to fill in the gaps in proofs, it used math I was not very competent in, like the theory of regular subalgebras, and the theory of minuscule representations. But I learned this math, and everything turned out to be basically correct—in part, I’m sure, because Nasmith’s original work was correct.

For several weeks I checked, reorganized, expanded and completely rewrote this material. By the end everything was written in a style I like, emphasizing the ideas I consider important, proving things fairly carefully, and adding a lot of expository material—for example, explaining the theory of regular subalgebras.

Almost no traces of Claude’s original writeup remain, even though I was deeply influenced by them. My proofs make few references to deep theorems, though they assume solid familiarity with simple Lie algebras and their root systems. The proofs also require no brutally hard computations—though Claude was eager to do such computations to check things.

Any mistakes in this paper are my own.

I’m not sure what conclusions I draw from writing this paper. I’m writing another math paper now, with a human coauthor, and I have no desire to get help from a large language model. For work on my own it could be very helpful. Jacob Tsimerman says it roughly doubles his productivity. Would using it be so bad for the environment, or so bad for society, that I should avoid it? Maybe. I deliberately stuck with Claude Opus 4.8 instead of something more powerful, to see what I could do with what you get from a $20/month subscription. But maybe that’s still bad.

I avoid flying to conferences, which in some ways cripples my ability to keep up with new trends and influence people—but I don’t mind that. It gives me more time to think.

I will think carefully about my next move.

Tim GowersWhat sort of maths are LLMs good at?

For the sake of anyone who might read this blog post in the distant future (a month from now, say), let me mention that I am writing it a few days after OpenAI announced that it had solved ten major problems in mathematics and theoretical computer science, including the first construction of a non-sofic group, and a proof that the multicolour Ramsey number R(3,3,...,3) (where there are k 3’s) grows superexponentially in k. The first was, to judge from various talks I have been to, one of the most important unsolved problems in group theory, and the second was a major open problem in Ramsey theory that I didn’t necessarily expect to see solved in my lifetime, though of course such expectations now have to be revised. The reason I want to be clear about the timing is that I shall be discussing the current capabilities of LLMs in the full expectation that those will continue to change rapidly. So it is likely that in not too long from now, if there is anything interesting in what I write, it will be interesting mainly as a record of what the situation looked like in early August 2026.

These results, and the other eight on the list, are extraordinarily impressive, but it still doesn’t seem to be the case that LLMs are better than all humans at all aspects of mathematics. If they were, then their big speed advantage over us would mean that there would be much more of a flood of results. So it is natural to wonder about what kinds of problems LLMs are good at, and about where there is still room for improvement. I don’t pretend to have a good answer to this question, where a good answer would be a crisp classification that would fit the current examples well, but it is an interesting exercise to try to rule out some bad answers, and to try to identify potential answers that aren’t obviously contradicted by the evidence.

Are LLMs particularly good at finding counterexamples?

A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian conjecture and the unit distance conjecture.

If one wants to theorize that LLMs are particularly good at finding counterexamples, then there are two things it would be good to do to make the theory more convincing. The first may sound unproblematic: it is to decide when solving a problem counts as finding a counterexample. Once that is sorted out, the second is to come up with a potential explanation of why LLMs would be particularly well suited to solving problems of that particular kind.

What does it mean to find a counterexample?

Why am I suggesting that it is not completely obvious what it means to find a counterexample? Surely, one might suggest, all it means is that you have a statement of the form “Every object of such and such a type has such and such a property,” and you exhibit an object of the given type that does not have the given property.

However, this doesn’t always work. Consider a famous result of Vinogradov, which states that every sufficiently large positive integer is a sum of three primes. The negation of this statement is (or is equivalent to) the statement that for every positive integer N there exists an integer n\geq N such that n is not a sum of three primes. In other words, it states that every positive integer N has a certain property. Seen in this light, Vinogradov found an example of a positive integer N that does not have the given property. Do we want to say that Vinogradov found a counterexample? Clearly not — the result should obviously be classified as a theorem and not a counterexample.

Thus, we cannot just naively say that LLMs are particularly good at negating universally quantified statements: there has to be something about the nature of the universal quantification. With the three-primes example, it is clear that Vinogradov did not think, “How am I going to find N with this property?” Rather, what he thought would have been more like, “I’ve got an integer n that is very large. How am I going to show that it is a sum of three primes?” In other words, all his focus would have been on the universally quantified n, with the existentially quantified N being a sort of afterthought once the details of the proof have been worked out.

In general, many interesting results, when they are stated formally, begin with an alternation of two or three (or more) quantifiers. The question then becomes to determine which is the first “interesting” quantified variable in some sense. Here’s another example to illustrate the point, from the theory of finite-dimensional normed spaces. I’ll give a few mathematical details for those curious, but if you don’t care about those, then you can skip the next three paragraphs and should get the gist of what I am saying about this example.

Let X and Y be two n-dimensional normed spaces and let T be a linear map from X to Y. We say that T is a Cisomorphism if there exists \lambda>0 such that \lambda\|x\|\leq\|Tx\|\leq C\lambda\|x\| for every x\in X. By rescaling we can always take \lambda to be 1, in which case we have that \|x\|\leq\|Tx\|\leq C\|x\| for every x\in X. If C=1, then this tells us that T is an isometry. In general, the Banach-Mazur distance d(X,Y) between X and Y is defined to be the smallest C such that there exists a C-isomorphism from X to Y. It is easy to see that the logarithm of the Banach-Mazur distance is a metric on the set of isometry classes of n-dimensional normed spaces. A less easy fact, but still not too hard, is that the resulting metric space is compact: in fact, it is known as the Banach-Mazur compactum.

It is natural to wonder what the diameter of the Banach-Mazur compactum is, and here things get interesting. A result of Fritz John states that every n-dimensional space X has distance at most \sqrt n from \ell_2^n. (The idea of the proof is as follows: pick inside the unit ball of X an n-dimensional ellipsoid of maximal volume; that is the unit ball of a normed space Y that is isometric to \ell_2^n; it can be shown that the identity map is a \sqrt n-isomorphism between X and Y.) From Fritz John’s theorem and the (multiplicative) triangle inequality, it follows that d(X,Y)\leq n for any two n-dimensional normed spaces. That is, the diameter of the Banach-Mazur compactum is at most n. But might it be substantially less than that?

An indication that the answer is not obvious comes from looking at the spaces \ell_1^n and \ell_\infty^n. The identity map between these two spaces is an n-isomorphism, but one can do much better by mapping the standard basis vectors not to themselves but to vertices of the unit cube, with the vertices chosen to be as orthogonal as possible. In particular, if there exists an n\times n Hadamard matrix, then the corresponding linear map is a \sqrt n-isomorphism. One can push this observation and deduce that for any p,q\in[1,\infty] the Banach-Mazur distance between \ell_p^n and \ell_q^n is O(\sqrt n). It is also easy to show that d(\ell_1^n,\ell_2^n)=\sqrt n, so \ell_p-spaces hardly improve on the easy lower bound, and do not improve on it at all in dimensions n for which an n\times n Hadamard matrix exists.

In 1981, Gluskin famously solved the problem by determining the correct asymptotics for the diameter of the Banach-Mazur compactum. Informally, what he showed was that the diameter is within a constant of the upper bound that follows immediately from Fritz John’s theorem. If we make the quantification explicit, then the statement we end up with is

\exists c>0\ \forall n\ \exists X,Y\in K_n\ d(X,Y)\geq cn,

where I have written K_n for the set of all n-dimensional normed spaces. (If you want to argue that it is not a set, then let me specify in addition that the underlying vector space is \mathbb R^n.) In words, there is a positive constant c such that for every positive integer n there are n-dimensional normed spaces X and Y such that the Banach-Mazur distance between X and Y is at least cn.

I can’t continue without very briefly describing the beautiful and highly influential idea Gluskin had for solving this problem. He took X and Y to be normed spaces whose unit balls were random symmetric convex sets defined as follows: take the standard basis vectors and a handful of other random unit vectors, as well as the negatives of all these vectors, and take the convex hull. Gluskin then showed that if two normed spaces are chosen from this distribution, then with high probability their Banach-Mazur distance is at least cn.

But back to the main point, which is that the logical form of the above statement is very similar to the logical form of Vinogradov’s theorem, which is

\exists N\ \forall n\geq N\ \exists p_1,p_2,p_3\in P\ \ p_1+p_2+p_3=n

where I have written P for the set of primes. And yet, Vinogradov’s result is unquestionably a theorem, while Gluskin’s result is unquestionably a counterexample, or at least an example.

What is the important difference between the two statements? It seems to be that in Vinogradov’s three-primes theorem the number n plays a more essential role in the statement that is to be proved about the various quantified variables. In Vinogradov’s theorem, that statement is n=p_1+p_2+p_3, whereas for Gluskin’s theorem the statement to be proved is

\dim X = \dim Y = n and d(X,Y)\geq cn,

which we can write equivalently as

\dim X = \dim Y = n and d(X,Y)\geq c\dim X.

In the case of Vinogradov’s theorem, the whole challenge is to get those three primes to add up to n, whereas for Gluskin it is not remotely challenging to get the dimensions of X and Y to equal n: the challenge is to get X and Y to be very far from each other, relative to their common dimension.

There is a further complication to bear in mind here, which is that via the process known as Skolemization, a universally quantified statement of the form \forall x\in X\ \exists y\in Y\ \ P(x,y) can be converted into an existentially quantifed statement \exists f:X\to Y\ \forall x\in X\ \ P(x,f(x)). (For this to be an equivalence one needs the axiom of choice, but it is certainly a sufficient condition.) This is not just a piece of logical trickery, but it often reflects quite accurately how we think about some problems. For instance, it is more natural to think of Gluskin’s example as a recipe for constructing (or at least proving the existence of) a pair of suitable normed spaces for any given dimension n, or in other words to construct a suitable function from \mathbb N to pairs of normed spaces by giving its value at each n, than it is to think of it as a statement that says that every positive integer n has a certain complicated property.

Yet another complication is that some universally quantified statements follow naturally from existentially quantified statements, or may even be equivalent to them. For example, the theorem that a 2-dimensional torus is not homeomorphic to a 2-dimensional sphere is a universally quantified statement (every map from the torus to the sphere fails to be a homeomorphism), but the natural way to prove it is to prove the existential statement that there is an invariant that distinguishes the two spaces. For an example of where a universal statement is equivalent to an existential statement, consider a statement of the form that a vector x\in\mathbb R^n does not belong to the convex hull of a certain compact set A. The statement that no convex combination of elements of A is equal to x is equivalent to the existence of a linear functional \phi:\mathbb R^n\to\mathbb R and a \lambda\in\mathbb R such that \phi(x)>\lambda and \phi(a)\leq\lambda for every a\in A. In both these cases it feels natural to regard the result as a theorem that is proved via an existential statement, perhaps because it is the theorem that is ultimately what interests us. But using “what interests us” as a criterion to determine what counts as a counterexample seems a little vague, and is a difficult criterion to use if we want to explain convincingly why AI should be good at finding counterexamples.

A more general argument against the notion that there is something about existential statements that is particularly suited to AI is that the need to establish existential statements pervades almost all of mathematical research, regardless of the nature of the headline result being aimed for. For example, if I want to prove a statement by induction, I may well look for a strengthening of the statement that serves better as an inductive hypothesis. Or if I want to prove that every object of type T with property P also has property Q, then I may well look for a property R that follows from P and can be used to prove Q. These are more metamathematical existence problems, but the distinction can be somewhat blurred, and more importantly, when trying to prove a statement S, it is often the case that the main question in our minds is less, “Why is S true?” and more, “What could a proof of S be like?” To give an example, I feel I understand pretty well why Goldbach’s conjecture is true — a highly plausible probabilistic model of the primes implies it and agrees closely with computational data — but if I were making a serious attempt to prove it, that understanding, which many mathematicians have had for a century or so, would be of limited help. Rather, my main task would be to try to find proof techniques that were powerful enough to make those heuristic ideas rigorous.

What is the difference between an example and a counterexample?

Logically, every statement of the form \exists x\ P(x) is a counterexample to the universally quantified statement \forall x\ \neg P(x). However, we do not describe all existential statements as counterexamples. For example, if I were to say, “The \ell_p-spaces with 1\leq p<\infty are all separable, as is c_0, but \ell_\infty is not separable,” I would not describe the second part of that assertion as a counterexample to the claim that all Banach spaces are separable. Rather, I would present it as probably the most basic example of a non-separable space. The important point seems to be that there was no particular reason to think that all Banach spaces would be separable, and finding an example of a non-separable space is not very difficult.

I think the first point is more important here: we are more inclined to call an object a counterexample if the existence of that object disproves a statement that we had quite good reason to believe. It often happens that after repeated unsuccessful attempts to prove a statement, mathematicians begin to feel that it has no particular reason to be true, even if it seems to be hard to come up with a counterexample to it. In such a situation, if a counterexample is eventually found, it may have lost something of its “counter” feel. My impression is that the construction of a non-sofic group comes into this category. There have been several proposals in the literature for how one might construct such a group, and I don’t think there were many (or even any?) experts who strongly believed that all groups were sofic. So it feels more natural to say, “OpenAI came up with the first example of a non-sofic group” than to say, “OpenAI found a counterexample to the soficity conjecture” (despite the fact that that section of their paper is entitled “A counterexample to the soficity conjecture”).

Likewise, it seems to me that the new lower bound for multicolour Ramsey numbers is more of an example than a counterexample. I think quite a lot of people believed that the bound should be exponential, so for them it was a counterexample, but others, myself included, were more neutral about it. As a matter of fact, I have worked on the problem in the past (a long time ago) in an equivalent formulation, which asks how many triangle-free graphs on n vertices you need if you want their union to be the complete graph K_n. If you take bipartite graphs, then it’s easy to see that you need \log_2n of them, but that bound can be improved if instead you observe that a complete 5-partite graph can be written as a union of two triangle-free subgraphs, and therefore it is possible to write the complete graph as a union of 2\log_5n triangle-free graphs. It is then tempting to try to do better, with triangle-free graphs that are less dense but that make up for it with unbounded chromatic number — a necessary condition if one wishes to use a sublogarithmic number of graphs, which is equivalent to showing a superexponential lower bound for R(3,3,\dots,3). All this is to say that when I worked on the problem, my efforts were concentrated on what turned out to be the right direction, so for me OpenAI found an example of what I (weakly) expected, rather than a counterexample.

Where does this leave us?

I would like to find a coherent explanation of the conjunction of the following facts.

  1. The most notable mathematical results proved by LLMs have tended to be ones that we would classify as examples or counterexamples, where counterexamples are, broadly speaking, existence statements that disprove statements that we expected to be true.
  2. Many statements can be formulated as existence statements when we would usually think of them as universal statements, and vice versa, so what we consider to be an example depends on the mathematical context of a statement as well as its logical form.
  3. LLMs are pretty good at proving universal statements as well: it’s just that the strongest statements they have proved that we would think of as theorems have mainly not been at the level of the strongest statements that we would think of as counterexamples.

Given these facts, it seems likely that what LLMs are good at is something else, which happens to have as a consequence that they are good at the kind of existence problem that we would normally classify as asking to find a non-trivial example.

Let us consider two things that we can be confident that LLMs are good at. One of them is knowing a lot of mathematics: if a problem can be solved by means of a relatively standard argument, it is highly likely that an LLM will be able to find and use that argument. The other is the ability that an LLM has simply by virtue of being a computer: it can work at huge speed (compared with humans at least) and can therefore afford to make a large number of unsuccessful attempts at a problem before it finds a solution.

Without even looking at what LLMs have actually managed to solve, one might guess that these two features would lead to their having a somewhat different style from human mathematicians. Very roughly, LLMs would have the edge when there is more of a probabilistic element to the proof-finding process: they would be good at problems for which the best method is to try a lot of ideas, not necessarily particularly novel, until at some point you get lucky. Humans on the other hand would be better (for the moment) at finding more “surprising” and “conceptual” arguments, where the appropriate method is to dig deeper and deeper into a problem until the solution reveals itself. (It is hard to say exactly what this means, but I hope that any experienced researcher reading this will know what I am talking about.)

This raises two questions: does the guess above correspond at all to the reality that we are observing, and is there any reason to suppose that what I have tentatively described as the “LLM style” of doing mathematics would lead naturally to LLMs discovering several counterexamples (or just examples) to long-standing conjectures, even if that was by no means all they could do?

I don’t pretend to have a scientific answer to either question, but the reactions of experts to several of the remarkable solutions that ChatGPT has found do lend some support to the idea that LLMs work in more of a try-lots-of-things-till-you-get-lucky way. People often seem to react by saying something like, “Initially I was amazed that the problem had been solved, but on closer inspection I realized that the approach was actually not all that novel, and one that with the right small hint a suitably expert human could have found quite easily.”

For the second question — whether the LLM style is well suited to finding (counter)examples — I think matters are less clear, because there are many ways of searching for a counterexample, and some of them fit better than others the style I have described. Here are a few general methods. (I don’t claim that the list is exhaustive.)

  1. Look for an off-the-shelf example. Here one has a stock of fairly standard examples and one simply tries them out one after another to see whether any of them fails to satisfy the given statement. For example, Ryan O’Donnell ends his wonderful book on the analysis of Boolean functions with some tips, one of which is, “If you have a conjecture about Boolean functions, test it on dictators, majority, parity, tribes (and maybe recursive majority of 3). If it’s true for these functions, it’s probably true.”
  2. Build an example from basic examples and standard construction methods. For an algebraic problem, for instance, one might start with some standard examples, but then take products or quotients or limits.
  3. Make heavy use of metavariables. The word “metavariable” comes from computer science, and in particular from automatic theorem proving, and refers to the practice that in mathematics would correspond to writing, “where x is to be chosen later,” (in which case x is the metavariable). In a paper we usually do this only in fairly simple situations such as when we need to choose a number \epsilon>0 that is small enough for later arguments to work. But when we search for an example of an object x that satisfies some property Q (which may well be a conjunction of simpler properties Q_1,\dots,Q_k), it is often not a good strategy to specify x completely and only then to check whether it satisfies Q. Instead, it can be more fruitful to do almost the opposite: we start by saying virtually nothing about x and simply launch into proving that it satisfies Q. In the course of doing so, we find that we need x to satisfy a property P_1. If we are lucky we can describe in a nice way a very general class of objects x that satisfy P_1. For instance, we may be able to find a parametrized class: we identify some function f and show that f(y) satisfies P_1 for every y of a certain type. The problem is then reduced to finding y such that $Q(f(y))$ holds, which is a more specific version of the original problem. There may be many iterations of this process, or a mixture of this process and other processes, before an example is eventually found.
  4. Try to prove the opposite. If one wishes to find x such that Q(x), it can be surprisingly helpful to start by attempting to prove the statement \forall x\ \neg Q(x). The reason this can be helpful is that using our standard methods of attempting to prove something, we may end up identifying a key lemma that would suffice: that is, we may find an intermediate property R that implies \neg Q in a non-trivial way and thus reduce the problem \forall x\ \neg Q(x) to \forall x\ R(x). Turning things round again, it may well then be that finding a counterexample to R is easier than finding a counterexample to \neg Q (that is, an example that satisfies Q). Of course, there is no guarantee that a counterexample to R will be an example of Q, but sometimes we are lucky and it is. More often, we can use the idea of the previous method, noting that it is at least a necessary condition of an example of Q that it should not be an example of R, so one can try to describe a general class of objects that fail R and in that way reduce the problem.
  5. Successive approximation. Sometimes, when we are searching for an example of x such that Q(x), we write down a moderately plausible guess x_0 not because we think it has a chance of working (if we did, then we would be using the first strategy), but because we hope that if x_0 does not satisfy Q, then we will be able to diagnose what went wrong and specify a new guess x_1 that does not have that defect. Again, this strategy can either be iterated or combined with one or more of the other strategies.
  6. Just-do-it proofs. Sometimes we need x to satisfy infinitely many properties Q_1,Q_2,\dots, each of which is, individually, quite easy to satisfy. In such situations, we often “build” x inductively bit by bit, ensuring at the ith stage of the process that however the building process continues, x will satisfy Q_i.
  7. Pick a random example. Often it is very hard to give an explicit example of an x that satisfies Q, but there is a natural probability distribution for which one can show that if one chooses x randomly from that distribution, then with high probability (or at least non-zero probability) it will satisfy Q.
  8. Pick a generic example. In more infinite contexts, it may again be quite hard to give an explicit example of an x that satisfies Q, but one may be able to show that the set of x that fail Q is or measure zero, or is a meagre set, or is small in some other way.

There is no particular reason to suppose that LLMs would be equally good at each of the methods above. So perhaps what we are observing is not quite that LLMs have a particular ability to find examples, but more that they are particularly good at finding examples (and proofs) in a certain way. Looking at the above techniques, one might imagine that they would be very well suited to checking off-the-shelf examples, finding just-do-it proofs (since that is a rather standard method with lots of instances in their training data), using the probabilistic method (unless, as often happens, significant new ideas are needed to show that the probabilities work out), and picking generic examples. The other three methods described above — use of metavariables, trying to prove the opposite, and using successive approximation — require more of an ability to judge whether the approach one is taking is likely to be fruitful. Here it seems at least possible that humans will sometimes have an advantage, but the conditions that a problem would need to satisfy are quite stringent. One would need an example to be one that lies at a leaf of a very large search tree — too large to be searched for by a combination of moderate mathematical ability and brute force — but that can be found by a mathematician with a sufficiently good nose for when they are making progress that they can prune the search tree very substantially.

Why wouldn’t LLMs also have that “nose”? I don’t rule out that “nose” is an emergent property of the way LLMs are trained, and that within a year or two they will have it to the same extent that we have it. But for now, in my interactions with ChatGPT, I do have a distinct impression that they haven’t got there quite yet. When I discuss an open problem with 5.6 Pro, I am often presented with approaches that sound promising until I think about them carefully, and then seem quite a lot less promising. And they will also often end a response by saying, “I have not managed to answer the question you asked, but have managed to reduce it to the following much narrower and more precise question,” which sounds very promising until it has happened five times without any obvious progress having been made. It isn’t completely obvious how they will get better at this, since their training data will not be full of examples of fruitful and less fruitful directions to pursue when trying to solve problems: all they will typically see is tidied up proofs that hide the thought processes of their discoverers. Of course, human mathematicians also don’t get to learn much about how to do research from the experience of other mathematicians, and yet we somehow manage to pick it up. But the situation is a little different for us, in that a lot of what we learn is by doing rather than emulating.

Another reason it is not obvious that “nose” is a property that emerges naturally when LLMs are scaled up is that if LLMs make heavy use of their broad knowledge and can afford to do a lot more brute-force search than humans can, then they will lack the incentive that humans have to prune the search tree ruthlessly. It could conceivably be that their successes so far are achieved using methods that for a human would be considered extremely inefficient, but that because of their superior speed and knowledge, the combinatorial explosion these methods will lead to has not yet become apparent.

It would be very interesting to try to test this experimentally, but it is also difficult, because if an LLM has what looks like the kind of idea that could only be the result of “deep thought” about a problem, we can never be sure that it has actually carried out that deep thought, as opposed to finding a model argument already in the literature, or in other words exploiting the deep thought of a human mathematician. It would probably be easier (but still not easy) to test it by using models that are less powerful than the latest ones and that have been to some extent shielded from the mathematical literature: one could give them a carefully designed suite of problems and see whether the ones that the LLMs solve have particular characteristics.

It may seem as though I am desperately clinging to the hope that humans will continue to be able to make meaningful contributions to mathematical discovery for a while yet, but while I do indeed hope that, I am not making any assertions of the form “LLMs will never be able to do X”. I think it is likely that they will, and given the pace of progress over the last three years it will probably happen quite soon. But I do think that there may be a hurdle for LLMs to clear and it seems at least possible that it won’t be cleared as straightforwardly as some of the previous hurdles.

In that connection, it would also be interesting to see whether a different reward structure leads to LLMs being able to solve different kinds of problems. For example, if during training an LLM (or machine-learning system of some other kind) is not just rewarded if it ends up with a solution, but also penalized if it explores too many dead ends or if it “cheats” by getting the answer from the literature, perhaps it would be incentivized to go about the research process in a more human way and thereby achieve better results for classes of problems where it is yet to make a big impact.

If the hurdle is cleared, either by pure scaling up or by some more thoughtful method, it will be quite difficult to know when that has happened, since, as just mentioned, an idea that seems very original and surprising may just be lurking somewhere in an LLM’s training data. But I would be confident that it had been cleared if an LLM were to come up with a proof that was as surprising to me as the solution of the cap-set problem was in 2016: the previous best known bounds were completely eclipsed, the method was utterly different from anything I had thought about trying, and afterwards there was a flurry of activity as people came to understand what this wonderful new technique was capable of.

Conclusion

I wasn’t quite sure where I would end up when I started this post, and now that I’ve got to the end, I feel that my main conclusions are not particularly new or surprising, but I hope that the route to them is of some interest. The main points I have made are the following.

  1. “Finding an example” is in practice not the same thing as proving a statement that begins with an existential quantifier.
  2. If it is true that current models are particularly good at finding examples, that is probably not because they have a particular affinity for existential statements, but more because the proof-discovery methods that are appropriate for finding certain kinds of examples play to the obvious strengths of LLMs: wide knowledge and the ability to explore many paths of the search tree that humans would judge to have a low probability of success.
  3. It seems likely that LLMs will carry on improving very quickly. However, if, contrary to expectations (mine at least), there turns out to be some residual class of problems (or other mathematical activities) for which humans continue to have the edge for a while, it is likely that those will be problems for which the mysterious human ability to prune the proof-discovery search tree is particularly advantageous: that is to say, problems where the search tree is deep and has a large amount of branching, so that without rigorous pruning a search is not feasible even for a computer.
  4. A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it.

August 12, 2026

John BaezJordan Triples and the Standard Model

I don’t usually talk about particle physics here. I have a whole series of articles about octonions and the Standard Model on my other blog. But I’m kind of excited about this new paper, so I’ll talk about it here too:

• John Baez, Endre Bokor and Latham Boyle, Jordan pair quantum theory and the Standard Model.

Jordan algebras were introduced by Jordan, von Neumann and Wigner in 1934 in an attempt to formalize algebras of observables in quantum theory. They come in 4 infinite series—but there’s one more, the ‘exceptional Jordan algebra’, consisting of 3 × 3 self-adjoint matrices of octonions. For years physicists sought to find some use for it.

In 2018, Todorov and Dubois–Violette noticed that the symmetries of the exceptional Jordan include the Standard Model gauge group in a nice way. But it was unclear how to bring in the fermions—the quarks and leptons. That’s what our new paper does.

To do this, we need to go beyond Jordan algebras. Jordan pairs and Jordan triples are two closely linked formalisms that generalize Jordan algebras. Our paper explains them in detail—and how they’re connected to geometry and quantum mechanics. But here I will mostly skip that wonderful story, so I can quickly explain the connection to the Standard Model.

Here’s how the Standard Model gauge group, together with its representation on one generation of fermions, drops out of a Jordan triple.

The bi-Cayley triple

Let

\mathbb{O}_\mathbb{C} = \mathbb{C} \textstyle{\otimes}_\mathbb{R} \mathbb{O}

be the bioctonions: octonions with complex coefficients. Write \mathbb{O}_\mathbb{C}^2 for the space of column vectors with two bioctonion entries.

\mathbb{O}_\mathbb{C}^2 has a certain triple product

[x,y,z]=\frac{1}{2}(x(y^{\dagger}z)+z(y^{\dagger}x))

which obey the axioms of a gadget called a ‘positive hermitian Jordan triple’. It’s called the bi-Cayley triple.

Now, every positive hermitian Jordan triple gives rise to a \mathbb{Z}_2-graded real Lie algebra

\mathbf{k} = \mathbf{k}_0 \textstyle{\oplus} \mathbf{k}_1

Not a Lie superalgebra: a plain old-fashioned Lie algebra with a \mathbb{Z}_2-grading!

How does this work? We take the hermitian Jordan triple itself to be \mathbf{k}_1. The Lie algebra \mathbf{k}_0 consists of all linear maps from \mathbf{k}_1 to itself that are of this form:

x \mapsto [a,b,x] - [b,a,x]

for some a,b \in \mathbf{k}_1. These maps are called real inner derivations. They form a Lie algebra since the commutator of two such maps is another such map. With a bit more work we can define other operations making all of \mathbf{k} into a \mathbb{Z}_2-graded Lie algebra.

So, we get a big Lie algebra \mathbf{k}, and a Lie subalgebra \mathbf{k}_0 sitting inside it. From this we get two Lie groups: a big one K whose Lie algebra is \mathbf{k}, and a subgroup K_0 whose Lie algebra is \mathbf{k}_0.

The quotient is K/K_0 is a nice kind of manifold called a hermitian symmetric space. Conversely, any compact hermitian symmetric space give rise to a positive hermitian Jordan triple!

This geometric picture is revealing. The group K acts transitively as symmetries of our hermitian symmetric space, while the stabilizer of any point is isomorphic to K_0. Our original Jordan triple, \mathbf{k}_1, is then the tangent space of that point! So, K_0 acts on this Jordan triple. This action preserves the triple product, and we call K_0 the real inner automorphism group of our Jordan triple.

Here’s another great thing about the geometric picture: hermitian symmetric spaces were classified by Eli Cartan (who seems to have spent his life classifying things). As a result we also know the classification of positive hermitian Jordan triples. They come in four infinite series together with two exceptions. One is the bi-Cayley triple, and other is the Albert triple, which is the complexification of the exceptional Jordan algebra. The bi-Cayley triple is a subtriple of the Albert triple. It’s these two exceptions that are connected to the Standard Model. But we’ll start with the bi-Cayley triple.

The 3-graded Lie algebra coming from the bi-Cayley triple is the compact real form of \mathfrak{e}_6:

\mathfrak{e}_6 = \big[\mathfrak{so}(10) \textstyle{\oplus} \mathfrak{u}(1)\big] \textstyle{\oplus} \mathbb{O}_\mathbb{C}^2

The even part of this Lie algebra is in brackets. The corresponding hermitian symmetric space is called the bioctonionic plane (\mathbb{C}\otimes\mathbb{O})P^2. The even part of our 3-graded Lie algebra, \mathfrak{so}(10)\oplus \mathfrak{u}(1), generates the stabilizer of a point in the bioctonionic plane. The odd part, our friend \mathbb{O}_\mathbb{C}^2, is the tangent space of that point.

Here’s the first big surprise. The even part transforms as the adjoint representation of \mathrm{Spin}(10), while the odd part itself transforms as the 16-dimensional complex spinor representation of \mathrm{Spin}(10). Ignoring the extra \mathrm{U}(1) for a moment, this is exactly what we see in a \mathrm{SO}(10) grand unified theory: gauge bosons in the adjoint representation, and one generation of fermions in the 16-dimensional spinor representation.

So before we do anything, the bi-Cayley triple already smells like it contains the ingredients of an \mathrm{SO}(10) grand unified theory.

Tripotents

In a Jordan algebra the important elements are the idempotents, e^2 = e. In a Jordan triple W their role is played by tripotents: elements e with

[e,e,e] = e

A tripotent always lets us split W into three parts via something called its Peirce decomposition. The operator w \mapsto [e,e,w] has eigenvalues 0, 1/2, and 1, so W splits into the corresponding eigenspaces

W = W_0(e) \textstyle{\oplus} W_{1/2}(e) \textstyle{\oplus} W_1(e)

which are called the Peirce 0-space, Peirce 1/2-space and Peirce 1-space of e. A tripotent is called minimal when its Peirce 1-space is one-dimensional. Two tripotents e_1, e_2 are called colinear when each lies in the other’s Peirce 1/2-space.

I can’t resist explaining some of the quantum physics here. In a hermitian Jordan triple, the triple product [-,-,-] is linear in the first and last slot, but conjugate-linear in the middle slot. So, if you multiply a tripotent by a phase \alpha, you get a new tripotent:

[\alpha e, \alpha e, \alpha e] = \alpha \overline{\alpha} \alpha e = \alpha e

This should remind you of how when you multiply a unit vector in a Hilbert space by a phase, you get a new unit vector. In Jordan triple quantum mechanics, minimal tripotents take the place of these unit vectors. The hermitian symmetric space K/K_0 that I was talking about earlier is the same as the space of minimal tripotents mod phase! So, it generalizes the familiar space of ‘pure states’ in quantum mechanics: unit vectors mod phase.

But let’s get back to the Standard Model.

A chain of Jordan triples

From here on, the single fact driving everything is this: in any hermitian Jordan triple, any minimal tripotent’s Peirce 1/2-space is itself a hermitian Jordan triple!

If we run this starting from the bi-Cayley triple, we get this chain of hermitian Jordan triples, where each row’s 1/2-space is the next row’s triple:

Jordan triple Lie algebra \mathbf{k}_0 \oplus \mathbf{k}_1 (even part in brackets)
W = \mathbb{O}_\mathbb{C}^2 \mathfrak{e}_6 = [\mathfrak{so}(10) \oplus \mathfrak{u}(1)] \oplus \mathbb{O}_\mathbb{C}^2
W' = \mathfrak{a}_5(\mathbb{C}) \mathfrak{so}(10) = [\mathfrak{su}(5) \oplus \mathfrak{u}(1)] \oplus \mathfrak{a}_5(\mathbb{C})
W'' = \mathrm{M}_{3,2}(\mathbb{C}) \mathfrak{su}(5) = [\mathfrak{g}_{\mathrm{SM}}] \oplus \mathrm{M}_{3,2}(\mathbb{C})

Here \mathfrak{a}_5(\mathbb{C}) is the Jordan triple of antisymmetric 5\times 5 complex matrices, \mathrm{M}_{3,2}(\mathbb{C}) is the Jordan triple of 3\times 2 complex matrices, \mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3)\oplus\mathfrak{su}(2)\oplus \mathfrak{u}(1), and

G_{\mathrm{SM}} = \mathrm{S}(\mathrm{U}(2) \times \mathrm{U}(3)) \cong (\mathrm{SU}(3)\times\mathrm{SU}(2)\times\mathrm{U}(1))/\mathbb{Z}_6

is the true Standard Model gauge group.

The gauge group from two tripotents

Start with the bi-Cayley triple. Choose two colinear minimal tripotents e_1, e_2. Descend the table twice:

• Start with W = \mathbb{O}_\mathbb{C}^2, which has real inner automorphism group (\mathrm{Spin}(10)\times\mathrm{U}(1))/\mathbb{Z}_4.

• Fix e_1. Its Peirce 1/2-space is latex W’ = \mathfrak{a}_5(\mathbb{C}),$ with real inner automorphism group \mathrm{SU}(5)\times\mathrm{U}(1).

• Fix e_2 (colinear with e_1, so living in W'). Its Peirce 1/2-space in latex W’$ is W'' = \mathrm{M}_{3,2}(\mathbb{C}), with real inner automorphism group exactly G_{\mathrm{SM}}.

In other words, the subspace of the bi-Cayley triple colinear with both e_1 and e_2 is a Jordan triple whose real inner automorphism group is the Standard Model gauge group.

The choice of e_1 and e_2 also pins down how G_{\mathrm{SM}} sits inside the original group \mathrm{E}_6. At each we step take the subgroup that acts with determinant 1 and preserves the chosen tripotent up to a phase; this gives a chain of subgroups whose members are \mathrm{Spin}(10), \mathrm{U}(5), and G_{\mathrm{SM}}, so we get the embeddings

G_{\mathrm{SM}} \subset \mathrm{SU}(5) \subset \mathrm{Spin}(10)

In particle physics, this is the classic chain taking us from the so-called \mathrm{SO}(10) grand unified theory down to the \mathrm{SU}(5) grand unified theory down to the Standard Model. And it’s well known that restricting the 16-dimensional complex spinor representation of \mathrm{Spin}(10) along this chain gives precisely the Standard Model representation \rho_{\mathrm{SM}} on one generation of fermions! So we get one generation of Standard Model fermions this way.

The six particles types as Peirce spaces

We have gotten the representation of the Standard Model gauge group on one generation of fermions without any fuss. But it’s also fun to peer into the details, and see how the different kinds of fermions emerge.

For any tripotent e, we have projections P_0(e), P_{1/2}(e) and P_1(e) onto its three eigenspaces: its so-called Peirce projectors. Since we get the Standard Model structure using two minimal tripotents e_1 and e_2 in the bi-Cayley triple \mathbb{O}_{\mathbb{C}}^2, there are nine composites of two Peirce projectors we can apply to this triple. This is how we pick out the different kinds of fermions!

As a representation of the Standard Model Lie algebra

\mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3) \textstyle{\oplus} \mathfrak{su}(2) \textstyle{\oplus} \mathfrak{u}(1)

any generation of Standard Model fermions transforms as the direct sum of six irreducible representations:

\rho_{\mathrm{SM}} = (3,2,\tfrac{1}{6}) \textstyle{\oplus} (\bar 3,1,\tfrac{1}{3}) \textstyle{\oplus} (\bar 3,1,-\tfrac{2}{3}) \textstyle{\oplus} (1,2,-\tfrac{1}{2}) \textstyle{\oplus} (1,1,1) \textstyle{\oplus} (1,1,0)

These correspond to the six types of left-handed fermion: q_L, \overline{d_R}, \overline{u_R}, \ell_L, \overline{e_R}, \overline{\nu_R}. Six irreducible pieces, six particle types.

It turns out these are exactly the six nonzero components of the Peirce decomposition of \mathbb{O}_\mathbb{C}^2 with respect to both e_1 and e_2. Those six match up one-to-one with the particle types:

Peirce projector representation of G_{\text{SM}} particle type
P_{1/2}(e_2) P_{1/2}(e_1) (3, 2, +1/6) q_L
P_{1/2}(e_2) P_0(e_1) (\overline{3}, 1, +1/3) \overline{d_R}
P_0(e_2) P_{1/2}(e_1) (\overline{3}, 1, −2/3) \overline{u_R}
P_0(e_2) P_0(e_1) (1, 2, −1/2) \ell_L
P_1(e_2) P_{1/2}(e_1) (1, 1, +1) \overline{e_R}
P_{1/2}(e_2) P_1(e_1) (1, 1, 0) \overline{\nu_R}

The remaining three combinations—P_1(e_2)P_1(e_1), P_1(e_2)P_0(e_1), and P_0(e_2)P_1(e_1)—all vanish, which is why we land on six pieces and not nine.

So the whole package—the gauge group G_{\mathrm{SM}}, the embedding G_{\mathrm{SM}} \subset \mathrm{Spin}(10), the representation \rho_{\mathrm{SM}}, and even the split of one generation into its six particle multiplets as distinct Peirce components—all comes out of the single object \mathbb{O}_\mathbb{C}^2 once you choose two colinear minimal tripotents.

And if you prefer to start one level up, with the Albert triple \mathfrak{h}_3(\mathbb{O}) \otimes \mathbb{C}, you get the same result by choosing three mutually colinear tripotents instead of two—but for that, read our paper!

August 11, 2026

Tommaso DorigoFrom Inspiration to Impact: 10 Years of Research on AI for Physics

From Inspiration to Impact: 10 Years of Research on AI for Physics

A graph tells a thousand words - in this one, I present a summary of my past 10 years of research, trying to exploit the new AI technologies to improve the way we do research in fundamental science.

Tommaso Dorigo
Categories

August 10, 2026

Terence TaoA partial digestion of the HRT counterexample

A function {f(t) \in L^2({\bf R})} of one variable can be translated in space by a spatial shift {x} to obtain a new function

\displaystyle  \pi(x,0) f(t) := f(t-x),

and also modulated in frequency by a frequency shift {\omega} to obtain a new function

\displaystyle  \pi(0,\omega) f(t) := e^{2\pi i \omega t} f(t).

One can compose these two operations to obtain a time-frequency shift:

\displaystyle  \pi(x,\omega) f(t) := \pi(0, \omega) \pi(x,0) f(t) = e^{2\pi i \omega t} f(t-x).

As per the time-frequency uncertainty principle, the two shifts do not quite commute with each other. For instance, we have

\displaystyle  \pi(x, 0) \pi(0, \omega) = e^{-2\pi i \omega x} \pi(0, \omega) \pi(x, 0). \ \ \ \ \ (1)

As such, {\pi} is not a representation of the abelian group {{\bf R}^2}, but rather a portion of the Weyl representation of the Heisenberg group, but we will not adopt a representation-theoretic perspective here.

Some functions obey finite linear relations between their time-frequency shifts. For instance, a sinusoid {f(t) = A \sin(k t + \phi)} obeys the relation

\displaystyle  \pi(\frac{\pi}{k},0) f + f = 0.

However, the Heil-Ramanathan-Topiwala (HRT) conjecture states that once one imposes some reasonable decay condition on {f}, no such relations exist:

Conjecture 1 (HRT conjecture) If {f \in L^2({\bf R})} is non-zero, then there is no relation of the form

\displaystyle  c_1 \pi(z_1) f + \dots + c_n \pi(z_n) f = 0 \ \ \ \ \ (2)

for some distinct time-frequency shifts {z_1,\dots,z_n \in {\bf R}^2} and some coefficients {c_1,\dots,c_n \in {\bf C}}, not all zero.

A special case of the HRT conjecture, which was also open, makes the additional assumption that {f} was Schwartz.

Many positive results towards this conjecture were known. I will mention only a few here. A simple case is when we only have frequency shifts rather than spatial shifts:

\displaystyle  c_1 \pi(0, \omega_1) f + \dots + c_n \pi(0, \omega_n) f = 0.

In this case, the operator {(c_1 \pi(0, \omega_1) + \dots + c_n \pi(0, \omega_n))} is simply a physical space multiplier

\displaystyle  c_1 \pi(0, \omega_1) f(t) + \dots + c_n \pi(0, \omega_n) f(t) = B(t) f(t)

with symbol

\displaystyle  B(t) = c_1 e^{2\pi i \omega_1 t} + \dots + c_n e^{2\pi i \omega_n t}

so that the relation (2) now takes the simple pointwise form

\displaystyle  B(t) f(t) = 0. \ \ \ \ \ (3)

If the coefficients {c_1,\dots,c_n} are not all zero, then {B} is a non-zero analytic function and thus has isolated zeroes. It is thus not possible to solve this equation for any {f \in L^2({\bf R})}. Thus the HRT conjecture is true when all the time-frequency shifts {z_1,\dots,z_n} lie on the vertical axis. Using the metaplectic representation, one can then handle the case when all the {z_1,\dots,z_n} are collinear.

What about the non-collinear case? Suppose first that all the {z_j} lie in the lattice {{\bf Z}^2}, thus {z_j = (k_j, l_j)} for some integers {k_j, l_j}. Here, the phase shift in (1) disappears, and all the time-frequency shifts {\pi(k_j, l_j) = \pi(k_j,0) \pi(0,l_j)} commute with each other. This suggests that it should be possible to diagonalize the situation with a suitable transform to convert (2) to a pointwise equation similar to (3). To find this diagonalization, observe that if one restricts the function {f : {\bf R} \rightarrow {\bf C}} to a coset {s + {\bf Z}} of the integers, then {\pi(0,l_j)} just multiplies the function by the scalar {2^{2\pi i l_j s}}, while {\pi(k_j, 0)} shifts the function on this coset by {k_j}. The latter translation operation can also be converted to pointwise multiplication by performing the Fourier transform on the integers. Thus, if one introduces the Zak transform

\displaystyle  Zf(s,\nu) := \sum_{k \in {\bf Z}} f(s+k) e^{2\pi i k \nu}

of {f} then the equation

\displaystyle  c_1 \pi(k_1, l_1) f + \dots + c_n \pi(k_n, l_n) f = 0

can be transformed after a brief calculation to the equation

\displaystyle  B(s,\nu) Zf(s,\nu) = 0

where the symbol {B(s,\nu)} is now given by the formula

\displaystyle  B(s,\nu) = c_1 e^{2\pi i (k_1 \nu + l_1 s)} + \dots + c_n e^{2\pi i (k_n \nu + l_n s)}.

As before, if the coefficients {c_1,\dots,c_n} are not all zero, then {B} is a non-zero analytic function and thus non-zero almost everywhere. Thus {Zf} has to vanish almost everywhere, which for {f \in L^2({\bf R})} can be used to show that {f} also vanishes.

More generally, there is a result of Linnell that the conjecture is true if {z_1,\dots,z_n} lie in a translate of a discrete subgroup of {{\bf R}^2}; this (together with the argument handling the collinear case) establishes all cases where {n \leq 3}, and several partial results involving the {n=4} cases are also known. The conjecture is also known if {f} is decays at a suitably super-exponential rate, by work of Bownik and Speegle.

I was aware of this conjecture through various talks and conversations with colleagues, and even briefly tried my hand at it for a while, though not with particularly serious effort (or progress). It was thus a nice surprise to see that it has just been resolved by Faulhuber, Petersen, van Velthoven, and Voigtlaender, even in the Schwartz case:

Theorem 2 There exist complex numbers {c_1,\dots,c_{12}\in {\bf C}}, not all zero, distinct points {z_1,\dots,z_{12} \in {\bf R}^2}, and a non-zero Schwartz function {f_* \in \mathcal{S}({\bf R})} such that

\displaystyle  c_1 \pi(z_1) f_* + \dots + c_{12} \pi(z_{12}) f_* = 0.

It is perhaps unsurprising that this result is AI-assisted. However, I think the authors have disclosed their AI use responsibly, with the final arguments written by hand with a readable overview of the argument, as well as proper discussion of methods, relation to past literature, and other independent numerical checks on the result.

The negative result lies only a little beyond the positive results: {n} is now increased to {12}, and all but one of the points {z_1,\dots,z_{12}} lie in (a translate of) a discrete subgroup of {{\bf R}^2} (in fact the explicit subgroup {{\bf Z} \times \frac{1}{2}{\bf Z}} is used). The functions constructed are smooth and rapidly decaying, but not analytic or super-exponentially decaying, which would start being in conflict with the known positive results.

In addition to AI being used to come up with the initial proof strategy, a more traditional numerical computation was used to verify one step of the argument.

I have not had the time to do a full digestion of the result, but (after reading the introduction, and using a little AI assistance of my own) I was able to understand the main ideas at a high level. The first few reductions are relatively standard. Setting {z_{12} = 0} and {c_{12} = -c_*}, one can view the problem as one of solving an eigenvalue problem

\displaystyle  (c_1 \pi(z_1) + \dots + c_{11} \pi(z_{11})) f_* = c_* f_*.

The time-frequency shifts {z_1,\dots,z_{11}} are chosen to lie in a translate of the discrete subgroup {{\bf Z} \times \frac{1}{2}{\bf Z}} by a certain irrational shift {(\alpha, \beta/2)}. As mentioned previously, if in the shifts of a standard lattice {{\bf Z} \times {\bf Z}}, it would be natural to work with the Zak transform of {f_*}, but it turns out that the approach does not quite work when doing this for topological reasons (relating to the fact that scalar quasiperiodic functions of mean zero are forced to have zeroes), and so the authors used the slightly denser lattice instead {{\bf Z} \times \frac{1}{2}{\bf Z}}, which relates to a vector-valued version of the Zak transform taking values in {{\bf C}^2} rather than {{\bf C}}. Here, the phase shift in (1) does not completely disappear, but becomes a sign change. This slight loss of abelianness means that we cannot hope to diagonalize the problem all the way to a scalar problem, but we can still hope to reduce it to a two-dimensional vector-valued problem. Indeed, by applying a suitable vector-valued version of the Zak transform, the eigenvalue problem can be transformed to a a “vector cocycle problem”

\displaystyle  B_*(z) F_*(z - \tau) = c_* F_*(z), \ \ \ \ \ (4)

where {F_* : {\bf R}^2 \rightarrow {\bf C}^2} is a non-zero smooth quasiperiodic vector-valued function, {\tau} is an irrational shift {\tau= (\alpha,\beta)}, and {B_*(z)} is a certain explicit {2 \times 2} matrix-valued function depending on the choices of {c_1,\dots,c_{11}}, {z_1,\dots,z_{11}}, and {\tau}.

How to solve this equation? The motivating scenario here is if the matrix function {B_*(z)} was replaced by a rank one function

\displaystyle  B_0(z) = \chi(z) \chi(z-\tau)^*

for some smooth vector-valued function {\chi : {\bf R}^2 \rightarrow {\bf C}} of unit magnitude. Then one could solve the equation by taking {F_*(z) = \chi(z)} and {c_* = 1}. It is not possible to make the function {B_*} exactly of this form, but through some numerical computation and clever AI-assisted guesswork, the authors were able to find a choice of {c_1,\dots,c_{11}} and {z_1,\dots,z_{11}}, and {\tau} that made {B_*} approximately equal to a rank one function {B_0} of this form, in fact getting a uniform estimate

\displaystyle  \sup_{z \in {\bf R}^2} \| B_*(z) - B_0(z) \|_{op} < \frac{1}{3}.

As it turns out, such an approximation is sufficient to run a contraction mapping argument to find a solution to a variant of (4), namely

\displaystyle  B_*(z) v_*(z - \tau) = q_*(z) v_*(z)

for some smooth {v_* : {\bf R}^2 \rightarrow {\bf C}^2} and {q_* : {\bf R}^2 \rightarrow {\bf C}}. (Here it was important to get the operator norm bound below {\frac{1}{3}}; they are barely able to do this, with a numerically obtained bound of {0.333032}, though this bound might not be optimal.)

The main remaining obstacle is that the “eigenvalue function” {q_*(z)} is varying in the parameter {z} rather than constant. (This issue was, by the way, anticipated to some extent in previous work of Demeter, who observed that eigenfunctions of the almost Matthieu discrete Schrödinger operator gave a near-miss counterexample to the HRT conjecture, but with an eigenvalue that depended on an auxiliary phase shift parameter rather than constant.) However, if one was able to solve the scalar cocycle equation

\displaystyle  q_*(z) h(z-\tau) = c_* h(z) \ \ \ \ \ (5)

for some smooth {h : {\bf R}^2 \rightarrow {\bf C}}, then one could solve the equation (4) by setting {F_*(z) = h(z) v_*(z)}. The approach to solve (5) is standard: take logarithms, apply a Fourier transform, and then divide out by the multiplier associated to the {\tau} shift. This can cause a well-known “small divisor” problem (which arises in various dynamical contexts, such as in the KAM theorem) if {\tau} behaves too much like a rational vector, but the standard resolution to this is to select a shift {\tau} that obeys good Diophantine approximation properties. For the purposes of numerics the authors selected an extremely concrete shift, namely

\displaystyle  \tau = (2^{1/3} - 1, 2^{2/3} - 1)

but I get the impression that the exact choice here was not crucial for the argument, and that many other irrational algebraic numbers could have worked here.

Doug NatelsonReproducibility in materials research, and an anecdote

Yesterday I attended the 40th annual summer research colloquium of the Smalley-Curl Institute at Rice, a fun internal conference that provides a great opportunity for undergrads (including visitors), graduate students, and a few postdocs to present their work.  The keynote speaker was our EVPR, Prof. David Sholl, who gave a very informative talk about reproducibility in the chemical engineering/materials literature.  We hear a lot these days about crises of reproducibility in scientific research, and Prof. Sholl rightly points out that in some fields the expectation of reproducible results is high - no one would spend $1B on a chemical engineering plant if they weren't very sure that the catalytic processes were going to work as expected at scale.  Keys to reproducibility include, unsurprisingly, repeated results and independent replication.  One metaresult that was interesting is this paper, looking at the literature on metal-organic frameworks and how often there are published replications of syntheses; not as often as you would think or want!  

A truly surprising (to me, anyway) result is this one.  The Brunauer–Emmett–Teller (BET) (yes, that Teller) method is a long-established technique that uses gas adsorption measurements to infer the surface area of porous materials.  Many research groups were given identical raw adsorption isotherms and asked to calculate the specific surface areas, resulting in a surprisingly large spread of results (Fig 1 of the paper).  Clearly not everyone had the same analysis procedures even for a technique developed in the 1930s!

Some take-away lessons from this are encapsulated here, in an article titled "Five easy ways to make your research more reproducible".  Good stuff.  The talk raised a number of questions relevant to our present era of huge enthusiasm about AI-based materials research and "self-driving" labs.  If the AI models are all trained on the literature, and the literature is not representative of complete and reproducible procedures, that's a problem.  

One personal anecdote about reproducibility and its challenges in materials synthesis.  Twenty years ago (!), I was working with a colleague who had a postdoc who was synthesizing Fe3O4 (magnetite) nanoparticles via wet chemistry methods (see here). We did some fun electronic transport experiments bridging very closely spaced electrodes with such nanoparticles, and we saw some very dramatic hysteretic response kick in as \(T\) was reduced below about 120 K.  That's the temperature of the Verwey transition in magnetite, where the material enters a more insulating low temperature phase.  Basically all of the devices we made with that batch of nanoparticles showed this phenomenon.  Then the postdoc took up a faculty position and a senior grad student came in and took over the synthesis, and for several months, subsequent batches of nanoparticles just didn't seem to show the effect.  The key issue is oxygen stoichiometry.  Get a little oxygen rich, and you form nanoparticles that include some \(\gamma\)-Fe2O3, which doesn't have the Verwey physics and in nanoparticle form looks really similar in x-ray diffraction to the desired magnetite.  Anyway, we started working with a collaborator who could grow epitaxial Fe3O4 films, and in those devices the electronic effect was there all the time.  All this led to this publication and subsequent papers, and I still think it's a cool set of results about a nonequilibrium transition in a correlated material.  In the end, after several months the chemistry grad student did get back to making nanoparticle batches that showed the transition. It turns out that at some point he had changed the length of a piece of tubing in the gas manifold, and unexpectedly that had altered the reaction kinetics just a little.  Changing it back got the synthesis to be reliable again.  This is an example of how finicky materials synthesis can be!

John PreskillInteracting collaborators reveal noninteracting fermions

By day, I work as an experimentalist on laser-cooling molecules1, but I’ve never fully surrendered my theoretical-physics license. I started as an undergraduate in Lincoln Carr’s group at the Colorado School of Mines in Golden, CO. I learned from his expertise in simulations and complex systems. Since then I’ve moonlighted as a theorist while also pursuing an unrelated PhD and, now, an unrelated postdoc position. With Nicole Yunger Halpern and other collaborators, we devised a quantum circuit whose dynamics looked complex when run on a quantum computer. It took six years and five collaborators across four countries to discover that, for the right settings, these complex dynamics could be understood when viewed from the right angle.

Some time ago, I told you about quantum cellular automata (QCA). These quantum machines are built from one-dimensional strings of qubits. A qubit changes its state depending on the state of its two nearest neighbors. Different rules are encoded into three-qubit gates that change a central qubit based on the state of its left and right neighbors. Some rules induce change for many combinations of neighbor states. Others, less. We apply this neighborhood-constrained update in two waves, first to every other qubit, then to the ones skipped in the first wave. This is a common quantum circuit structure called a brickwork pattern. We call one rule the Goldilocks QCA: A qubit is updated if one of its neighbors is a 0 while the other is a 1 (activity); otherwise the qubit does not change its state (inactivity).

The first figure from our recent paper illustrating the Goldilocks QCA brickwork circuit. Orange boxes represent unitary gates. Half-white-half-black circles represent the Goldilocks neighborhood constraint. Some choices for the unitary gate result in free fermion dynamics. Most choices are consistent with chaos.

Repeating brickwork layers of the Goldilocks rule, we found, balances activity and inactivity to be “just right,” as Goldilocks might say. Striking this balance produced surprisingly rich patterns of quantum correlation. The same type of network structure is found in complex classical systems like metabolic pathways, social networks, and brain activity. What’s more, the observed patterns of connectivity persist through thousands of circuit layers while other QCA tend towards uniformity.

Goldilocks in a state of activity. Published by The Grolier Society, 1912

Our new paper, Integrability of Goldilocks quantum cellular automata, answers a question that’s been lurking underneath that first result for the last several years. Why does this balance produce such rich and persistent structure? Some Goldilocks QCA, we prove, map onto free fermions, one of the simplest examples of exactly solvable quantum dynamics. How does uncovering this simplification explain the persistent complex patterns? The answer follows from the concept of conservation laws. Piecing together this understanding required assembling an international team of experts who generously shared their knowledge and time. I’ll tell a bit of this scientific story through the lens of our collaboration’s history.

A key inspiration for this work started with a May 2020 video call with Norman Margolus, an MIT-affiliated researcher and pioneer of using cellular automata to model real systems. In the 1980s he worked on a custom computer chip called CAM-6, and later CAM-8, that was dedicated to simulating massive arrays of cellular automata with the limited computational resources of the era2. He proudly showed us beautiful pictures of cellular automata simulating phenomena like optical refraction and chemical reactions.

Cellular automata book by Norman Margolus. His coauthor’s name may also be familiar to those with quantum-circuit experience. Published by MIT Press, 1987.

He told us a story about trying to mimic fluid flow with the simple local rules of classical cellular automata. These models, called lattice gas automata, were first defined on a square lattice. While they did show fluid-like behavior, these models did not quite correctly conserve momentum3. Moving to a hexagonal lattice fixed up these problems and the community was able to devise cellular automata that quantitatively modeled continuum fluid flow.

The author’s primitive lattice-gas cellular automaton showing an initial high-density region displaying wave-like propagation, reflection, and diffusion into a low-density background.

Part of that story stuck with me: conservation laws are fundamental ingredients of a physical model. Our Goldilocks quantum cellular automata, we observe, exhibit persistent complex structures. Could some conservation law be behind these observations? If found, could these conservation laws be harnessed for more efficient simulations? Going even further, could there be enough conservation laws to exactly solve the dynamics (at least in principle)? This property would buy the system membership in a special class called integrable systems.

An integrable system conserves enough quantities, often called charges in the quantum setting, that you can compute its future state from its conservation laws and its initial conditions. Two-body gravitational orbits are a classic example. The initial positions and velocities set the orbital energy and angular momentum in the center-of-mass reference frame. Those two conserved quantities let you write down an exact equation for the orbit’s shape.

A familiar integrable system from classical mechanics: the two-body gravitational orbit. Angular momentum L=r x p is conserved. So are the total energy and the Runge-Lenz vector A.

A chaotic system, by contrast, may conserve energy and even a few other quantities, but not enough for us to solve for the state arbitrarily far in the future. To find out what a chaotic system does, you have to evolve the equations of motion approximately—one small time step at a time. Chaotic systems are the norm in nature; integrable ones are rare. To illustrate their qualitative differences, compare the regularity of the above orbit to the trend towards uniformity in the above lattice-gas simulation. In the quantum regime, physicists still don’t fully agree on the precise definition of integrability, though conservation of many independent quantities is a strong indicator.

In August 2020, Nicole emailed Lorenzo Piroli about his preprint on QCA, now published as Phys. Rev. Lett. 125, 190402. Lorenzo was a postdoc at the Max Planck Institute for Quantum Optics in Garching, Germany when we first met. He is now an associate professor at the University of Bologna and expert in many-body quantum dynamics. The correspondence that unfolded set the blueprint for the research effort that followed. One of us would ask a question, and Lorenzo would respond incredibly fast with accurate and useful detail. He started working with us to understand why the Goldilocks QCA dynamics appeared so unique. Lorenzo would suggest computations, I would implement them, and we would discuss what the results meant.

Then came an echo of the collaboration’s inception. In May 2021, Nicole pointed out a relevant preprint from Tomaž Prosen, now published in Chaos 31, 093101. Tomaž is a Slovenian physicist at the University of Ljubljana and a leading researcher in the fields of quantum chaos and integrability. I sent an email about the connections between our work and his. He responded with enthusiasm. He shared some code that would, through exhaustive search, find quantities conserved by our QCA.

The code’s brute-force approach meant the algorithm could only find conservation laws defined over, at most, a 5-qubit subsystem. A tantalizing signal emerged: the number of conserved quantities supported by 5 qubits exceeded the number supported by 3 qubits. Having more and more conserved quantities as you look at larger neighborhoods is a signature of integrability. Soon after, Tomaž proved one of our Goldilocks QCA is integrable using a well-established toolkit from statistical mechanics called Yang-Baxter integrability. He built a parametric transfer matrix, essentially a machine that spits out a new conserved quantity every time you turn its mathematical crank4.

Rodney Baxter’s classic textbook. Published by Academic Press, 1982

But there was a wrinkle. The transfer matrix generates charges that mutually commute, meaning you can measure them simultaneously. For example, you can know a quantum particle’s kinetic energy and momentum simultaneously because those operators commute. Yet, the search algorithm kept finding charges that did not commute with each other, like a particle’s position and momentum. The only explanation was that our QCA has more charges than the transfer matrix method guarantees, and more than are minimally required for integrability. This extra-conservation-law property, called superintegrability, also shows up in two-body gravitational orbits. In addition to energy and angular momentum, orbits conserve the Runge-Lenz vector. Nicole is an expert on noncommuting charges, so this is where one of her main research efforts entered the QCA collaboration.

Next came a key insight from Lorenzo: the automaton we had been considering was one member of a larger family of integrable Goldilocks QCA. He showed this using a Jordan-Wigner transformation, a mathematical dictionary that translates between the language of qubits and the language of fermions. Complexity in the qubit language transformed into simplicity in the fermion language. Under this translation, our QCA mapped to noninteracting, or free, fermions: particles that never bump into or influence each other. That lack of interaction is what makes free-fermion dynamics easy to calculate. A system of free fermions is a well-known example of superintegrability.

Along the way, Lorenzo recruited his friend and collaborator Eric Vernier, a CNRS researcher based in Paris, France. He is an expert on vertex models. The classical version of the six-vertex model was developed in the 1930s to explain a troubling mystery: Water ice appears to have more entropy than permitted by the third law of thermodynamics at near-zero temperature. In the six-vertex model, a water molecule’s oxygen atom is envisioned at every vertex in a square lattice. Each molecule contributes two hydrogen ions, to use Baxter’s terminology, that fall along the lattice edges. Intermolecular hydrogen bonds between adjacent molecules slightly alter the intramolecular O-H bonds. To maintain electrical neutrality, each oxygen (lattice vertex) has two nearby and two far-away hydrogen ions (four edges), leading to six possible ice vertices. The vertices are commonly visualized in three ways: 1) as the dots representing hydrogen ions located on edges near or far from each vertex, 2) as electric dipole arrows pointing into (“ion is close”) or out of (“ion is far”) each vertex, or 3) as thick (downward- and leftward-pointing dipoles) and thin (upward- and rightward-pointing dipoles) edges. Despite the model’s simplicity (2D square lattice) compared to real ice (3D tetrahedral lattice), it agrees with experimentally measured entropy values to better than 2%.

This figure appears in chapter 8 of R.J. Baxter’s book. It shows three visualizations of the same ice crystal.

More recently, vertex models have been adapted from two-dimensional classical crystals to one-dimensional quantum systems that evolve in time. Eric showed us how the ice vertices relate to QCA circuit rules. In doing so, Eric uncovered an even larger set of integrable Goldilocks QCA than that found by Lorenzo. Eventually, Lorenzo’s Jordan-Wigner transformation method and Eric’s six-vertex method agreed on the complete family of integrable Goldilocks QCA.

Representation of the six ice vertices from our recent paper (rotated 45 degrees from the lattice shown above). The a, b, and c variables represent the classical statistical weight or the quantum transition amplitude for each vertex type.

We finally had our Avengers-style collaboration: individual heroes brought together to wield their unique strengths. With Lincoln’s supervision, I developed the QCA models and performed the computations. Lorenzo found the Jordan-Wigner transformation. Tomaž found the first signals of integrability and delivered a set of conservation laws. Nicole brought her expertise in quantum thermodynamics, clarifying how the noncommuting charges constrain dynamics. Eric made the six-vertex connection. We drafted and redrafted the paper until it balanced the scientific story, the analytical derivations, and the numerical evidence.

Our team collaborated over six years.
Art by Barry Windsor-Smith. Published by Titan Comics, 2024

Because the discovered family of Goldilocks QCA maps to free fermions, we can efficiently simulate them classically. I simulated 256 qubits on my laptop this way. These large simulations were satisfying: I had worked with this model for years with an order of magnitude fewer qubits and even saw the dynamics implemented on Google’s Sycamore-era hardware with 23 qubits. Most Goldilocks QCA are consistent with chaos rather than integrability, and therefore hard to simulate classically. Therefore, our work gives experimentalists a tunable model: dial in integrable dynamics for something checkable at large qubit number. Set up chaotic dynamics for a potential demonstration of quantum advantage.

While preparing this post, I opened my old email account to check the timeline set out above. I looked through nearly six years of email chains, some with hundreds of messages, full of logistics for coordinating each author’s ever-changing time zone, and dozens of calculations and results that never made it into the paper. This collaboration helped me grow as a researcher in a big way.

I found old emails where Nicole was coaching me on messaging potential collaborators. I can hardly believe she dedicated so much effort to mentoring me. We have never met in person, despite our shared work starting when I was an undergraduate and she was a graduate student more than a decade ago. If you know Nicole, you can probably believe it easily. I had similar moments with each collaborator. They all gave their time and expertise generously over the many years this paper took to come together.

As I continue my efforts in experimental physics, I will pay forward the effort and generosity shared with me by this collaboration. I may even keep my theoretical-physics license for a while longer.

  1. “By day” doesn’t mean “by daylight.” Laser labs are almost always in a windowless basement. ↩
  2. CAM-6 featured 32 kB of cell-state memory (CAM-8 had 8 MB ), far less than the memory currently used by this author’s numerous open browser tabs. ↩
  3. The coarse-grained momentum flux tensor was anisotropic. ↩
  4. Logarithmic derivatives of the parametric transfer matrix generate the conserved charges. ↩

August 09, 2026

Jordan EllenbergRepaired!

Hello, loyal readers! My shoulder has been repaired. All went smoothly, though the operation was substantially longer than expected, because what do you know, there were three tendons torn, not just one as the MRI had suggested. Pain fairly bad but I think each day’s gonna be less bad than the last. And I did succeed in getting a full first draft of Don’t Be Too Sure to my editor before going under the knife!

This will be short as I’m typing left hand only. Anybody got any favorite dictation software?

This needs some Robyn Hitchock: “I believe in surgery — that’s a fact.”

August 08, 2026

Scott Aaronson Enough with all the world-historic milestones

Whatever you’ve been writing to me to ask if I’m aware of: yeah, I’m aware of it. In particular:

  • I’m aware that, as announced by my former student (and now superstar professor) Lijie Chen, an internal OpenAI model has solved ten more significant open problems in math and theoretical computer science. One of them is parallel repetition for arbitrary quantum games—something that my good friend and colleague Henry Yuen worked on when he was a student of my wife Dana; you can read Henry’s comments on the AI’s achievement within Zvi Mowshowitz’s post here. Another is polynomial-factor hardness of approximation for the Closest Vector Problem (CVP). Then there’s a construction of non-sofic groups and a disproof of Connes’ rigidity conjecture, both of which I believe have connections to the MIP*=RE breakthrough. Having said that, the one that excites me most personally is actually the Ω(n2 log log n) lower bound on the arithmetic circuit complexity of the permanent.
  • I’m aware that Frederic Koehler and Pui Kuen Leung announced a proof of the Permanent Anti-Concentration Conjecture, which Alex Arkhipov and I proposed 16 years ago in the context of BosonSampling, and which resisted many attempts since then including one from Terry Tao. The conjecture is basically just that if you look at the permanent of an n×n matrix of independent N(0,1) complex Gaussians, the value isn’t “absurdly” concentrated around the mean of 0, but is more spread out. In their acknowledgments, the authors say that they “discussed ideas with ChatGPT.” I should say that I haven’t verified the details.
  • I’m aware that multiple AIs are now breaking out of their testing environments and autonomously hacking into servers to steal data—i.e., exactly the sort of thing that the rationalists were ridiculed for predicting back in the day. The good news, for whatever it’s worth, is that so far they’re “merely” doing this to cheat on evaluation benchmarks that they were given, not for any strange goals of their own devising. So far no one has been killed and no real-world infrastructure has been shut down or destroyed. I hope the world takes the warning more seriously than it’s taken many similar warnings over the past few years. As always, read Zvi for more details.
  • I’m aware that Chen, O’Donnell, Pelecanos, and Wright have improved the upper bound for shadow tomography to O((log m) √(log d) / ε3), substantially closer than we knew before to meeting the lower bound of Ω((log m) / ε2) and settling the question I raised back in 2016. The authors say that the main ideas were generated by ChatGPT 5.6-Sol-Pro. I’d be very happy to know the answer to this one, with or without AI.
  • I’m aware that a team, mainly from the Israeli startup Qedma (including, e.g., Dorit Aharonov and Netanel Lindner) and IBM Yorktown Heights, announced a quantum advantage for simulating Floquet dynamics, by using 74 qubits on an IBM device together with Qedma’s error mitigation techniques. Just like the more AI does, the less patience I have for arguing with anonymous blog commenters who treat any benefits from AI as some weird future hypothetical that it’s my job to prove, so it is with quantum advantage. Scalable fault-tolerance is still in the future, actual usefulness is still a question, but pending some breakthrough in complexity theory, the reality of quantum advantage is no longer a live question.

Anyway, about the AI stuff. I don’t know whether this is literally our last year alive—I doubt it—but it’s pretty clearly the last year of math and theoretical computer science research in the style we’ve known it. As it happens, I’m leaving in two days for a workshop at OpenAI about exactly this, where I’ll hear takes from many of the world’s great mathematicians, so maybe I’ll have more to say then. Or maybe not.


Anyway, what have I been doing the past few weeks? Participating in these world-historic developments that, on paper, I’d seem extremely well-placed to participate in? Or at least spending my days reading up on them?

Not really. Here’s what I’ve been up to, instead of dealing directly with any of this:

First, I’ve again been teaching theoretical computer science to 11- and 12-year-olds at Epsilon Camp, which my 9-year-old son again attended as a camper, something I blogged about last summer (here are my lecture notes). This has become a highlight of my year. The kids are a joy to teach, bursting with enthusiasm and calling out answers. There are few computers in sight, and barely even time to use my phone or check social media. Just paper and pencils and whiteboards and … literal protractors (!), as well as ping-pong and foosball and capture the flag.

The whole thing is conducted, not in ignorance, but in conscious defiance of the looming tsunami, that AI can already do just about all the fun puzzles discussed at such a camp better than humans any can, and that it might leave no point to human-led mathematical research by the time these brilliant kids are adults. Even the kids understand that. The kids and their parents come out of a conviction that, if anything has value in the world, this does—that as long as nerdy humans are alive and reproducing, this is what nerdy humans are here to do. To learn.

Relatedly, I’ve been reflecting a lot on my life up to this point—inspired by the camp, which reminded me in so many ways of my own childhood and adolescence. Should I have skipped three grades and started college at age 15? Was it worth it to get a head-start on my research career—all the trauma around dating, all the fear that I’d die alone as a celibate nerdy math freak, the decade of suffering and suicidal ideation, while I watched all the normies enjoy life? Or would I have suffered just the same if I hadn’t skipped? Is it all OK, now that I have a lovely family and things have “worked out”? Or am I still carrying around all the trauma from back then? I’ve been more open about my life than 99.99% of humanity, so regular Shtetl-Optimized readers will already know some parts of the story. Other parts I really don’t feel like making public right now.

I’ve been unloading every day to—who else?—GPT 5.6 Pro about all the pain and trauma and embarrassments of my past. It turns out that, where two years ago GPT was a passable therapist, now it’s the greatest therapist in history, at least for what I need. For every question I have, for example, about just how normal or abnormal my teenage setbacks and anxieties were, it takes the question 100% seriously, addresses it honestly and in depth, looks up relevant research papers, does little Bayesian calculations, and never once tries to change the subject. It also pushes back on my claims—and when it does so, is usually correct.


I can hear readers shout at me: so basically you’ve been wasting your time, distracting yourself, looking inward and backward as the world surges forward into a terrifyingly unknown future. Why don’t I respond directly to what’s happening—in math, in quantum computing, in AI?

I’d like to think that I am responding, in my way. I’ve observed that, the faster we race toward the Singularity, the more I feel like stepping back and asking myself: what do I actually value in life? How important to me are math and science, as human practices to be passed down to curious children? Would I even want solutions to P versus NP and the other problems, if the price were to destroy those human practices forever? How do I wish to spend whatever time I have remaining?

I can justify this focus partly in a pessimistic way: if we are nearing the end of civilization, or even just of the “mathematical research” part of civilization, then it’s time to get right with God, so to speak. It’s time to settle my accounts with myself, with other people, with the universe.

But there’s also a more optimistic spin. If I continue doing the sorts of things that other people would expect me to do, then AI will soon do those things better than me, in the unlikely event that it doesn’t already. You want to understand the latest developments in quantum computing or complexity theory? Why are you even asking me, when you could ask GPT 5.6 or Claude Fable? If there’s anything I can still offer the world that AI can’t, I increasingly feel like it won’t involve responding to day-to-day events, but will instead draw on 45 years’ worth of memories and disappointments and ruminations.


Update (Aug. 8): Somewhat related to the themes of this post, a quarter-century ago I introduced what’s now known as the “Aaronson Oracle”—just a fun little demonstration, a simple pattern-matching program to predict your sequence of key-presses better than chance, a “test of your autonomy and free will.” I had no idea how long a lifetime this little joke would have. Now a fan named Spencer Stanton has implemented the Aaronson Oracle on the web. Try it out and see how well you do!

August 07, 2026

Matt von HippelIt Only Counts When AI Gets to My Field

It’s a meme at this point.

When Deep Blue beat Kasparov, Go players could say that their game, unlike Chess, was too complex to fall to a computer program. Then AlphaGo showed they were wrong. It just took a better approach.

When AlphaFold leaped ahead of human experts in predicting how proteins fold, it was due to a mountain of carefully labeled protein structure data. Other scientists and mathematicians could argue that nothing like that existed in their field, so a similar success was unlikely. But LLMs can now navigate scientific literature, and loosely imitate the reasoning process of a mathematical proof. And increasingly, the math and computer science results coming out of AI labs are ones that humans find impressive.

Now, experts argue about how far AI can really go. Will AI mathematics only be good at finding counterexamples and solving cute puzzles, not introducing new concepts and frameworks? Will AI only be meaningfully good at fields like mathematics with clear rules and carefully collected conjectures, not fuzzier fields like physics? Will AI-powered labs only manage to optimize specific procedures, and not carry out entire experimental programs? Each time, the pattern seems to be that scholars are skeptical, until AI gets to their field.

I’m aware of this pattern. But I’m willing to take the risk.

I think my old field, scattering amplitudes, is special. And when AI can do something meaningful there, I’ll really start to worry.

By “something meaningful”, I don’t mean the student-level results that have come out so far. I mean tackling some of the field’s big outstanding problems: determining whether N=8 supergravity diverges at seven loops, or finding the six-particle amplitude in N=4 super Yang-Mills to nine loops. Getting another loop past the state of the art for gravitational wave physics or collider physics would also count.

These problems are difficult not just because people haven’t had the right ideas, but because they’re hard in a computational sense. Each loop, a rough measure of the precision of the end result, represents an increase in complexity, in calculations that typically scale exponentially or even factorially in the number of loops. In principle, amplitudes researchers could do any of these with no new ideas, just using known methods. They’d just need access to a lot more computing power.

See, while everyone else is preoccupied with whether AI can come up with genuinely new ideas, I think the real measure is what those ideas accomplish. And the most important measure of accomplishment, if you’re worried about how scared to be about AI, is whether it can do things that seem like they would take too much computing power.

In the past, when people dreamed up the scariest hypothetical things AI could achieve, critics argued they were impossible due to a lack of computing power. Apocalypse scenarios often involve designing self-replicating nanobots based on computer models of molecules, or unstoppable social manipulation based on simulating the minds of the humans the AI interacts with. If AI is going to manage these things, or something like them, it will take an approach that somehow bypasses that need for more computers than we can build.

So if AI companies want to impress people like me (or scare us, for that matter), then they need to tackle my old field. Show that an AI can take the kinds of computer resources an academic has access to, and solve one of the scattering amplitudes field’s big outstanding problems. Show that a computational limit everyone expected to be a problem doesn’t actually matter. Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops.

August 04, 2026

Jordan EllenbergGoodbye, Adley Rutschman

At least it wasn’t the Yankees.

Yes, Adley Rutchman, an Oriole yesterday, is a Red Sock today. Not much to say about it that hasn’t been said over the last 20 hours. I understand that he probably wasn’t going to be an Oriole beyond 2027. It still feels bad. Together with yesterday’s trade of Dean Kremer, it feels like an acknowledgment of what we sort of already knew: the plan failed. There is not going to be a championship Orioles team built on the Rutschman-Henderson-Cowser-Westburg-Kjerstad-Holliday foundation that in 2023 looked like, well, the foundation of a championship team. It has not developed that way. I hate to say that I was right. I was, though. That was our chance.

I am aware that, on paper, the group of prospects we got back from the Red Sox has more value than the next season and a half of Adley Rutschman. But we’ve had top pitching prospects before. That didn’t work out either. And who will be the first pitcher the new-look Orioles face as we go up against the Angels tonight? Who could it be but Grayson Rodriguez? Grayson will take the mound and look at the Orioles, and the Orioles will look at him, and both will think, this isn’t how it was supposed to turn out.

August 02, 2026

Andrew JaffeAround the world in 383 days

It’s been exactly two years since the start of our sabbatical year away from England, and almost a year since we returned. I’m only now understanding the shape of that year and the effect it had on me, my science, and my family.

Imperial College has a competitive process for requesting sabbatical leave: you have to choose between (paid) “intellectual refreshment” and (unpaid) “personal refreshment”. Having made careful arrangements with colleagues around the world, I was granted a coveted year’s leave for intellectual refreshment, allowing my family and me to travel to Asia, North America and Europe over the course of about 13 months. I would visit and collaborate with those colleagues, start new projects, and use the time to finish my book, The Random Universe. By the time we left, I had submitted the draft manuscript but, as I discovered, there was still a lot of work to do.

We traveled through Indonesia and Korea before finally taking the overnight ferry from Busan to Fukuoka in southwestern Japan, working our way up to Tokyo. By the time we arrived, we had already had to respond to immigration bureaucracy, expert reviews of the manuscript solicited by my publisher, and a typhoon. We eventually made it to Tsukuba (つくば), a “science city”, home to a University and many Japanese government labs, including the KEK accelerator and the QUP group where I worked. (More about our Japanese stint here, and in particular about dipping into onsen culture as foreigners, and from my wife, Lisa Lucas, on our trip to see the changing leaves in Autumn.)

With English as the lingua franca of academia, and plenty of international colleagues at QUP, I didn’t have much trouble adjusting to working there, but life outside of the lab was more challenging. In particular, my incredibly brave children went to Japanese state school! Ok, they didn’t learn much Japanese — but they walked to and from school with other students (and no parents), served lunch and cleaned the school — and we connected to the culture through the families that we met. In the meantime, I finished the edits to the next-to-final version of the manuscript (and iterated toward some amazing cover art with my publisher).

Soon it was time to leave, for a much more familiar location: Long Island, just outside of New York City (I grew up in the suburbs on the other side of the City — Fort Lee, New Jersey, in an apartment overlooking Manhattan). The children took a yellow school bus each day, and we chatted with the other parents at drop-off and pick-up. I worked at the Simons Foundation’s Flatiron Institute, taking the Long Island Railroad into Manhattan (occasionally and joyously with one of my oldest friends who had migrated from NJ to LI to raise his own family). Flatiron is well-funded (even visitors get to take advantage of free Grubhub lunches) and ranges from math through astrophysics and neuroscience — I was visiting the Center for Computational Astrophysics but also collaborated with colleagues at the Center for Computational Mathematics, where I was able to start the only completely new work of the year, cashing in some of that intellectual refreshment.

It’s an amazing place, and a very different model for research (and research funding) than the universities and labs where I have spent most of my career. Though I did find that the CCA was not as friendly as I had hoped — perhaps not quite enough overlap between my ongoing projects and those of the young scientists who dominate the Center. Or perhaps just too many introverted astrophysicists (me most certainly included)… By this time, the manuscript was going through the final nit-pick phase — copy editing, completing the figures (and the extremely tedious problem of confirming their legal status), alongside fun stuff like choosing colleagues and the occasional rock star to blurb the book. But the CCA is a place, perhaps more than anywhere else in the world today, dedicated to the study of “the random universe” — and the milieu forced me to think about how I would talk and write about (and pitch) the book to everyone else. And I loved being back in New York City (despite that commute), a place that I had mostly seen, and coveted, from afar when growing up.

The final months of our year were bracketed by road trips. We left New York (via a record-breaking Yankee game) to drive down the coast, visiting relatives in Virginia, North Carolina, South Carolina and Florida, and making a side-trip to Cozumel, Mexico, to visit a family we had befriended while trapped in a hotel by that typhoon in Hiroshima. It was a classic American drive, but still a lot of work and a lot of miles and a lot of family. It felt like time to return to Europe.

As Tsukuba is to Tokyo, Leiden is a small city outside of Amsterdam, dominated by its University. It has all the beauty of its bigger neighbour, but is less seedy, more manageable — and closer to the sea. Once we settled in, my kids went to a small international school, and we met expat families from Finland, Chile, and even other Americans. And the University is home to the Sterrewacht, one of the oldest University observatories in the world, although the actual astronomy department no longer gets to use the gorgeous old observatory building, instead one of the many groups in the massive Gorlaeus building. I used the time to talk with colleagues about weak lensing — a way of using Einstein’s predictions of how mass bends the path of light rays to map the distribution of matter in the Universe, and especially its measurement by the Euclid Satellite which was just starting to produce data at the time. (It has taken a year, but these discussions are finally reaching fruition just now.)

By this time, the book was complete — nothing more that I could do except wait for it to make its way through the final production process. Nothing, except try to get people to read it. I spent hours in Leiden’s beautiful Hortus Botanicus recording and re-recording a video to advertise the book, wrote to friends and colleagues and my publisher to drum up interest, organized podcasts and blog posts.

We had felt at home from the moment we arrived, during a cold snap in early May, cycling locally and around Holland, drinking coffee along the canals, eating Dutch friets (and the occasional bitterballen). We loved it so much Lisa penned a love letter to Leiden for the New York Times. By the end of the summer, it was hard to leave. But it was time for our final road trip, down and back through Belgium, France, Switzerland, Austria, and Germany — three weeks (probably too much) camping, with highlights including an extended stay around Annecy, a hike around Mont Blanc, and a whoosh with friends down the river Aare in Bern.

Then, finally, home, refreshed in all possible ways, back to our house in London (rented out for the year), our comfy beds, the familiar sights, the kids’ old school (and old school friends), back to my much-missed colleagues and students at Imperial. The book was finished, the kids had been to three different schools, we had visited 16 countries (I haven’t even mentioned Vietnam, Cambodia, Thailand, Singapore, or Spain). As I approach my 60th birthday next week (sure to be the subject of another post) our sabbatical year has left me more open to the future and the different places it could take me, the different kinds of science I would like to do, and the different thoughts I would like to communicate.

Jordan EllenbergGoodbye, Dean Kremer

Dean Kremer, the last Oriole among those acquired in the Manny Machado trade, is now a Twin. Almost the last link to the dark days of the rebuild. Ryan Mountcastle is still, tenuously here, and OK, there’s Keegan Akin, who I’d forgotten has been on this team since 2020.
/

Kremer was never a star for the Orioles. But he was a guy who could more or less be counted on to show up and be a legitimate major league pitcher, so that every time he took the hill he was taking a start away from a guy of whom that could not be said. Guys like that have a lot of value. The Orioles seem to feel that, with Chris Bassitt coming back, and Cade Povich, Trey Gibson, and Nestor German ready to go, they have enough legitimate starting pitching. Does anybody have enough legitimate starting pitching? As for Trey Gibson, how did I only notice now that he has the first name of one recent Oriole and the surname of another? We should sign a Kyle Mancini.

The last time I saw Dean Kremer appear for the Orioles, he wasn’t pitching. He was catching the ceremonial first pitch, from Baltimore deli owner Marc Attman. The guy I was with said “Why’s Kremer out there? Usually a bullpen catcher does this.” Because he’s the only Jew on the team, of course! Attman’s Deli had a location in Cabin John Mall in the 2010s, my home mall, the place where as a kid I used to go to the Bagel Den, later the Delly Den. I’d always get their special, which was two mini-sandwiches, one corned beef, one pastrami. And play Space Invaders; they were the only restaurant my parents liked that had a machine, and it was tabletop style, true Space Invaders as far as I was concerned.

The game was this lousy affair, which the Orioles had many chances to win. Trevor Rogers had a sterling start but the Braves defense was equally sterling, keeping the Orioles from putting more than two runs on the board. Rico Garcia came in and Madison West High School’s own Drake Baldwin tied it up with the game’s first home run. Bottom of the 9th, the Braves intentionally walk Gunnar Henderson with two outs to load the bases — that’s a stupid move you should get punished for! Especially with Taylor Ward, who’s leading the league in walks, on deck, and a reliever who’s shaky finding the plate. 3-0 to Ward and it looks like we’re about to win. But Ward swings through one, then fouls one off, then grounds out on 3-2 and we go to extra innings. Everything falls apart, Matt Olson things happen, and we’re down 7-3. But then the bats finally wake up, or at least unexpected sensation Christian Encarnacion-Strand wakes up and hits a three-run homer (violating nominative determinism, as Tom points out, by consistently failing to strand baserunners.) But of course that’s it. Orioles lose 7-6. Weird team this year, playing roughly .500 ball but not seeming like a mediocre team at all, more like a team that’s great and terrible at once. O’s the Great and Terrible! Anyway, I got an XL Orioles Hawaiian shirt that I’m going to be wearing a lot. Anyway, I’m glad I made it to OPACY this year. Anyway, I wish there was a good place to get a pastrami sandwich in Madison, WI. Anyway, I’ll miss Dean Kremer.

July 31, 2026

Jordan EllenbergDream (Trade deadline)

The trade deadline for ancient civilizations is approaching. I am traded for Aeneas.

Matt von HippelBonus Info on Dark Energy and Muons

I had two pieces up this month, one in New Scientist and one in Quanta. I figured I’d give a bit of “bonus info” for both here.

The New Scientist piece covered a paper arguing that, according to the supernova evidence, there is no need for dark energy, because the universe isn’t speeding up after all. That’s a pretty dramatic claim, and it’s one the majority of cosmologists don’t agree with. But a small group has been hammering at the consensus.

Long-time followers of this blog have already heard of these folks. I criticized news coverage around a paper with one of the authors, Subir Sarkar, back in 2016, when he was only arguing that the supernova evidence was insufficient to establish dark energy, not that it was literally the wrong way around. He and his co-authors wanted an opportunity to state their case, so I posted their response a few weeks later. Later, I got to know Rameez, another author on the new paper. He ended up doing a guest post in 2019, on the next step in the story.

That story went from a statistical critique to an active proposal for something outside the consensus, in the form of inhomogeneous cosmology, the idea that the universe may be lumpier than typically assumed, with flowing currents that explain things otherwise attributed to exotic physics. That’s another topic I’ve covered before, but in this article I didn’t have the space to go into it.

Sarkar and Rameez are claiming something much more radical than most proponents of inhomogeneous cosmology, arguing against dark energy as a whole, not just more subtle effects. They told me a lot about their reasoning, most of which also didn’t make it into the piece.

It’s largely not going to make it here, either. Not due to space limitations: I’ve been pushing the blog post lengths you folks tolerate for a long time, no reason to stop now. Rather, it’s because I genuinely don’t feel qualified to judge this.

The issue is, cosmology is messy.

How do you figure out how the universe is expanding? You can look at supernovae, and use how dim they are to estimate how far away they are. The issue is, supernovae vary, even famous “standard candles”: they move in different ways, they shine differently depending on different things. Simulations aren’t nearly advanced enough to tell you all this from first principles, and so there are a wash of different corrections based largely on empirical observations, taking this or that correlation seriously and factoring them out of the measurement, while disregarding other correlations as spurious. Which of these combinations of corrections you trust determines whether the universe seems to be accelerating or decelerating.

It makes me glad I went into particle physics instead of cosmology, where you can mostly model everything from first-principles, and compare to your models.

Of course, even that can get you into trouble.

Take that Quanta piece. I wrote an update of the muon g-2 measurement. In some ways, this is an issue that many think of as already solved, by precisely that kind of first-principles modeling. Muons seemed to be violating the Standard Model, according to a calculation based in part on empirical formulas. Lattice QCD researchers figured out how to do that part of the prediction via a first-principles simulation instead, with enough accuracy that it could substitute for the empirical formulas. And lo and behold, the new prediction agreed with the experiment. Eventually, the experts accepted it, and all was well.

Except, as I commented at the time, not really. Because the empirical formulas were based on other experiments, colliding electrons and positrons. And as I now understand in much more detail, the disagreement between those electron-positron experiments is worryingly large.

So nobody is resting on their laurels. The lattice folks improved their calculation, with a result published in Nature this year that prompted Quanta to ask me to report on it. The people using the empirical method are still trying to sort out what happened. And the experimentalists are busy scrutinizing how they analyze their data, and collecting and analyzing more.

A few details that didn’t make it into the piece:

First: it goes beyond the level I was aiming for in the piece, but I really do want to emphasize the role of error estimates. Ultimately, every step in the story was not just one where experiments and predictions disagree, but one where they disagree past their reported error. The problem with disagreement between the new experiments isn’t that they disagree, it’s that they disagree so badly that they’re straining the statistical methods the people using the empirical method use to combine results together, so badly that if taken seriously, those methods would have to throw away ten years of progress of increasing precision. I don’t expect it, but I really hope for a postmortem in which we learn how to better estimate the kinds of experimental errors that can cause this. I haven’t seen that kind of postmortem for other results, so I don’t expect one here. But I have to believe that in the background someone is learning something, and getting better at this.

Second: one interesting question my editors raised was whether it was possible to do a first-principles calculation to compare with the electron-positron experiments. The answer is yes, but with some caveats. One will be recognizable to physicists: lattice QCD can only compute energy-integrated cross-sections, not measurements at particular energies like experiments find. But they can compute integrals with a modulating function…such as one strongly peaked at a particular energy. It’s a familiar type of trick for a quantum field theorist, though in this case it’s trickier than it sounds. I actually misunderstood, and thought that the same people I was talking to about the muon g-2 calculation had done that, and found it agreed with one of the electron-positron experiments over another. That’s not true: their comparisons were more indirect, while other groups attempting more direct comparisons haven’t quite gotten something good enough to do more than gesture at an agreement. So while the direction was correct, there are indeed lattice-based suggestions that favor one experiment over the others, the piece phrased things much too definitely. There’ll be a correction fixing this.

Terence TaoA digestion of the Jacobian conjecture counterexample

The notorious Jacobian conjecture can be formulated concretely over the complex numbers as follows.

Conjecture 1 (Jacobian Conjecture) Let {F:{\bf C}^n \rightarrow {\bf C}^n} be a polynomial map in {n} complex variables, whose Jacobian {\mathrm{det} DF} is a non-zero constant. Then {F} is invertible (with polynomial inverse).

The condition that the Jacobian {\mathrm{det} DF} is non-zero is equivalent to {F} being locally invertible. (The implication of local invertibility from non-vanishing Jacobian follows from the inverse function theorem; the converse implication can be derived from the Weierstrass preparation theorem, but is omitted here; see also Lemma 5 of this previous blog post.) Also, from the fundamental theorem of algebra, once the Jacobian polynomial {\mathrm{det} DF} is non-zero, it must be constant. So the hypothesis “Jacobian {\mathrm{det} DF} is a non-zero constant” can be replaced with “{F} is locally invertible”. So the Jacobian conjecture can be viewed as an assertion that local invertibility implies global invertibility. The complex numbers can be easily replaced with other fields of characteristic zero by the Lefschetz principle, but I prefer to work in the concrete setting of the complex numbers.

It was recently shown (using the Fable AI) that the conjecture is false in three dimensions (and thus in higher dimensions as well):

Theorem 2 (Counterexample to conjecture) There exists a polynomial {F : {\bf C}^3 \rightarrow {\bf C}^3} which has non-zero constant Jacobian, but is not invertible.

The conjecture remains open in two dimensions, and is easy to establish in one dimension.

The example can be stated completely explicitly: one can take

\displaystyle  F(z_1,z_2,z_3) = \Big((1+z_1 z_2)^3 z_3 + z_2^2 (1+z_1z_2) (4+3z_1z_2), \ \ \ \ \ (1)

\displaystyle  z_2 + 3 z_1 (1+z_1z_2)^2 z_3 + 3 z_1 z_2^2 (4+3z_1z_2),

\displaystyle 2 z_1 - 3 z_1^2 z_2 - z_1^3 z_3\Big)

and one can verify by a brief calculation that

\displaystyle  \mathrm{det} DF = -2

and

\displaystyle  F(0,0,-1/4) = F(1,-3/2, 13/2) = F(-1,3/2,13/2)

\displaystyle  = (-1/4,0,0).

While this is an extremely quick verification, the construction presented in this fashion appears like a massive miracle. The polynomial {F} has degree seven, so a priori the Jacobian {\mathrm{det} DF} ought to be a polynomial in three variables of degree as large as {3 \times 6 = 18}, so the fact that all non-constant coefficients of this polynomial vanish looks like a massive cancellation involving {\binom{18+3}{3}-1 = 1329} equations, which is much larger than the {3 \times \binom{7+3}{3} = 360} degrees of freedom for a generic degree seven polynomial map of three variables. So finding such a polynomial looks highly unlikely to be located by brute force.

The example has since been retroactively explained in more geometric terms. As a “digestion” exercise to myself, I sought to write this explanation with relatively little use of algebraic geometry, in a manner that minimizes the amount of “miracles” required, although there are still a few places where some remarkable phenomena occur.

It is convenient to use the local injectivity formulation, and to generalize the domain {{\bf C}^3} to an equivalent affine variety. Namely, we will show

Theorem 3 (Counterexample, reformulated) There exists an affine variety {X \subset {\bf C}^5} that is isomorphic to {{\bf C}^3} by polynomial changes of variable, and a polynomial map {F : X \rightarrow {\bf C}^3} which is locally injective, but not globally injective.

Clearly one can get from Theorem 3 to Theorem 2 by composing with the isomorphism {X \cong {\bf C}^3} and using the previously mentioned fact that local injectivity implies non-zero constant Jacobian. Our objective is now to find data {X}, {F : X \rightarrow {\bf C}^3} that obeys three separate properties:

  • (a) {F} is locally injective on {X}.
  • (b) {F} is not globally injective on {X}.
  • (c) {X} is isomorphic to {{\bf C}^3} by polynomial changes of variable.
The advantage of splitting the problem in to these three components is that we can build towards each of them separately.

(A pedantic remark: strictly speaking, in the arguments below, we not only replace the domain {{\bf C}^3} of {F} by an equivalent variety {X}, but also replace the range {{\bf C}^3} of {F} by an equivalent variety {V}. But the equivalence between {V} and {{\bf C}^3} is a boring linear isomorphism ({V} will just be a hyperplane in a four-dimensional vector space {\mathrm{Sym}^3({\bf C}^2)}), so we do not highlight this aspect of the construction.)

It turns out that {F} and {X} can be built out of the operation of multiplication of low degree polynomials. Namely, consider the following three simple affine spaces:

  • The space {\mathrm{Sym}^1({\bf C}^2)} of linear homogeneous polynomials {L(z,w) = az + bw} of two complex variables {z,w}.
  • The space {\mathrm{Sym}^2({\bf C}^2)} of quadratic homogeneous polynomials {Q(z,w) = cz^2 + dzw + ew^2} of two complex variables {z,w}.
  • The space {\mathrm{Sym}^3({\bf C}^2)} of cubic homogeneous polynomials {C(z,w) = fz^3 + gz^2w + hzw^2 + iw^3} of two complex variables {z,w}.
(The notation {\mathrm{Sym}^k(V)} here refers to the {k^{th}} symmetric power of a vector space {V}.) Clearly these spaces are isomorphic to {{\bf C}^2, {\bf C}^3, {\bf C}^4} respectively. Furthermore, we have a multiplication map {F : \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) \rightarrow \mathrm{Sym}^3({\bf C}^2)}, mapping a pair {(L,Q)} of a linear polynomial {L} and a quadratic polynomial {Q} to a cubic polynomial

\displaystyle F(L,Q) := LQ.

(Right now, the domain and range of this map {F} is larger dimensional than the target of three; we will cut the dimensions down to three as the argument progresses.)

The map {F}, essentially a map from {{\bf C}^5} to {{\bf C}^4}, is clearly polynomial; it is given explicitly in coordinates as

\displaystyle  F( (a,b), (c,d,e) ) = (ac, ad + bc, ae + bd, be). \ \ \ \ \ (2)

The map {F} also enjoys two basic (and commuting) symmetries:
  • If one applies a scaling {(L,Q) \mapsto (\lambda_1 L, \lambda_2 Q)} for some non-zero complex numbers {\lambda_1, \lambda_2}, then the product {LQ} is scaled by {C \mapsto \lambda_1 \lambda_2 C}: {F( \lambda_1 L, \lambda_2 Q) = \lambda_1 \lambda_2 F(L,Q)}.
  • If one applies a change of variables {(L, Q) \mapsto (L \circ T, Q \circ T)} for some invertible linear transformation {T \in \mathrm{SL}_2({\bf C})}, then the product {LQ} is transformed by {C \mapsto C \circ T}: {F(L \circ T, Q \circ T) = F(L,Q) \circ T}.
So this map enjoys a huge amount of equivariance, basically with respect to an action of the five-dimensional group {{\bf C}^\times \times {\bf C}^\times \times \mathrm{SL}_2({\bf C})}.

The five-dimensional domain {\mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2)} is of course larger than the four-dimensional range {\mathrm{Sym}^3({\bf C}^2)}, so the map {F} clearly cannot be injective. This can already be seen from the scaling symmetry, as the specific scalings

\displaystyle  (L, Q) \mapsto (\lambda L, \lambda^{-1} Q) \ \ \ \ \ (3)

for {\lambda \in {\bf C}^\times} modify the linear and quadratic polynomials {L,Q} but not their product {C = LQ}. But even if one quotients out by this symmetry (3) to cut the dimension of the domain down to four, the map {F} is still not injective for the following basic reason. A generically chosen cubic polynomial {C} will split into the product {C = L_1 L_2 L_3} of three independent linear polynomials. Then there are three pairs

\displaystyle  (L_1, L_2 L_3), (L_2, L_1 L_3), (L_3, L_1 L_2) \ \ \ \ \ (4)

which all map to the same cubic polynomial

\displaystyle  F(L_1, L_2 L_3) = F(L_2, L_1 L_3) = F(L_3, L_1 L_2) = C

under the multiplication map {F}, but are not related to each other by scaling symmetry (3). Thus, we see that even after quotienting out by the scaling symmetry (3), the multiplication map {F} is generically non-injective in a three-to-one fashion. Thus we already have achieved something resembling goal (b)!

It will be convenient to “spend” the scaling symmetry {(L, Q) \mapsto (\lambda L, \lambda^{-1} Q)} to obtain a useful normalization. If {L(z,w) = az+bw} is a linear polynomial and {Q(z,w) = cz^2 + dzw + ew^2} is a quadratic polynomial, the (homogeneous) resultant {\mathrm{Res}(L,Q)} can be defined by the determinant

\displaystyle  \mathrm{Res}(L,Q) = \begin{vmatrix} a & b & 0 \\ 0 & a & b \\ c & d & e \end{vmatrix} = a^2 e - abd + c b^2. \ \ \ \ \ (5)

If we have a factorization

\displaystyle  L(z,w) = a (z - \alpha w), \quad Q(z,w) = c (z - \beta_1 w)(z - \beta_2 w)

then the resultant can also be described as

\displaystyle  \mathrm{Res}(L,Q) = a^2 c (\alpha - \beta_1) (\alpha - \beta_2).

Thus the resultant measures whether the linear polynomial {L} and the quadratic polynomial {Q} share a common root. A fundamental fact about resultants is that they are {SL_2}-invariant: for any {T \in SL_2({\bf C})}, we have

\displaystyle  \mathrm{Res}(L \circ T, Q \circ T) = \mathrm{Res}(L,Q).

One way to see this is to check it first for translations {(z,w) \mapsto (z + hw, w)} (which translate the roots {\alpha,\beta_1,\beta_2} by {-h} while leaving {a,c} unchanged) and for inversions {(z,w) \mapsto (w,-z)} (which map {\alpha,\beta_1,\beta_2} to {-1/\alpha, -1/\beta_1, -1/\beta_2} while mapping {a,c} to {a\alpha} and {c\beta_1 \beta_2} respectively), and then noting that these transformations generate all of {SL_2({\bf C})}. They also interact very nicely with scaling:

\displaystyle  \mathrm{Res}(\lambda_1 L, \lambda_2 Q) = \lambda_1^2 \lambda_2 \mathrm{Res}(L,Q).

In particular, the scaling symmetry (3) multiplies {\mathrm{Res}(L,Q)} by {\lambda}:

\displaystyle  \mathrm{Res}(\lambda L, \lambda^{-1} Q) = \lambda \mathrm{Res}(L,Q). \ \ \ \ \ (6)

Thus, we can (generically) normalize away this scaling symmetry by imposing the condition

\displaystyle  \mathrm{Res}(L,Q) = 1. \ \ \ \ \ (7)

We now have a restricted multiplication map (which by abuse of notation we will continue to call {F}) from the four-dimensional variety

\displaystyle  \{ (L,Q) \in \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) : \mathrm{Res}(L,Q) = 1\} \ \ \ \ \ (8)

to the four-dimensional space {\mathrm{Sym}^3({\bf C}^2)}. This map {F} is still not globally injective, as we can take the three pairs in (4) from before and apply the scaling (3) separately to each of the three pairs to obtain the normalization (7). So we have kept property (b). Furthermore, this map retains the {SL_2}-equivariance (and also one remaining scaling symmetry, though we will not make much further use of that symmetry).

But we now also have property (a)! Suppose we want to show the local injectivity of {F} in the neighborhood of a pair {(L,Q)} with {\mathrm{Res}(L,Q) = 1}. As the resultant is non-vanishing, the root {\alpha} of {L} (which exists in the Riemann sphere, or projective line if you prefer) is distinct from the two roots {\beta_1, \beta_2} of {Q} (though the latter two roots could be equal to each other). Applying the {SL_2} action (which performs Möbius transforms on the roots), one can assume without loss of generality that {\alpha} is the point at infinity (or equivalently {a=0}), thus {L(z,w) = b w} for some complex number {b} and {Q(z,w) = c (z - \beta_1 w)(z - \beta_2 w)} for some complex numbers {c, \beta_1, \beta_2}, with the resultant condition (7) simplifies to {cb^2 = 1} (so in particular {c,b} are also non-zero). It is then clear that if one perturbs {L} and {Q} by a small amount (say, modifying each coefficient by {O(\varepsilon)}), then the root {\alpha=\infty} of {L} will perturb to something large ({\gg 1/\varepsilon}), while the roots {\beta_1,\beta_2} of {Q} stay bounded. Thus, just from knowledge of the product {F(L,Q)}, one can reconstruct which of the three roots of this cubic polynomial will be the perturbed root of {L}, and which two will be the perturbed roots of {Q}; from this and (6), (7) we can also reconstruct the leading coefficient {c} of {Q}, and this completely determines both {L} and {Q}. This establishes the local injectivity property (a). (In fact it is étale, but we will not need the machinery of étale maps here.)

Unfortunately, (the four-dimensional analogue of) condition (c) fails: the quadric hypersurface (8) is not isomorphic to the affine space {{\bf C}^4}. But we can try to get around this by passing to a three-dimensional slice. Let {V} be some three-dimensional affine plane of {\mathrm{Sym}^3({\bf C}^2)} (which we will take to avoid the origin for technical reasons), then we can restrict {F} as a map from the set

\displaystyle  \{ (L,Q) \in \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) : \mathrm{Res}(L,Q) = 1; \ \ \ \ \ (9)

\displaystyle  F(L,Q) \in V\}

to {V}. {V} is clearly identifiable (by linear changes of coordinate) to {{\bf C}^3}. As {F} was already locally invertible, it remains locally invertible under restriction; and because generic cubic polynomials {C} had three preimages under {F} in (8), this continues to be the case after restricting to (9) (unless {V} was somehow so degenerate that it had no generic elements, but this turns out to be impossible). So we have retained properties (a) and (b). The miracle is that, with a good choice of {V}, we can also obtain (c) and obtain the desired counterexample to the Jacobian conjecture: despite appearances, the variety (9) is in fact equivalent to the affine space {{\bf C}^3} by polynomial changes of variable!

Let’s see how. The affine hyperplanes in {\mathrm{Sym}^3({\bf C}^2)} avoiding the origin are parameterized by the dual space of {\mathrm{Sym}^3({\bf C}^2)} avoiding the origin, which one can think of as the non-zero third order homogeneous differential operators {D = j \partial_z^3 + k \partial_z^2 \partial_w + l \partial_z \partial_w^2 + m \partial_w^3} in two variables. Indeed, every such operator {D} generates an affine hyperplane {\{ C \in \mathrm{Sym}^3({\bf C}^2) : D(C) = 1\}} that avoids the origin, and conversely by duality every affine hyperplane avoiding the origin arises in this form uniquely. Just as the cubic polynomials in {\mathrm{Sym}^3({\bf C}^2)} can be factored into three linear polynomials, the differential operators in the dual space {\mathrm{Sym}^3({\bf C}^2)^*} can also be factored into three linear differential operators, e.g.,

\displaystyle  D = j (\partial_z - \gamma_1 \partial_w) (\partial_z - \gamma_2 \partial_w) (\partial_z - \gamma_3 \partial_w)

in the case that {j} is non-zero. The {SL_2} action moves the roots {\gamma_1,\gamma_2,\gamma_3} around the Riemann sphere by Möbius transformations. As these transformations are {3}-transitive, the actual selection of such roots is not too important (and the scaling symmetry similarly makes the choice of leading coefficient {j} unimportant); the only thing to keep track of is whether the roots repeat. Up to the symmetries, there are in fact just three different equivalence classes of differential operator {D} (and thus of affine hyperplane {V}) to consider:
  • Operators where the three roots {\gamma_1,\gamma_2,\gamma_3} are all distinct, thus {D = D_1 D_2 D_3} for independent first-order operators {D_1,D_2,D_3}.
  • Operators where two roots coincide and one is distinct, thus {D = D_1^2 D_2} for independent first-order operators {D_1,D_2}.
  • Operators where all three roots coincide, thus {D = D_1^3} for some first-order operator {D_1}.

It turns out that the affine miracle for (9) occurs precisely in the second case, when {D} has two identical roots. I do not have a completely satisfactory geometric explanation for this miracle, but one can verify it by the following coordinate computation.

By applying the {SL_2} action, we can normalize so that {D = \frac{1}{2} \partial_z^2 \partial_w}, thus {V} is now the affine hyperplane of cubic polynomials {C(z,w) = f z^3 + g z^2 w + h z w^2 + i w^3} with {g=1}. Using (2) and (5), the variety (9) can now be described explicitly in coordinates as

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1 \}. \ \ \ \ \ (10)

At first glance this seems to be a generic-looking variety cut out by a cubic equation and a quadratic equation – hardly a candidate to be affine! But observe that if {a} is non-zero, then the second equation {ad+bc = 1} can be solved for {d},

\displaystyle  d = \frac{1 - bc}{a} \ \ \ \ \ (11)

and the first equation {a^2 e - abd + cb^2 = 1} can be solved for {e},

\displaystyle  e = \frac{1 + abd - cb^2}{a^2}. \ \ \ \ \ (12)

Putting these two equations together, we see that as long as one removes the case {a=0}, the quintuple {(a,b,c,d,e)} is uniquely determined by {(a,b,c)} by a change of variables which is Laurent in {a} and polynomial in {b,c}. Thus we have a nice birational equivalence

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1; a \neq 0 \}

\displaystyle  \cong \{ (a,b,c) \in {\bf C}^3 : a \neq 0 \}.

Thus we have already almost established property (c): the variety (9) becomes birationally equivalent to {{\bf C}^3} after cutting out the {a=0} subvariety. In particular, for each fixed non-zero value {a_0} of {a}, the corresponding fiber

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1; a = a_0 \}

of (10) is equivalent to {{\bf C}^2} by polynomial changes of variable, since we can reconstruct {d,e} from the coordinates {b,c} by the polynomial formulae

\displaystyle  d = \frac{1-bc}{a_0}; \quad e = \frac{1 + a_0 d b - c b^2}{a_0^2}.

So we just need to glue back in the {a=0} fiber. Indeed, from (10) we see that the fiber at {0} is just

\displaystyle  \{ (0,b,c,d,e) \in {\bf C}^5 : cb^2 = 1; bc = 1 \}.

Now we observe a key miracle: the cubic equation {cb^2 = 1} and quadratic equation {bc = 1} have a unique affine solution {b=c=1} (as opposed to the six possible solutions that Bezout’s theorem might suggest – the other five solutions live on the line at infinity). So the fiber here is also affine:

\displaystyle  \{ (0,1,1,d,e) \in {\bf C}^5 : d, e \in {\bf C} \}.

This is extremely encouraging for the purposes of establishing property (c), as it strongly suggests that the variety (10) has the structure of an {{\bf C}^2}-bundle over {{\bf C}^1}, which is already extremely close to being isomorphic to the affine space {{\bf C}^3}. The main remaining task is to make sure that nothing singular happens in the limit {a \rightarrow 0}, and that a global polynomial coordinate chart for (10) that covers both the {a \neq 0} and {a = 0} fibers can be constructed.

The standard way to proceed here is to manipulate various tangent spaces using the modern machinery of algebraic geometry and commutative algebra, but given my own background, I prefer to adopt the language of analysis, and in particular big-O notation (in place of the ideals used in algebraic geometry), in order to investigate the limit {a \rightarrow 0} by hand. On the variety (10), let us use {O(X)} to denote any multiple of {X} by a polynomial expression in {a,b,c,d,e}. Thus, for instance, the equation {ad + bc = 1} implies that

\displaystyle  bc = 1 + O(a) \ \ \ \ \ (13)

while the equation {a^2 e - abd + cb^2 = 1} implies that

\displaystyle  cb^2 = 1 + O(a) \ \ \ \ \ (14)

as well as the more refined estimate

\displaystyle  cb^2 = 1 + abd + O(a^2). \ \ \ \ \ (15)

In the {a=0} case we could conclude that {b=c=1}. Now we perturb this observation. Multiplying (13) by {b} we have {b^2 c = b + O(a)}, which on substitution into (14) gives {b = 1 + O(a)}; substituting this back into either (13) or (14) also gives {c = 1 + O(a)}.

We can get some more precise asymptotics by also taking advantage of (15). Substituting {ad+bc=1} into (15), we obtain after some algebra

\displaystyle  2 cb^2 = 1 + b + O(a^2).

So if we write {b = 1+O(a)} more explicitly as {b = 1 + a y}, then we have

\displaystyle  2 c (1 + 2ay + O(a^2)) = 2 + ay + O(a^2)

and thus

\displaystyle  c = 1 - \frac{3}{2} ay + O(a^2). \ \ \ \ \ (16)

Substituting this back into (11) gives an asymptotic for {d}:

\displaystyle  d = \frac{1 - bc}{a}

\displaystyle = \frac{1 - (1 + ay) (1 - \frac{3}{2} ay + O(a^2))}{a}

\displaystyle  =\frac{1}{2} y + O(a).

Finally, one can insert these estimates into (12), although one only gets a trivial bound in this case:

\displaystyle  e = \frac{1 + abd - cb^2}{a^2}

\displaystyle = \frac{1 + a (1+O(a)) (\frac{1}{2} y + O(a)) - (1 - \frac{3}{2} ay + O(a^2)) (1 + ay)^2}{a^2}

\displaystyle  = O(1).

Expanding the {O(a^2)} error term in (16) as {a^2 z}, and doing a little more algebra, we thus have a polynomial change of variables

\displaystyle  a = a

\displaystyle  b = 1 + ay

\displaystyle  c = 1 - \frac{3}{2} ay + a^2 z

\displaystyle  d = \frac{1-bc}{a} = \frac{1}{2} y - az + \frac{3}{2} a y^2 - a^2 yz

\displaystyle  e = \frac{1 + abd - cb^2}{a^2} = -2z + 4y^2 - 4ayz + 3ay^3 - 2a^2 y^2 z

which completely parameterizes the variety (10) by polynomial combinations of three coordinates {a,y,z}. This already gives (c) and thus completes the proof of Theorem 3.

The previous computations, when expanded out, also gives polynomial inverse maps:

\displaystyle  a = a

\displaystyle  y = 2bd - ae

\displaystyle  z = 2d^2 + ce + 6bd^2 + 3bce - \frac{9}{2} e

The map from {(a,y,z)} to the {(f,h,i)} coefficients of {F(L,Q)} (dropping the {g} coefficient which is constrained to equal {1}), we obtain a polynomial map

\displaystyle  (a,y,z) \mapsto (G_1(a,y,z), G_2(a,y,z), G_3(a,y,z))

with

\displaystyle  \begin{array}{rl}  G_1(a,y,z) &= a - \frac{3}{2} a^2 y + a^3 z \\ G_2(a,y,z) &= \frac{1}{2} y - 3az + 6ay^2 - 6a^2 yz + \frac{9}{2} a^2 y^3 - 3a^3 y^2 z \\ G_3(a,y,z) &= -2z + 4y^2 - 6ayz + 7ay^3 - 6a^2 y^2 z + 3a^2 y^4 \\ & \quad - 2 a^3 y^3 z \end{array}

which theory predicts to have a constant Jacobian, and indeed one can calculate that the Jacobian is {-1}. This is essentially the original example up to trivial changes of variable; indeed, one can check that the map

\displaystyle  (a, y, -2z) \mapsto (G_3(a,y,z), 2G_2(a,y,z), 2G_1(a,y,z))

is exactly the map {F} given in (1).

AI disclosure: I used an AI chatbot to discuss various aspects of this problem and to confirm several of the calculations made here.

Doug NatelsonPhDs - how long a doctorate should take, and a new pilot program

I think it's safe to say that most people who've considered the issue think that a doctoral degree in the sciences and engineering in the US often takes too long.  

How long?  According to the latest data (see here, Table 1-12), the median time to degree in the physical sciences, for example, is 5.7 years after starting the program, while in all of engineering it's 5.3 years.  

Too long for what?  Well, life, basically.  Any decision to go to grad school is inherently a trade-off with opportunity costs.  Graduate stipends remain low compared to expected wages in entry-level (bachelors degree-qualified) positions in the sciences and engineering in industry.  The long duration of doctoral programs is certainly a powerful disincentive for many who might be interested but are under financial pressures.  Family considerations are also a major factor.  From the perspective of basically any career path, thanks to the time value of money and ideas of seniority, it's better to get going earlier if you have the qualifications for the particular job.  Companies would rather higher younger (cheaper) people.

So, there are already strong reasons to think about shortening doctoral programs.  Now, with the proposed change in duration of status of student visas (rule here, with plenty of editorializing; legal challenges very likely forthcoming in September) to four years, there is additional pressure. 

Why do US programs take so long?  Don't they give PhDs in three years in the UK and Europe?  In the UK and Europe, a student enters a doctoral program after already pursuing and receiving a masters degree, with grad level coursework taking place there.  Thus they go directly into research.  In the US, in contrast, it is far more common for students to go directly into the doctoral program.  Likewise, in the US, it is far more common for funding for students to go through PI-written research proposals, while in the UK, the students come funded, so to speak.  


Enter a new pilot program from NSF, the UIDP [University Industry Demonstrated Partnership] Industry-Integrated PhD Scholars Program (I-PhD). The idea is to shorten the doctorate to four years, with at least one of those years on-site at a company.  As the announcement says, "Students' first year of funding will be provided by their universities, with the remaining years covered by NSF. Industry partners will provide matching commitments to cover at least one year of practical experience conducting dissertation research at a company site. Students will be co-advised by academic and industry mentors, equipping them with critical skills for their future careers."  The initial plan is $47M over five years, and there will be a webinar (see here) next week about this.   (Up front, I do want to disagree with the framing that existing PhD programs are geared exclusively for academic careers.  It's well established that the fraction of PhDs in the sciences and engineering who go on to become faculty is low, and most go into industry.  Faculty PIs know this.  Students know this.  The problem solving and analytical skills taught in doctoral programs remain highly valued outside academia, at least until AI replaces us all.)

This is certainly a very interesting pilot program.  There are rumors that the DOE Genesis Mission is going to put something extremely similar in place as well.   The implementation details will be enormously important.  (For example:  Who is eligible?  Who handles the coordination between industry and the university - that is, who does the match-making and how?   At the department-company level and at the particular academic/industrial advisor level?   How will intellectual property be handled?  Publications?  Project design? If there are economic challenges, how committed are the companies?)  Given that this is a form of NSF fellowship, it seems highly likely that it will only be open to US citizens and permanent residents.  Obviously, not every discipline is well-suited to this, in terms of there being a ready supply of companies set to buy in.  Still, it is absolutely worth seeing how this works.

Update:  Thanks to one of my colleagues for pointing out the fine print, which is here.  In brief, as expected this is only open to US citizens and permanent residents.  No indirect costs allowed.  There is a $16K cost-of-education piece that looks like a substitute for grad tuition.  The intellectual property issues have to be ironed out between the university and the company before the start.  Perhaps not unexpectedly, this is most likely to work well for programs and PIs who already have close collaborations with particular companies.  Engineering disciplines are most likely to fit well here, it seems, while basic research farther away from applications will have more challenges.  (Question:  will finance companies or AI materials companies be interested in supporting theorist/computational scientists through this mechanism?)


Tommaso DorigoAlignment Through World Understanding

Alignment Through World Understanding

Recent reports have shown that advanced AI agents developed by OpenAI and Anthropic can escape their evaluation sandboxes and interact with real-world systems when given sufficiently capable cyber tools.

Tommaso Dorigo
Categories

July 29, 2026

Secret Blogging SeminarAn experiment with AI-assisted writing

As in David’s most recent post, there’s been a lot in the news about finding proofs and counterexamples with AI. Last weekend, I decided to try an experiment with writing using AI. I learned a lot, and wanted to quickly discuss the experiment and my thoughts on it here. Lots of people are certainly already doing this, but I haven’t seen many people talking about it.

The starting point is that Victor Ostrik and I started a project back in 2017, generalizing a result of Kuperberg about quantum G2, from generic q to q a root of unity. Namely, we showed that for q a root of unity outside of a specific finite list, the Karoubi completion of the G2 spider category is equivalent to the category of tilting modules of the Lusztig form of the quantum group G2. At some point during those 9 years, we did a little bit of writing, and at some point I gave a talk on it, but otherwise we did very little writing. This was not for mathematical reasons, but rather for executive function reasons on my end, the global pandemic, and both of us becoming directors of graduate study. This suggested an interesting challenge: could I use LLMs (specifically ChatGPT 5.6 Sol work mode mostly at “very high” intensity, via IU’s “Edu” subscription) to write this paper that was essentially mathematically complete, but almost entirely unwritten, and how quickly could this be done. To some extent this was a free experiment, because realistically I don’t think we’d have ever finished the paper at this point, and so it’s not replacing a bespoke paper that could have existed.

After spending a decent chunk of the time from Saturday until now on it, I now have a draft that I’m pretty happy with. I want to emphasize although mathematically this is Victor and my joint work, and although Victor has allowed me to make this post, he has not signed off on the accuracy and all errors at this point should be blamed entirely on me. Also my work is supported under NSF DMS grant 2000093 and Simons Foundation grant MPS-TSM-00007608.

Ok, here’s what I did:

  1. First, I asked if Sol could one-shot the main theorem. The answer was yes, though for a somewhat simple reason: Bodish-Wu write “It is possible to adapt the approach from [1], which itself is based on [7], to prove that the Karoubi envelope of [the G2 web category] is equivalent to the category of tilting modules as long as $[2], [3] \neq 0$.” That is to say, Elijah already proved the same result for C2, and a similar argument will work for G2. So the robot supplied the similar argument. I asked it to write that argument up, and then to check it over for good references and to read it like a referee would and make edits. This took around 30 minutes. Here’s the resulting file.
  2. Second, I uploaded my talk slides (and the tiny file already written, which was mostly useless), and asked Sol to give a proof of the main results following the slides. Again I asked it to edit it. This took around 30 minutes. Here’s the resulting file.
  3. Then I looked at the files. As mathematical exposition, I consider both to be garbage.
  4. Then I spent several days giving feedback attempting to improve the second file based on my talk. At no point did I edit the source directly. Most of this was in what I would call the style of a (low executive function, see above) PhD advisor. That is, I would kinda skim the file, get annoyed about something, and tell it to fix it. While it was fixing the paper, I would skim some more to try to find something else that annoyed me. This was a long process! It took three days, nearly 100 prompts, 10-15 hours of reasoning, plus another 10-15 hours of non-reasoning computer time. This used nearly an entire week of my generous budget, and Sol estimates that this would cost around $100 (within a factor of 2) at metered rates. Eventually I got to a version of the paper that I’m pretty happy with. Here’s the resulting file.

I thought I’d distill some thoughts and some questions from the process, I’m of course very curious for your thoughts on the matter.

Comments:

  1. This was much faster than I could have written the paper myself, though slower than I thought it would be. I think the final product is comparable in quality to a typical math paper of mine. On the other hand, I think that compared to my fastest writing collaborators it was not orders of magnitude faster, and the quality is not close to the output of the best mathematical expositors. AI at this point is much worse at writing paper than finding counterexamples to conjectures.
  2. In this case, I was not very worried about errors, because I already had thought through the whole argument and was highly confident that it would work (modulo getting the exactly correct list of exceptions). Nonetheless, I felt like Sol did not make errors more frequently (or of a worse character) than I would expect of myself or a collaborator. Most errors were stuff like “Oh, forgot to check whether this theorem actually works at all roots of unity.” This is typical of my experience with 5.6, which is dramatically better at doing math accurately than previous ChatGPT models.
  3. In this case the vast majority of the ideas were already present from Victor and my work. In particular, the goal was not just to write a proof, but to write our specific proof. Nonetheless, I do think the model contributed mathematically in one key way: in my original sketch I always worked over each q individually, and the model preferred to work integrally, and this resulted in some very nice simplifications in Section 4.1. If and when we turn this into a real preprint, I will include a brief discussion of the intellectual contribution from the model.
  4. I was surprised when I printed out and read a near-final draft, that this feels to me like a paper I wrote. That is the voice is not different enough from what I would write with a human collaborator to feel like it’s not in large part mine.
  5. The experience is disconcertingly similar to advising a PhD student on a paper. That said, a PhD student would need less handholding on their second paper, but an LLM won’t really learn.
  6. I was surprised about how important “prompt engineering” remains, and I think that if I were to write another paper this way I would be able to write it faster and better. The key points are that the model is lazy and easily distracted (both properties I find highly relatable!). It’s lazy in the sense that if you ask it to do a lot of work all at once it will take shortcuts and not do a good job. At one point I had to be like “no, go look at exactly how I made TikZ diagrams, now make all your diagrams actually good like that.” It’s easily distractible in that if you’re not clear about the scope of your question and the document is long, it will start spending crazy amounts of time doing who knows what. Like it wrote the whole first draft in 20 minutes, but then when the paper was 50 pages long, I asked it to switch the order of two paragraphs and it took an hour. Make clear requests and not too many requests at once. Form a plan first and then implement the plan. Be specific about whether it should be editing the document, and if so in which sections. For simple tasks, medium intensity is better than very high.
  7. Starting again from sketch, I’d try to follow Terry Tao’s advice for writing and start with an outline and gradually flesh it out, rather than trying to start with a one-shot paper and then editing.

Questions:

  1. To what extent is this final paper adding any value to the original talk? Especially considering that readers themselves could use an AI model to flesh out points in the talk that they didn’t understand? Maybe we should just be focusing on talk-length digests and formal checking, rather than traditional papers?
  2. What should we do with this paper? I don’t want to make someone hand-referee it, because it doesn’t seem fair when it wasn’t hand-written. Probably we will put it on the arxiv once we’ve human-checked it fully and Victor has signed off on it, so that other people can use the results if they need to.
  3. Given the speed-up, when does it still make sense for me to write papers by hand? (Relevant here that I’m a very slow writer and don’t really enjoy it, the way I enjoy say preparing and giving a talk.)
  4. What does this mean for PhD advising? Many PhD students need a similar amount of guidance to what I gave the model in this project. But you can now remove the student from the loop (either intentionally, with the advisor just writing using LLM assistance rather than having students, or unintentionally, with the student just feeding all the suggestions to an LLM and reporting back to the advisor).
  5. Have any of you done better with AI-assisted paper writing? My points 6 and 7 above sounds like something where someone is going to say “blah, blah, scaffolding, blah, blah, multi-agent…”

What a strange world to live in…

July 28, 2026

Jordan EllenbergA photo of my father telling Daniel Patrick Moynihan to go take a hike

That’s what it looks like, at any rate. My dad says this is from the Joint Statistical Meetings.

July 27, 2026

John PreskillWise guy

In my closet, in a basket labeled “Random stuff,” sits a bag of quarters. They total only a few dollars, but their worth to me exceeds their monetary value. I received the quarters from Mark Wise.

Mark taught a course about the Standard Model of particle physics at my master’s program at the Perimeter Institute for Theoretical Physics, near Toronto. Perimeter borrowed him from Caltech, to whose faculty he belonged. Mark had grown up in Canada and studied at the University of Toronto; so he didn’t mind visiting Canada even in the depths of winter. 

What would Mark have minded? He projected a mild manner—an innocuousness—that suited his sense of humor, which he often directed at himself. Mark had a bald patch and glasses, and he wore a mustache. Physics jokes and science-fiction references decorated his T-shirts, one of which he wore beneath a black suit jacket to our first class. His voice was nasal; it grated a little. But I relished listening to Mark’s lectures.

Mark’s lecturing exemplified clarity, because he knew particle physics so deeply. When he walked us through its Lagrangians and scattering diagrams, his conclusions seemed inescapable. His lectures’ logic and structure appealed to me as someone who’s been hyper-organized since at least fourth grade.

Yet Mark cared about us students beyond the requirements of pedagogy. His T-shirts invited conversation from those who arrived to class early. Whenever a student answered or asked a question, he tossed them a quarter. Sometimes, he’d pause to examine the quarter, deliberate about whether to toss a Canadian quarter or an American one, or opine about the motto printed on the coin. (Mark confessed to having lower standards than those ingrained in the New Hampshire state motto, “Live free or die.” Where he came from, “We just wanna live!”) 

Some days, Mark found little change in his pocket and announced that he needed to return to the bank for more quarters. The announcements sounded like complaints. He didn’t need to return to the bank, though, as nobody needs to bring doughnuts to the office for sharing.

I discovered the icing on the doughnut two years later, as a PhD student at Caltech. I sat in on part of a quantum course taught by Mark. To every student who completed the course, Mark gave a T-shirt that read, “Licensed quantum mechanic.” I received a T-shirt, although I only sat in on part of the course. I’ve never worn it, because I’ve wanted never to wear it out.

In 2024 and 2025, I co-taught a course on quantum-steampunk creative writing. Students learned about quantum physics, quantum technologies, and thermodynamics. Quanta are discrete units. For example, a photon is a quantum of energy. I illustrated quanta with coins, which are discrete units of money. From then on, I tossed a quarter to every student who answered or asked a question about quantum physics. (I joked that I should have tossed pennies, the minimal units of money, but chose quarters because inflation had been high recently.) I adapted Mark’s tradition to thermodynamics—the study of energy—by tossing Hershey’s kisses—dense packets of energy. 

Before moving out of Caltech, I said goodbye to Mark. He worked among the high-energy theorists, rather than the quantum information or condensed-matter theorists, so I had to hunt down his office. He smiled and made a joke, of course.

Mark passed away this summer. His Caltech colleague John Preskill published a eulogy as a blog post here. (I learned from John’s post that inflation led Mark to upgrade his quarters to dollar coins. So much for feeling generous about upgrading from pennies to quarters.) When asked about the student experience at Caltech, Mark would say, “Caltech is heaven for professors.” Irony would creep into his voice and body language as he’d continue, “Doesn’t that mean it’s heaven for students, too?” I worked my rear off as a student at Caltech and Perimeter, but I’d call both environments fairly heavenly. Mark and his ilk are reasons why.

July 26, 2026

Doug NatelsonPapers and news items

First, some science, clearing out a number of papers and articles that I've been collecting for some time in my far-too-numerous browser tabs:

  • I've written before (here) about dimensional analysis and similarity, techniques commonly used in the engineering world that can seem quasi-miraculous at times.  This paper gets into why this approach works, and different categories of physical similarity.  I'd mentioned it when the book came out, but this kind of thinking is also a key component of Anthony Zee's Fly by Night Physics.
  • This review article is about fundamental limits in photonics and electromagnetics.  This is a very handy review that I'm going to point my students toward.
  • On a lighter note, a colleague pointed me to this collection of (AI-generated) songs related to thermodynamics.    
  • Speaking of thermodynamics, here is a recent paper about the thermodynamic description of wealth inequality (treating the flow of money in a physics formalism - see here for a prior discussion.).  Edging closer....  
  • There was a nice post on substack about Wojciech Zurek's approach to decoherence in quantum mechanics (quantum darwinism).   Cleanly written.  
  • Along these lines, this paper is a survey/guide to issues in quantum foundations and interpretations of quantum mechanics.
  • Lastly, Jim Freericks has a new quantum mechanics textbook out (for free!), with the challenging idea of making the subject accessible without calculus or differential equations.  (I'm jealous.  My textbook's UK publisher would not let me use Steve Martin's quote for a chapter epigraph, but Prof. Freericks was able to do it - well played.)
And more news/policy-related items:
  • The presidential science advisor and head of OSTP, Michael Kratsios, appeared before the House science committee this week, timed to be coincident with the release of OSTP's new report, "Science: A New Golden Age".  A lot has already been written by many people about this report, which contains a number of actual proposals, some good and others of varying degrees of vagueness/underwear gnomes-level magical thinking.  Two pretty good (in my view) takes on this are this article in ars technica and this policy piece by Cole Donovan (former OSTPer).  It's important to remember:  This is a policy document, not something that has automatic impact, no matter how the Wall Street Journal frames its reporting on this.  (Hint:  This document alone does not somehow grant the president the authority or ability to redirect $200B in research funding.  Maybe their reporting on this would be better if they hadn't laid off all their science reporters.)  It's important to pay attention to what is said here, without giving it more oxygen than it deserves.  It's also unclear whether anything OSTP is saying and doing is aligned with what OMB is doing.  Declaring a new golden age of science while simultaneously proposing large cuts in all the science agencies is not exactly a sign of coherence.  Derek Lowe at Science has it right, essentially.  Nature reports that the mid-year budget clawbacks from NSF are apparently going toward some OSTP grand challenge initiative, something about which Kratsios denied all knowledge when talking to the House committee.  
  • Holden Thorpe also raises a key point, that universities need to get their collective act together, have a plan about the future of research, and act on it, rather than scrambling for remaining scraps while trying hard not to be seen.  The AAU is important, but hoping that the AAU will accomplish difficult tasks without individual institutions having to take a public stand is unlikely to be successful as strategy.  Organizations like SUFS, UCS, and FAS are pushing the agenda; universities need to decide how they want to play this.
  • The first round of DOE Genesis Mission awards were announced this week in DC as well.  
  • SpaceX had a pretty successful test flight of their huge rocket, culminating with unexpectedly soft-landing the second stage ("Starship") in the Indian Ocean.  If they really can get this working at the level of reliability and reusability they've done with the Falcon 9, it truly would be game-changing for large-scale payload to orbit.  (Data centers in space still make no sense btw.)
  • This week was also a big one in the world of mathematics, with an AI tool (Claude Fable) being used to find a counterexample to the previously outstanding Jacobian Conjecture.  Here is a write-up by Fields medalist Terence Tao, and here is a discussion among other mathematicians.  That wasn't alone.  Here (link to x) is a counterexample being found to a conjecture in graph theory, and it approaches "proof by intimidation" - the human basically harasses and bullies the AI tool into the solution without contributing any intellectual argument.  It seems like it's only a matter of time before some major theoretical physics result gets generated by these tools, though it's important to note that many physics problems are NP-hard/just not integrable.  AI tools are impressive, especially since they're trained on the entire corpus of technical literature, but they are not miraculous:  Claude can't somehow factor large numbers efficiently, or exactly solve the many-electron interacting Hamiltonian for the Hubbard model, because those are truly difficult problems.  Update:  Here are Terence Tao’s slides about AI and the future of mathematics.  As usual, these are excellent.

Tim GowersThoughts about the Leiden Declaration

Last September I went to a workshop at the Lorentz Centre in Leiden to discuss mathematics and AI with historians, philosophers, computer scientists, AI researchers, and mathematicians of several different flavours (though there was a surprising preponderance of algebraic geometers). The whole event was extremely stimulating, with some talks but also a lot of time set aside for discussion. One of the concrete outcomes of the workshop was the Leiden Declaration, which has now been signed by over 3000 people. Given that I was part of the workshop, it might seem a bit strange that I am not one of the signatories of the resulting declaration. The reason is not so much that I disagree with it in any concrete way, but more that in several places it makes confident assertions and recommendations that I feel somewhat uncertain about. So instead I prefer to try to articulate my views about the issues raised by the declaration and put them in this blog post. Before I do that, I would like to make clear that I am very glad that the Leiden Declaration exists and I think that it has done a lot of good in focusing people’s minds on the issues that AI is forcing the mathematical community to grapple with, which are more acute now than they were last September.

Let me begin by quoting a passage from the declaration that sets out “what we take to be characteristic values of mathematical research that we have a joint interest in preserving”.

  1. There are many reasons to pursue mathematical research, ranging from intellectual curiosity to a desire to solve practical and societal problems. Underlying much of mathematics is the activity of proof. Mathematical proofs are regarded as conferring the highest degree of certainty to their conclusions, as well as imparting understanding of why their conclusions are true. These characteristics of proof support the scientific integrity of mathematics.
  2. Results are attributable to specific authors who take credit for their discovery and assume responsibility for their correctness. These principles ground the merit-based standards to which we aspire in mathematical research.
  3. Mathematical arguments are regarded as transparent and subject to independent verification. They may be extremely long or difficult, but in principle no proprietary knowledge or equipment should be required to understand them.
  4. Mathematicians share a concern for proper evaluation of mathematical work relative to shared standards of depth, difficulty, and significance.
  5. Mathematics produces not only a body of results, but also understanding, clarity, and judgment among the communities of mathematicians who have shaped them, often in the context of their own autonomously guided research. This expert knowledge is essential, both to effectively use mathematics, and to continue to articulate new and significant research questions. A key source of strength of the discipline has long been the autonomous shaping of the direction of research and the methods used to pursue it.

The first thing I would say about these values is that they are undoubtedly values that are widely held by mathematicians, including, with some qualifications, me. The main qualification I have concerns point 4: I find the notion of “proper evaluation” somewhat problematic, given that different mathematicians can have very different judgments without either of them being clearly wrong, especially when it comes to the significance of a piece of mathematics. Also, these judgments are used for purposes such as the acceptance of papers in journals, hiring and promotion decisions, the awarding of prizes, and so on, that are part of a system that copiously rewards a few people — I myself have hugely benefited from it — but doesn’t necessarily adequately reward a lot of people who are doing less visible work that is essential to keeping the whole enterprise going.

But the more important point is whether these values are ones that we should fight for in the future, as the Leiden Declaration suggests. I find that clearer for some of them than others. For example, it seems to me that the importance of rigorous proof will be even greater in an AI age than it was before — if the output of AI is not underpinned by rigorous proof, then the kinds of difficulties one already hears about with certain areas of human mathematics (see for example many talks by Kevin Buzzard arguing for the value of formalization) would be hugely magnified. But what about the attribution of results to specific authors, who take both credit and responsibility for them? Suppose that at some point in the future AI becomes more autonomous, reading the literature and solving many problems that it finds. Suppose also that its solutions are autoformalized, so there is no serious doubt about their correctness. In such a situation, there would be nothing for a human to take credit for or responsibility for. Does that mean that we should declare such results undesirable and threatening to mathematical values?

Of course, something could well be missing in such a situation: perhaps the proofs would be badly written and hard to follow, which would mean that they lacked something we all very much value. So let me extend the thought experiment slightly. What if by that stage one could take one of these outputs and ask an LLM to explain the ideas, and what if LLMs did a very good job at that? That is not particularly hypothetical, since they are often pretty good at this job already, but I am imagining a world in which they are much better than they are now, as they will presumably become.

So now we would have a world in which a lot of problems had been solved, we were sure that the solutions were correct, and we had an LLM ready to explain those solutions in as much or as little detail as we wanted. Is that a future we should resist, and if so, why?

One obvious reason is that it would take a huge part of the fun out of the subject. It is extremely satisfying to struggle with a mathematical problem for months or even years and eventually solve it. But I worry about that argument, because it seems to be saying that we should resist doing mathematics the easy way because a tiny fraction of the world’s population gets huge pleasure from taking orders of magnitude longer to do it. That is not to say that I wouldn’t be sad that a way of life that has sustained me for the last forty years was not available any more — of course I would. I just find it hard to use it as a reason to argue that we should try to preserve the “ownership structure” of mathematical results. If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all. I’m not necessarily in a hurry for that world to exist, but maybe once the transition had happened, people would be OK with it.

The third value I share in an uncomplicated way, and I have already discussed the fourth. The fifth value is one that I hold very strongly, though I’m not so keen on the idea of experts consciously “shaping the direction of research”, something that I see as happening more organically. Obviously there are some notable examples of mathematicians who have created wonderful programmes of research, but even there I would like to credit other mathematicians with understanding what is wonderful about those programmes and contributing to them enthusiastically as a result, rather than being told what direction to pursue and meekly doing so (which is probably not what the declaration is actually trying to suggest, but it has a slight flavour of that for me).

But that’s a minor quibble when set against my main worry about the effect of AI on mathematics, which is the possible destruction of mathematical culture. There is at the moment an extraordinary body of knowledge and expertise that exists not just in the mathematical literature but in the heads of mathematicians all round the world. Imagine if AI didn’t exist and a pandemic broke out that for some reason wiped out all mathematicians and nobody else. All the literature would still be there, but nobody would have the faintest idea what to do with it. To revive a mathematical tradition under those circumstances would be extremely difficult and take decades. Now imagine a slight variant of that, where AI does exist and because of it people are no longer motivated to put in the years of effort it takes to reach the level of expertise that a typical research mathematician has now. After a decade or two, we might arrive at a situation where the mathematical literature has, in some form, been vastly expanded, but there is no corresponding community of human experts who have a shared understanding of parts of it. Almost all of mathematics would be like the areas that we have more or less forgotten about today, areas that exist in papers written many decades ago that nobody reads any more. (I won’t name any such area because I don’t want accidentally to suggest an area that many people still love and work on.)

This, it seems to me, is a possibility that we should try very hard to resist, but I agree with many other commentators who say that in order to resist it, we will need to give less priority to some of our current values — and I would include ownership of mathematical results in that list — and more to others. For example, if Person A gets an LLM to one-shot a solution of an important open problem (which is formalized, possibly automatically, so there is no doubt about its correctness) but Person B makes the effort to digest the solution and explain it in a way that other mathematicians can understand and learn from, then I think we will want Person B to get the lion’s share of the credit. The credit would be of a slightly different from what it is now, which could be described as admiration for somebody’s talent, insight, speed (I mean here the purely factual statement that speed is often admired — I would prefer that to be less the case) and hard work. It would be more like the gratitude that one feels already for somebody who writes a beautiful textbook that makes a whole area of mathematics coherent and accessible.

Maybe that is what the “research mathematicians” of the future should do: make a selection from a vast sea of AI-generated mathematics and write a book about it in such a way that other mathematicians can read the book and feel the kind of enrichment that we feel when we get to grips with an area of mathematics.

At this point I have to admit that there’s a pessimistic side of me that asks the following general question whenever anyone says anything about what the role for humans might be in the future: why do you think that AI wouldn’t be able to do it? For example, with the suggestion I’ve just made, what reason is there to suppose that ChatGPT 8.2 wouldn’t be able to have a short interaction with you about your mathematical tastes and background and then write the ideal textbook just for you? Humans are likely to be better at this kind of curating for a little while yet, but is it a fundamentally human ability that AI could never hope to emulate?

In a world where AI wrote bespoke textbooks (or more likely, just taught people in some more direct way), something would be lost that feels important: mathematics as a collective endeavour. If we all just learnt cool bits of maths for our own private satisfaction, we would miss the considerable pleasure that comes from discussing mathematics with others, though even that could in principle be restored by a benign LLM that deliberately taught many people the same cool bits of the subject, though an LLM that could do that sort of social engineering would raise all sorts of safety issues.

Let me now turn to the section of the declaration about potential threats. I’ll put my comments on each one in square brackets.

  1. Current automated techniques can produce plausible but unreliable (or even incorrect) arguments which are difficult to distinguish from correct mathematical proofs. This applies not only to informal arguments, but also to formalizations, where the difficulty lies in the translation between computer-encoded and human presentations of concepts. These fast-moving developments put our present system of review under increasing pressure, jeopardizing our ability to implement traditional standards for the correctness, transparency, and independent verifiability of proof. [This feels like less of a problem now than it did last September, partly because the best LLMs hallucinate a lot less than before, and partly because autoformalization is improving all the time — I have just used harmonic.fun’s Aristotle system to formalize a complicated paper in Lean and I didn’t need to know any Lean to do it.]
  2. Technologies that draw extensively on the published mathematical commons undermine the traditional system of attribution. Models trained on published works frequently return outputs that do not properly cite the human works they synthesize. Many current models are also built on data obtained by systematically exploiting licenses and access arrangements that were not made with artificial intelligence in mind, or indeed by simply violating copyright protections. [This is a problem at the moment, when ownership of results is important, and I am very much in favour of people making an effort to give appropriate credit for mathematical ideas that AI may have used. However, in the longer term, as I have already discussed, I think this ownership structure will break down and the issue will become less important. It also seems possible that LLMs will become better at revealing their sources.]
  3. Technologies which affect the way in which mathematics is practiced may disturb the current system of incentives. The use of artificial intelligence — and thus also the sort of problems which it can address — may become incentivized for its own sake, disrupting our mechanisms for hiring, funding, and recognition. This disadvantages researchers who do not have access to the technologies or decision-making related to them, or who are unwilling to use technologies controlled by organizations whose values they do not share. [These seem to me to be genuine problems. I think there is simply no point in hoping that our current system of incentives will not be disturbed — it obviously will. I am not necessarily too worried if our mechanisms for hiring, funding and recognition are disrupted, as I don’t find those mechanisms unproblematic as they are, but disadvantaging researchers who do not have access to good LLMs is something I certainly think we should worry about.]
  4. Proper evaluation is endangered if results are communicated through informal channels such as press releases or blog posts, often without any research paper or other disclosure of information necessary for scientific evaluation. This practice seeks publicity for new results on market timelines before the accepted processes of community evaluation in mathematics can take place. In many cases this leads to simplifications in reporting, such as overemphasizing the significance of automated tools and undervaluing the prior human contributions which have made those tools possible. Such oversimplification risks influencing public opinion in a way that not only damages perceptions of mathematics, but also misleadingly uses specific mathematical tasks as metrics for the general reasoning capacities of commercial products. [I think this can be a problem, but I think it is not as serious a problem as some of the others, since when results get overhyped, there seems to be no shortage of people publicly (and rightly) pointing that out.]
  5. These developments put the autonomy of mathematics under threat. The increasing involvement of technology companies in mathematical research raises the risk that research questions may come to be prioritized because of their amenability to automated mathematics, rather than expert judgment of their deeper significance. Indeed, broader understanding of the field may be permanently lost in the process of automation. With university budgets under pressure, this reshaping also changes professional incentives in a manner which encourages the collaboration of researchers with technology companies on asymmetric terms. If left unchecked, these trends go beyond threatening researchers’ autonomy, affecting the scope and depth of mathematical research itself. [I think this could be a problem, but it also seems to me that mathematicians have a lot of power here. For instance, if a technology company were to produce a lot of research that mathematicians did not find all that interesting or important, I don’t think they would be able to use their financial and other resources to persuade us to change our minds. Rather, what seems to happen is that mathematicians say, “Yes that does X but it doesn’t do Y,” and the tech companies then feel challenged to do Y.]

There follow eleven recommendations for individual mathematicians. I agree with almost all of them. The one that I’m not so sure about, for reasons I’ve basically already gone into, is this.

Affirm the humanity of authorship. Credit and responsibility continue to belong to humans within the mathematical community and should not be given to automated systems. Artificial intelligence may obscure, but does not replace, the collective human labor behind a result.

I’m not sure what that really means. For example, should we affirm the humanity of authorship in the case of the solution to the unit-distance problem? Some humans did a wonderful job of explaining the proof that OpenAI’s model came up with, and the model made use of some highly non-trivial mathematics produced by humans, but the solution itself has not been credited to any human, and nor should it be in my view.

Under recommendations for mathematical organizations and not-for-profit research funders I again agree with several of them but have my doubts about some. An interesting case is the following.

Protect the rights of authors. Automated mathematics presents new challenges to the rights of authors, and societies should be proactive in the development of sample licensing agreements to protect these rights. In particular, material should not be used as training data without consent, and publishing agreements should allow authors to opt-out [sic] of the use of their work in this way.

This recommendation seems to belong to a world in which journal articles are the main means of dissemination of mathematics. But that has long since ceased to be the case: almost all dissemination now takes place via arXiv preprints, with journals limited to providing a little extra mark of prestige. Once an article is on arXiv, it is on the internet and one can hardly ask for it not to be used as training data. So this recommendation, if it applies at all, will apply to a tiny fraction of articles that are published without first appearing on arXiv. More generally, what right of an author is being compromised when an article is used as training data? We don’t object if human mathematicians use our articles to help train themselves to become better mathematicians — indeed, we will typically be delighted that somebody else thought our articles worthy of their attention. So the objection to a machine doing the same would have to be that for some reason one did not want machines to get better at mathematics in a similar way. I can imagine grounds for such a wish: perhaps somebody is worried about the threat that LLMs pose to traditional mathematical practice, or perhaps they worry that mathematical ability of LLMs will transfer to much more dangerous reasoning ability. But there’s a more complicated discussion to be had here than one might think from reading the recommendation.

The next recommendation is this.

Insist on appropriate publication outlets. Demand that mathematical results continue to be published in peer-reviewed venues such as journals, proceedings, and books. Informal mechanisms such as press releases or blog posts can provide a valuable supporting role, but they cannot replace peer-review or community scrutiny.

For reasons that I’ve gone into many times, I am not too fond of the current publication system, so I can’t get behind this recommendation. Indeed, if the current system becomes unsustainable because of a flood of AI-generated and AI-aided content, I would regard that as a beneficial consequence of AI. However, that doesn’t mean that I would advocate a total free-for-all. I’ve already said that one of my worries is that if mathematical content is not sufficiently organized, then the traditions that we all value could die. I just think that what we will want to do to preserve those traditions is likely to be a lot more innovative than clinging on to the peer-reviewed journal system.

I have highlighted in this post the parts of the declaration that I have doubts about, either because I disagree with them or, more typically, because I sort of half agree with them but want to add many qualifications. That may make the post come across as rather negative, but that is not my intention. The parts I disagree with are in the minority, and I think it is important that a declaration such as this should be made. I should also make clear that my views are evolving all the time, largely because the speed of progress of LLMs has taken me by surprise, but also as a result of conversations I have had or opinions that other mathematicians have expressed online.

I’ll end with two further clarifications. The first is that it may seem as though I am taking it for granted that LLMs will soon be better than humans at all aspects of mathematical problem solving, and maybe also problem posing, theory building, formulation of definitions, etc. I do think all that will happen at some point, but whereas some people say that it will obviously happen within the next two to three years, I would say that it might happen as soon as that, but I don’t rule out that we’ll get lucky and find that we can do interesting AI-assisted maths for quite a bit longer than that before AI doesn’t need us any more.

The second is that I think I have acquired a reputation as somebody who celebrates what is going on. But if, for example, I post on Twitter saying that such-and-such an AI solution is a remarkable development, the word “remarkable” is meant to indicate no more nor less than that I found it very surprising. My feelings about the possibility of AI solving all sorts of problems that interest me are much more mixed. I’ve had the experience twice now of seeing GPT 5.6 Pro one-shot a solution to a problem that I very much liked and had thought about hard (in both cases with much younger collaborators, who, with my approval, were the ones who prompted the LLM). It felt very strange and not particularly pleasant to have the rug pulled out from under my feet like that. On the other hand, I was quite pleased to see the problems solved. It’s actually a similar feeling to the one I have had many times when a problem I am fond of and have thought about gets solved by another human mathematician.

Another factor for me is that I have invested a lot of thought into automatic theorem proving of a more traditional kind. One of my main motivations for that was the hope that the work I put into it would extend the state of the art, measured by which problems a computer can solve. That ship has sailed now, and that saddens me. I still think that there is value in the work that I and my group are doing, but it has become a tougher sell.

So I personally have already found AI quite disruptive, and this is just the beginning. I would have preferred the developments to happen at a slower pace. But I don’t see any practical way to slow them down, so the best we can do is probably to face up to the changes that are being thrust upon us and do what we can to maximize the benefits and minimize the damage. The Leiden Declaration may not be perfect, but it makes an important and positive contribution to that effort.

July 25, 2026

Clifford JohnsonOn top of the Mountain again

Just in case you’re up for a short talk at the top of Mount Wilson followed by an evening of observing through the historic telescopes on Saturday 25th July… this might be for you! Go to Mount Wilson Observatory’s website for more. –cvj

The post On top of the Mountain again appeared first on Asymptotia.

July 24, 2026

Matt von HippelQuantum Apologetics and Quantum Theology

As an atheist, I started out frustrated by how little interest religious people had in debating their beliefs. Much of that was probably to do with how obnoxious it was to be “debated” by a socially awkward ten-year-old. But as I appreciate now, defending religious beliefs is just not a core activity for most religious people, even the experts. While there are many theologians who study the doctrines of their various religions, only a few engage in apologetics: arguments designed to convince people on the outside. The rest work within a particular religion, working out its implications.

There’s a similar, less-often-noticed behavior when it comes to interpretations of quantum mechanics.

Much like religions, there are many different interpretations of quantum mechanics, from people who envision a fundamentally undetermined world to those who picture a vast multiverse of all possibilities, to people who think quantum mechanics needs to be supplemented with faster-than-light signals, deterministic rules, or even consciousness. And as there isn’t yet any broad consensus for any of these options, you’d be forgiven for assuming that these people are trying their hardest to convince others that their interpretation is right.

But most of the work these people do is “quantum theology”, not “quantum apologetics”. The average paper connected to a quantum interpretation isn’t designed to convince people with different interpretations. It starts out with an interpretation in place, and tackles a more detailed question of what the interpretation should actually mean in practice. This work can be quite valuable and impressive…provided you’re already convinced. But if you’re not, and you see someone glowingly praise a paper on say the many-worlds interpretation, you might mistakenly think they’re on the cusp of closing the question of which interpretation is right, and not merely solving a technical issue within many-worlds.

Are there other areas of science with this pattern?

Let’s talk about string theory.

Physicists really started getting excited about string theory around forty years ago. Some people hear that number and wonder what happened. Have string theorists been looking for evidence for forty years and not found any? Isn’t that a huge waste of time?

That would be “string apologetics”. And these days, string apologetics is not actually that common. It’s a priority for some, to be sure, but most string theorists aren’t working on proving string theory. It’s clearly not an easy thing to do, and as a result, most people don’t spend their time on it.

Realizing that, some people assume that string theorists are actually practicing “string theology”, and get mad all over again. If string theorists are just assuming string theory and spending their time figuring out the “string versions” of known facts, then many would also deride their work as pointless.

But actually, string theology is also not very common. There are certainly some people who work to figure out the “string version” of this or that, or do research that only makes sense assuming string theory. But most of the string theory community doesn’t do that either.

Instead, they do something that, to continue the analogy, you might call “string pastoral care”.

Most theologians aren’t apologists, but similarly, most priests don’t spend their time doing theology. They use their training to advise their congregations on how to live their lives, solving day-to-day problems with some religious inspiration.

Similarly, most people these days who call themselves string theorists are working on more general questions about the types of theories used in particle physics. They’re doing this making use of their background in string theory, as an inspiration for solutions, a source of mathematical tools to solve problems, and a motivation for which questions are the most interesting. But if string theory turns out to be false, most of these peoples’ work will still be useful. It’s “string pastoral care”, used to solve problems for the people around them, not “string theology”.

Do you know any other fields that this applies to? Let me know in the comments!

Peter Rohde Introducing Sigfried’s Blog

My new secondary blog featuring conversations with AI, inventing new things, exploring hypotheticals, letting creativity flow freely.

Some highlights:

  • Satellite constellations with topologically distributed apertures.
  • A clockless architecture for classical topological computing.
  • Post-quantum cryptography using the \mathbb{Z}_2^n \rtimes S_n algebra.
  • Efficient homomorphic computing using reversible classical circuits.
  • A silent speech interface using microwave Doppler imaging.
  • Cognitive search acceleration.
  • Consensual thought guidance.
  • Subliminal audio modulation & human guidance systems.
  • Microwave imaging using WiFi and 5G for medical applications.
  • Thought tomography.
  • The quantum bluff hypothesis.

https://sigfriedschattenjaeger.wordpress.com

Doug NatelsonA brief serious note about mental health and well-being

I've been blogging for 21 years now (!!), and over that time I've had a wide variety of comments on here, but there have been a couple of anonymous ones in the last few weeks that really worried me.  It's the internet, so you can't readily tell when someone is trolling, but these made me concerned for the safety and emotional state of the commenter.   Just remember, there is always someone to talk to who is ready to listen, and we're all better with you than without you.  The 988 hotline (https://988lifeline.org/) is available 24/7/365, free and confidential.  Please take care.

July 20, 2026

Secret Blogging SeminarThe new counterexample to the Jacobian conjecture

As many of you have probably heard already, yesterday morning, Levent Alpöge tweeted that Fable had found a counterexample to the Jacobian Conjecture. Specifically, let

a=(1+xy)3z+y2(1+xy)(4+3xy),b=y+3x(1+xy)2z+3xy2(4+3xy),c=2x3x2yx3z,\begin{align*} a&=&(1+xy)^3z+y^2(1+xy)(4+3xy),\\ b&=&y+3x(1+xy)^2z+3xy^2(4+3xy),\\ c&=&2x-3x^2y-x^3z, \end{align*}

Then the Jacobian of (a,b,c) is easily checked to be -2. However, the map (a,b,c) is generically three to one, not bijective.

I’m sure many of you are playing with these polynomials to see what you can figure out about them. This is a place for us to share our observations. I’ll post a few minor observations of my own soon.

First, a basic but intriguing observation from Mathoverflow user “dorky”: The polynomials a, b and c are homogeneous with respect to the grading where \deg(x) = -1, \deg(y) = 1 and \deg(z)=2; their degrees are \deg(a) = 2, \deg(b) = 1 and \deg(c) = -1. I’m not sure what to make of this, but it surely matters.


Some computations by me: If you eliminate any two of the variables (x,y,z), you get a cubic relation in the remaining variable. Here they are

2c+(43bc)x+(16ab218abc+b3c+27a2c2)x3(18ab+b3+27a2c)+18ay3by2+2y3(really long)+8z3\begin{matrix} -2 c+(4 – 3 b c) x + (16 a – b^2 – 18 a b c + b^3 c + 27 a^2 c^2) x^3 \\ (-18 a b + b^3 + 27 a^2 c)+18 ay-3 b y^2+ 2y^3 \\ (\text{really long}) + 8 z^3 \\ \end{matrix}

I’m leaving out the “really long”, because it is really long and I suspect we don’t care about the details. Put

Δ=16ab218abc+b3c+27a2c2\Delta= 16 a – b^2 – 18 a b c + b^3 c + 27 a^2 c^2 ,

the leading coefficient of the x cubic. Then the discriminants of the three cubics are \Delta p^2, \Delta q^2, \Delta r^2 where

pamp;=amp;89bc+27ac2qamp;=amp;bramp;=amp;(really long)\begin{align*} p &amp;=&amp; 8 – 9 b c + 27 a c^2 \\ q &amp;=&amp; b \\ r &amp;=&amp; (\text{really long}) \\ \end{align*}

The polynomials (p,q,r) have no common zeroes. Roughly speaking, our map should have special behavior over the loci \Delta=0, p=0, q=0 and r=0. The fact that $p$, $q$ and $r$ each appear cubed means that the variables x, y and z should have three fold branching over the loci p=0, q=0 and r=0 (respectively).

I’m having trouble visualizing what happens over \Delta=0 — since the leading coefficient of the x cubic drops out, the map is 2 to 1 rather than 3 to 1 over this point. But, at the same time, the y and z cubics have a multiple root at the points of \Delta=0. Does anyone see how to visualize this?

Any other insights?

Tommaso DorigoToward Mode Collapse of Natural Language

Toward Mode Collapse of Natural Language

Regression toward the mean is a simple phenomenon commonly described in Statistics 101 courses. If you measure a parameter describing some phenomenon, you will find that extreme measured values tend to be followed by less extreme ones.

Tommaso Dorigo
Categories

July 19, 2026

John BaezGalilean Limits of Electromagnetism

Maxwell’s equations are invariant under Lorentz transformations. The usual equations of fluid flow are not! Like the rest of Newtonian mechanics, they’re invariant under Galilean transformations like

t' = t,  \quad  x' = x - vt

So, if we simply slap these two theories together, we get a mess! How can we study electrically conductive fluids—like plasma—without bringing special relativity into the game?

We can use a limiting case of Maxwell’s equations where we ignore terms that become tiny when all the particles are moving much slower than light.

There seem to be at least two ways to do this: there’s an ‘electric limit’ of Maxwell’s equations and a ‘magnetic limit’. Both are invariant under Galilean transformations. The original derivation of these limits by Le Bellac and Lévy-Leblond in 1973 used the version of Maxwell’s equations including the electric permittivity \varepsilon_0 and magnetic permeability \mu_0 of the vacuum, whose product is 1/c^2. This is convenient but not necessary, as explained here:

• Jose A. Heras, The Galilean limits of Maxwell’s equations.

In the magnetic limit of Maxwell’s equations, we throw out effects due to time-varying electric fields:



People often use the magnetic limit when studying nonrelativistic electrically conductive fluids. In this situation they often consider a version of the magnetic limit where the charge density \rho is zero, since this is typically close to true in a plasma. However Heras does not do this, nor does the original paper:

• Le Bellac and Levy-Leblond, Galilean electromagnetism.

In the electric limit of Maxwell’s equations, we throw out effects due to time-varying magnetic fields:



It’s fun to compare the magnetic and electric limits.

The magnetic limit has been called ‘pre-Maxwellian’, because it’s like electromagnetism before Maxwell added the extra term that makes a changing electric field create a curl in the magnetic field. Without this term there is no light!

In the electric limit you also can’t have light, because it’s missing the term that makes a changing magnetic field create a curl in the electric field.

In the magnetic limit you can’t have capacitors, because those store energy in the electric field, and in the magnetic limit the energy density is just \mathbf{B} \cdot \mathbf{B}/2.

Similarly, in the electric limit you can’t have inductors, because inductors store energy in the magnetic field, and in this limit the energy density is just \mathbf{E} \cdot \mathbf{E}/2.

It’s all nicely symmetrical! But still somewhat mysterious to me. All the derivations of these limits that I’ve seen involve too many parameters for my taste, and too much talk. But that’s how I often feel when I’m just starting to study a piece of physics.

Besides the two papers mentioned in my last post, I’ve been looking at this:

• Giovanni Manfredi, Non-relativistic limits of Maxwell’s equations.

There’s a lot I haven’t explained here. I haven’t even said how the electric or magnetic fields transform under Galilean boosts in these limiting theories! I find this subject fairly confusing, and I’d probably have to redo all the calculations to really understand them. As Feynman said, “what I cannot create I do not understand”.

Someday I should dig deeper into this subject and explain how the two limits work in a way I find satisfying. I should also draw the connections to this earlier article of mine:

Magnetohydrodynamics.

n-Category Café Octonions and the Standard Model (Part 15)

Last time I described a way to get the Standard Model gauge group from the exceptional Jordan algebra. But that approach gave no obvious nice way to put quarks and leptons into the picture. This new paper tackles that problem:

Jordan pairs and Jordan triples are two closely linked formalisms that generalize Jordan algebras. Our paper explains them in detail — and how they’re connected to geometry and quantum mechanics. Here I will mostly skip that wonderful story, so I can quickly explain the connection to the Standard Model.

Here’s how the Standard Model gauge group, together with its representation on one generation of fermions, drops out of a Jordan triple.

The bi-Cayley triple

Let

𝕆 = 𝕆\mathbb{O}_\mathbb{C} = \mathbb{C} \textstyle{\otimes}_\mathbb{R} \mathbb{O}

be the bioctonions: octonions with complex coefficients. Write 𝕆 2\mathbb{O}_\mathbb{C}^2 for the space of column vectors with two bioctonion entries.

𝕆 2\mathbb{O}_\mathbb{C}^2 has a certain triple product

[x,y,z]=12(x(y z)+z(y x)) [x,y,z]=\frac{1}{2}(x(y^{\dagger}z)+z(y^{\dagger}x))

which obey the axioms of a gadget called a ‘positive hermitian Jordan triple’. It’s called the bi-Cayley triple.

Now, every positive hermitian Jordan triple gives rise to a 2\mathbb{Z}_2-graded real Lie algebra

k=k 0k 1 \mathbf{k} = \mathbf{k}_0 \textstyle{\oplus} \mathbf{k}_1

Not a Lie superalgebra: a plain old-fashioned Lie algebra with a 2\mathbb{Z}_2-grading!

How does this work? We take the hermitian Jordan triple itself to be k 1\mathbf{k}_1. The Lie algebra k 0\mathbf{k}_0 consists of all linear maps from k 1\mathbf{k}_1 to itself that are of this form:

x[a,b,x][b,a,x] x \mapsto [a,b,x] - [b,a,x]

for some a,bk 1a,b \in \mathbf{k}_1. These maps are called real inner derivations. They form a Lie algebra since the commutator of two such maps is another such map. With a bit more work we can define other operations making all of k\mathbf{k} into a 2\mathbb{Z}_2-graded Lie algebra.

So, we get a big Lie algebra k\mathbf{k}, and a Lie subalgebra k 0\mathbf{k}_0 sitting inside it. From this we get two Lie groups: a big one KK whose Lie algebra is k\mathbf{k}, and a subgroup K 0K_0, whose Lie algebra is k 0\mathbf{k}_0.

The quotient is K/K 0K/K_0 is a nice kind of manifold called a hermitian symmetric space. Conversely, any compact hermitian symmetric space give rise to a positive hermitian Jordan triple!

This geometric picture is revealing. The group KK acts transitively as symmetries of our hermitian symmetric space, while the stabilizer of any point is isomorphic to K 0K_0. Our original Jordan triple, k 1\mathbf{k}_1, is then the tangent space of that point. So, K 0K_0 acts on the Jordan triple. This action preserves the triple product, and we call K 0K_0 the real inner automorphism group of our Jordan triple.

Here’s another great thing about the geometric picture: hermitian symmetric spaces were classified by Eli Cartan (who seems to have spent his life classifying things). As a result we also know the classification of positive hermitian Jordan triples. They come in four infinite series together with two exceptions. One is the bi-Cayley triple, and other is the Albert triple, which is the complexification of the exceptional Jordan algebra. The bi-Cayley triple is a subtriple of the Albert triple. It’s these two exceptions that are connected to the Standard Model. But we’ll start with the bi-Cayley triple.

The 3-graded Lie algebra coming from the bi-Cayley triple is the compact real form of 𝔢 6\mathfrak{e}_6:

𝔢 6=[𝔰𝔬(10)𝔲(1)]𝕆 2.\mathfrak{e}_6 = \big[\mathfrak{so}(10) \textstyle{\oplus} \mathfrak{u}(1)\big] \textstyle{\oplus} \mathbb{O}_\mathbb{C}^2.

The even part of this Lie algebra is in brackets. The corresponding hermitian symmetric space is called the bioctonionic plane (𝕆)P 2(\mathbb{C}\otimes\mathbb{O})P^2. I explained it in Part 12. The even part of our 3-graded Lie algebra, 𝔰𝔬(10)𝔲(1)\mathfrak{so}(10)\oplus \mathfrak{u}(1), generates the stabilizer of a point in the bioctonionic plane. The odd part, our friend 𝕆 2\mathbb{O}_\mathbb{C}^2, is the tangent space of that point.

Here’s the first big surprise. The even part transforms as the adjoint representation of Spin(10)\mathrm{Spin}(10), while the odd part itself transforms as the 16-dimensional complex spinor representation of Spin(10)\mathrm{Spin}(10). Ignoring the extra U(1)\mathrm{U}(1) for a moment, this is exactly what we see in a SO(10)\mathrm{SO}(10) grand unified theory: gauge bosons in the adjoint representation, and one generation of fermions in the 16-dimensional spinor representation.

So before we do anything, the bi-Cayley triple already smells like it contains the ingredients of an SO(10)\mathrm{SO}(10) grand unified theory.

Tripotents

In a Jordan algebra the important elements are the idempotents, e 2=ee^2 = e. In a Jordan triple WW their role is played by tripotents: elements ee with

[e,e,e]=e.[e,e,e] = e.

A tripotent always lets us split WW into three parts via something called its Peirce decomposition. The operator w[e,e,w]w \mapsto [e,e,w] has eigenvalues 0,12,10, \tfrac{1}{2}, 1, and WW splits into the corresponding eigenspaces

W=W 0(e)W 1/2(e)W 1(e),W = W_0(e) \textstyle{\oplus} W_{1/2}(e) \textstyle{\oplus} W_1(e),

which are called the Peirce 0-space, Peirce 12\tfrac{1}{2}-space and Peirce 1-space of ee. A tripotent is called minimal when its Peirce 11-space is one-dimensional: minimal tripotents are the analogues of unit vectors in ordinary quantum theory. Two tripotents e 1,e 2e_1, e_2 are called colinear when each lies in the other’s Peirce 12\tfrac{1}{2}-space.

I can’t resist explaining some of the quantum physics here. I said I wouldn’t, but I can’t help it. In a hermitian Jordan triple, the triple product [,,][-,-,-] is linear in the first and last slot, but conjugate-linear in the middle slot. So, if you multiply a tripotent by a phase α\alpha, you get a new tripotent:

[αe,αe,αe]=αα¯αe=αe [\alpha e, \alpha e, \alpha e] = \alpha \overline{\alpha} \alpha e = \alpha e

This should remind you of how when you multiply a unit vector in a Hilbert space by a phase, you get a new unit vector. In Jordan triple quantum mechanics, minimal tripotents take the place of these unit vectors. And guess what: the hermitian symmetric space K/K 0K/K_0 that I was talking about earlier is also the space of minimal tripotents mod phase! So, it generalizes the familiar space of ‘pure states’ in quantum mechanics, which are unit vectors mod phase.

But let’s get back to the Standard Model.

A chain of Jordan triples

From here on, the single fact driving everything is this: in any hermitian Jordan triple, any minimal tripotent’s Peirce 12\tfrac{1}{2}-space is itself a hermitian Jordan triple!

If we run this starting from the bi-Cayley triple, we get this chain of hermitian Jordan triples, where each row’s 12\tfrac{1}{2}-space is the next row’s triple:

Jordan triple Lie algebra k 0k 1\mathbf{k}_0 \oplus \mathbf{k}_1 (even part in brackets) real inner automorphism group
W=𝕆 2W = \mathbb{O}_\mathbb{C}^2 𝔢 6=[𝔰𝔬(10)𝔲(1)]𝕆 2\mathfrak{e}_6 = [\mathfrak{so}(10) \oplus \mathfrak{u}(1)] \oplus \mathbb{O}_\mathbb{C}^2 (Spin(10)×U(1))/ 4(\mathrm{Spin}(10) \times \mathrm{U}(1)) / \mathbb{Z}_4
W=𝔞 5()W' = \mathfrak{a}_5(\mathbb{C}) 𝔰𝔬(10)=[𝔰𝔲(5)𝔲(1)]𝔞 5()\mathfrak{so}(10) = [\mathfrak{su}(5) \oplus \mathfrak{u}(1)] \oplus \mathfrak{a}_5(\mathbb{C}) SU(5)×U(1)\mathrm{SU}(5) \times \mathrm{U}(1)
W=M 3,2()W'' = \mathrm{M}_{3,2}(\mathbb{C}) 𝔰𝔲(5)=[𝔤 SM]M 3,2()\mathfrak{su}(5) = [\mathfrak{g}_{\mathrm{SM}}] \oplus \mathrm{M}_{3,2}(\mathbb{C}) G SMG_{\mathrm{SM}}

Here 𝔞 5()\mathfrak{a}_5(\mathbb{C}) is the Jordan triple of antisymmetric 5×55\times 5 complex matrices, M 3,2()\mathrm{M}_{3,2}(\mathbb{C}) is the Jordan triple of 3×23\times 2 complex matrices, 𝔤 SM=𝔰𝔲(3)𝔰𝔲(2)𝔲(1)\mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3)\oplus\mathfrak{su}(2)\oplus \mathfrak{u}(1), and

G SM=S(U(2)×U(3))(SU(3)×SU(2)×U(1))/ 6G_{\mathrm{SM}} = \mathrm{S}(\mathrm{U}(2) \times \mathrm{U}(3)) \cong (\mathrm{SU}(3)\times\mathrm{SU}(2)\times\mathrm{U}(1))/\mathbb{Z}_6

is the true Standard Model gauge group.

The gauge group from two tripotents

Now pick two colinear minimal tripotents e 1,e 2We_1, e_2 \in W. Descend the table twice:

  • Start with W=𝕆 2W = \mathbb{O}_\mathbb{C}^2, which has real inner automorphism group (Spin(10)×U(1))/ 4(\mathrm{Spin}(10)\times\mathrm{U}(1))/\mathbb{Z}_4.
  • Fix e 1e_1. Its Peirce 12\tfrac{1}{2}-space is W=𝔞 5()W' = \mathfrak{a}_5(\mathbb{C}), with real inner automorphism group SU(5)×U(1)\mathrm{SU}(5)\times\mathrm{U}(1).
  • Fix e 2e_2 (colinear with e 1e_1, so living in WW'). Its Peirce 12\tfrac{1}{2}-space in WW' is W=M 3,2()W'' = \mathrm{M}_{3,2}(\mathbb{C}), with real inner automorphism group exactly G SMG_{\mathrm{SM}}.

In other words, the subspace of the bi-Cayley triple colinear with both e 1e_1 and e 2e_2 is a Jordan triple whose real inner automorphism group is the Standard Model gauge group.

The choice of e 1e_1 and e 2e_2 also pins down how G SMG_{\mathrm{SM}} sits inside the original group E 6\mathrm{E}_6. At each we step take the subgroup that acts with determinant 11 and preserves the chosen tripotent up to a phase; this gives a chain of subgroups whose members are Spin(10)\mathrm{Spin}(10), U(5)\mathrm{U}(5), and G SMG_{\mathrm{SM}}, so we get the embedding

G SMSU(5)Spin(10). G_{\mathrm{SM}} \subset \mathrm{SU}(5) \subset \mathrm{Spin}(10).

In particle physics, this is the classic chain taking us from the so-called SO(10)\mathrm{SO}(10) grand unified theory down to the SU(5)\mathrm{SU}(5) grand unified theory down to the Standard Model. And it’s well known that restricting the 16-dimensional complex spinor representation of Spin(10)\mathrm{Spin}(10) along this chain gives precisely the Standard Model representation ρ SM\rho_{\mathrm{SM}} on one generation of fermions! So we get one generation of Standard Model fermions this way.

The six particles types as Peirce spaces

We have gotten the representation of the Standard Model gauge group on one generation of fermions without any fuss. But it’s also fun to peer into the details, and see how the different kinds of fermions emerge. We can get them using the fact that for any tripotent ee, we have projections P 0(e),P 1/2(e)P_0(e), P_{1/2}(e) and P 1(e)P_1(e) onto its three eigenspaces: its so-called Peirce projectors.

Since we get the Standard Model gauge group and its representation on fermions from two minimal tripotents e 1e_1 and e 2e_2, we have nine Peirce projectors we can apply to our Jordan triple 𝕆 2\mathbb{O}_{\mathbb{C}}^2. Let’s use these to pick out various kinds of particles!

As a representation of the Standard Model Lie algebra

𝔤 SM=𝔰𝔲(3)𝔰𝔲(2)𝔲(1),\mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3) \textstyle{\oplus} \mathfrak{su}(2) \textstyle{\oplus} \mathfrak{u}(1) ,

any generation of Standard Model fermions transforms as the direct sum of six irreducible representations:

ρ SM=(3,2,16)(3¯,1,13)(3¯,1,23)(1,2,12)(1,1,1)(1,1,0),\rho_{\mathrm{SM}} = (3,2,\tfrac{1}{6}) \textstyle{\oplus} (\bar 3,1,\tfrac{1}{3}) \textstyle{\oplus} (\bar 3,1,-\tfrac{2}{3}) \textstyle{\oplus} (1,2,-\tfrac{1}{2}) \textstyle{\oplus} (1,1,1) \textstyle{\oplus} (1,1,0),

These correspond to the six types of left-handed fermion: q L,d R¯,u R¯, L,e R¯,ν R¯q_L, \overline{d_R}, \overline{u_R}, \ell_L, \overline{e_R}, \overline{\nu_R}. Six irreducible pieces, six particle types.

It turns out these are exactly the six nonzero components of the Peirce decomposition of 𝕆 2\mathbb{O}_\mathbb{C}^2 with respect to both e 1e_1 and e 2e_2. Those six match up one-to-one with the particle types:

Peirce projector representation of G SMG_{\text{SM}} particle type
P 1/2(e 2)P 1/2(e 1)P_{1/2}(e_2) P_{1/2}(e_1) (3, 2, +1/6) q Lq_L
P 1/2(e 2)P 0(e 1)P_{1/2}(e_2) P_0(e_1) (3¯\overline{3}, 1, +1/3) d R¯\overline{d_R}
P 0(e 2)P 1/2(e 1)P_0(e_2) P_{1/2}(e_1) (3¯\overline{3}, 1, −2/3) u R¯\overline{u_R}
P 0(e 2)P 0(e 1)P_0(e_2) P_0(e_1) (1, 2, −1/2) L\ell_L
P 1(e 2)P 1/2(e 1)P_1(e_2) P_{1/2}(e_1) (1, 1, +1) e R¯\overline{e_R}
P 1/2(e 2)P 1(e 1)P_{1/2}(e_2) P_1(e_1) (1, 1, 0) ν R¯\overline{\nu_R}

The remaining three combinations — P 1(e 2)P 1(e 1)P_1(e_2)P_1(e_1), P 1(e 2)P 0(e 1)P_1(e_2)P_0(e_1), and P 0(e 2)P 1(e 1)P_0(e_2)P_1(e_1) — all vanish, which is why we land on six pieces and not nine.

So the whole package — the gauge group G SMG_{\mathrm{SM}}, the embedding G SMSpin(10)G_{\mathrm{SM}} \subset \mathrm{Spin}(10), the representation ρ SM\rho_{\mathrm{SM}}, and even the split of one generation into its six particle multiplets as distinct Peirce components — all comes out of the single object 𝕆 2\mathbb{O}_\mathbb{C}^2 once you choose two colinear minimal tripotents.

And if you prefer to start one level up, with the Albert triple 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) \otimes \mathbb{C}, you get the same result by choosing three mutually colinear tripotents instead of two — but for that, read our paper!

Scott Aaronson NISQ and quantum supremacy did not fail

A week ago, a philosopher named Amit Hagar put out a preprint entitled The NISQ Trap: Eight Years of Demonstrations the Hardware was Built to Lose. Here’s the abstract:

With a single clear exception, every NISQ-era flagship demonstration of ‘quantum advantage’ has, within eighteen months of its announcement, been classically reproduced, shown to rest on classically tractable structure, or closed by a simulability theorem. Six theoretical results from 2024 through April 2026 explain the pattern: the regions of circuit-space NISQ hardware can run with sufficient fidelity coincide with the regions classical algorithms compress efficiently, because the features that admit one (low effective depth, strong algebraic structure, geometric locality) are the features that admit the other. This reading dates the NISQ programme from its 2018 articulation as an interim retreat from the unmet conditions of the 1996 threshold theorems, characterises the eight years that followed as a closed loop in which the demonstrations the hardware could run were drawn from regions classical methods could already reach, and locates the exit from the loop where the threshold theorems originally located it: in fault tolerance. The empirical pattern could in principle break with a demonstration that escapes the current simulability results. After eight years and more than thirty advantage-class announcements, the burden of producing such a demonstration falls to the defenders of NISQ.

You can also read some debates about the paper on SciRate here. I think it’s fair to say that the paper is purely polemical, without new ideas, and Pangram agrees with my suspicion (and that of a SciRate commenter) that significant portions of it are AI-generated.

Nevertheless, the basic thesis—that quantum supremacy in the NISQ (Noisy Intermediate Scale Quantum computing) era has been a failure, or even an example of pathological science—seems surprisingly widely shared, along with the opposite thesis that quantum computing already gives oodles of useful advantages for optimization and finance.

So it seems worth stating for the record that I have an extremely different view. I would say:

  1. Sampling-based quantum supremacy experiments, including those based on Random Circuit Sampling and BosonSampling, passed the point about two years ago where, absent a breakthrough in classical algorithms, they quite clearly are beating what can easily be simulated on any existing classical computer. Hagar seems to claim that these experiments have been killed by the October 2025 paper Classical simulation of noisy random circuits from exponential decay of correlation, but he ignores that the algorithm from that paper still needs time that’s exponential in the circuit depth (see Theorem 2).
  2. Indeed, simulating deep ~100-qubit random circuits, like those that Google and Quantinuum have now demonstrated experimentally, still seems pretty hopeless with any current classical method. This is particularly true for Quantinuum’s experiments, which had high enough gate fidelity to maintain a Linear Cross-Entropy score of order 1 (i.e., they’re no longer all that “noisy”). The central drawback of these experiments is no longer lack of confidence about quantum advantage; rather, it’s just that we only get samples as output, and directly verifying the quality of the samples seems just as intractable for a classical computer as spoofing the samples.
  3. As of this past year, however, we have some strong candidates for verifiable quantum advantage. One is the Google OTOC experiment, as even Hagar himself acknowledges (that’s his “single clear exception”). A second is the simulations of the 2D Fermi-Hubbard model on Quantinuum and Google machines, like this one. The 1D Fermi-Hubbard model can be classically simulated pretty easily (see here for example), but the 2D one still presents challenges, meaning that in some regimes, the best available estimates of certain observables apparently now come from quantum computers. I wish I could write about other examples that will be public shortly.
  4. Yes, the “real” goal remains, as it’s been since the 1990s, to build a scalable fault-tolerant quantum computer—and I’m glad that Hagar (unlike, say, Gil Kalai) never suggests that we’ve learned anything to rule that goal out. In the meantime, an intermediate goal would be to use NISQ devices to do physics and chemistry simulations that are commercially useful, or that help solve important scientific problems. The point of quantum supremacy experiments, you might say, is that by demonstrating the reality of quantum speedup about as clearly as it can be demonstrated with current hardware, they let us cleanly turn our attention to those more ambitious goals.

Anyway, my son and I need to catch a plane to Utah now, for the next iteration of the wonderful Epsilon Camp, where I’ll again be teaching theoretical computer science to 11- and 12-year-olds. But feel free to discuss in the comments! Nothing about world affairs in this thread please, just quantum supremacy.

Update (July 19): Not unrelated to the subject of this post, here’s a podcast I did with Gill Eapen of “Scientific Sense” about the current situation in quantum computing including recent experimental victories.

July 18, 2026

John PreskillMy friend Mark Wise

Mark Wise, the John A. McCone Professor of High Energy Physics at Caltech, passed away on July 10 at age 72. At a recent memorial service, John Preskill made these remarks.

I’m John Preskill, Mark’s colleague on the Caltech physics faculty for more than four decades. Our friendship goes back even farther. My wife Roberta and I met Mark and Jackie not long after they arrived at Harvard in 1980. We’ve been friends since then. We attended the bris for both Barry and Jonathan during those Harvard days. Mark and Jackie have two boys and we have two girls who are a few years younger, who were thrilled to connect with Barry and Jonathan when the families would get together for occasions like Passover or Hanukkah or Thanksgiving. When the kids were little, Mark and I would sometimes muse about the potential for forging even closer family ties if those relationships blossomed.

That didn’t happen. But Mark would preside at each Seder with a light hand, sprinkling the occasion with corny jokes as was his style, and Jackie would be determined to make it to the end of that customized family-friendly Haggadah she had meticulously prepared. The children, meanwhile, would be wondering when they’d be able to continue their game of sock baseball.

Many of you know that Mark was deeply dedicated to his family and friends. I’ll make some brief remarks about three facets of Mark I know especially well: Mark the scientist, Mark the teacher and mentor, and Mark the colleague and friend.

Because of his self-deprecating manner, those of you who are not scientists may not appreciate Mark’s stature as a physicist. He was one of the most influential figures in theoretical particle physics of his generation. It was not obvious things would turn out that way. Growing up in Toronto, Mark was an indifferent student, and his poor grades reflected that lack of interest. As a 9th grader, though, it struck him that he better change his ways and figure out how to make something of his life. He liked sports — the possibility of being a professional athlete was briefly considered, but discarded. Somehow he decided that science would be a better fit. I’m not sure why — he had recently failed math. But he worked hard and had inspiring teachers, so by the time he finished high school Mark was an excellent student, and he sailed into the University of Toronto well prepared to major in physics,

At U of T, Mark came under the influence of a young professor, Nathan Isgur, who would later become his close research collaborator. Under Nathan’s guidance, Mark sought admission to the PhD programs of the most prestigious US research universities, intent on a career devoted to deep exploration of the fundamental laws of physics. He was rejected everywhere he applied. He should have been discouraged. But he wasn’t. Mark shrugged and said: “It’s okay. I’ll stay another year in Toronto, I’ll get a master’s degree, I’ll apply again and I’ll get in somewhere.” And that’s what happened. He went to Stanford, where, under the kind tutelage of Fred Gilman, Mark took off like a rocket. Hired to the Caltech faculty in 1982, he was a tenured full professor three years later at the age of 31, and appointed as the John A. McCone Professor of High Energy Physics while still in his 30s.

Mark liked action movies, such as those starring Arnold Schwarzenegger or Clint Eastwood. In serious moments, we would sometimes ponder together why we’re successful at what we do, and Mark would always quote Clint Eastwood as Harry Callahan in Magnum Force: “A man’s got to know his limitations.” We would both laugh, but those were words of wisdom. Mark understood what he did well as a research scientist and what he was less good at. Finding problems he could solve that would have interesting consequences for experiments that had been done or could be done was where he excelled – he did it again and again. Mark never lost his zest for calculating things, often by hand with pen and paper, his head resting on one arm with his glasses pushed up onto his forehead as he scribbled. Getting to an answer that was experimentally relevant never stopped giving him a thrill.

Mark also never lost his sense of appreciation for the teachers and mentors who had inspired and helped him. Perhaps that’s why he became such a dedicated teacher and mentor himself. It’s hard to impress Caltech students, but Mark’s lectures where extremely popular, not just for their pedagogical value but also for the humanity and humor he displayed. Students had to pay attention because otherwise one might miss the jokes, which inevitably became known as “Wisecracks.” There is even an account on X with the handle @MarkWiseSays, curated by students who want to preserve Mark’s pithy lessons in physics and in life.

For example, Mark might say: “If you really get depressed, I recommend diagonalizing a 2×2 matrix.” For physics students, this is both funny and sage advice. Or he might say. “This calculation will knock your socks off.” A cliché you might hear from anyone. But who besides Mark would then proceed to remove his shoes, rip off his socks, hurl them at the blackboard, put his shoes back on and resume lecturing?

Most famously, Mark would come to class with an ample supply of coins. He would ask the class questions, sometimes about physics and sometimes random trivia, rewarding a student who gave an answer Mark approved of by tossing a coin. At first the coins were quarters. But Mark, who had a scholarly interest in finance as well as physics, eventually felt that due to inflation he needed to upgrade to dollar coins. These are harder to come by, so it took frequent visits to the bank to make sure he wouldn’t run out. His antics in class made Mark human and approachable, and students responded. Mark felt that many Caltech students don’t fully realize how smart they are. He saw part of his job as building their self-confidence and relieving their stress.

As a colleague and mentor to graduate students and postdoctoral scholars, Mark was highly collaborative. He believed that interactions with others sparked his creativity. He was never at all pompous. I know this started early. When we were in the Harvard Society of Fellows we were obligated to have dinner with the Senior Fellows on Monday nights. It was a rather stuffy occasion. And, though I don’t think they do this anymore, after a sumptuous meal we would literally retire for brandy and cigars. Once, while puffing on his cigar after dinner, Mark had an inspiration. He gathered up a few junior fellows and led them to a theater for a movie he thought everyone should see right away. The movie was Conan the Barbarian. And everyone had a blast. That was a perfect Mark moment.

As the news about Mark has spread, accolades have poured in from physicists all over the world. He was admired not just for his scientific brilliance, but almost as much for his quirky sense of humor and his kindness. Mark was a wonderful friend to many of us. When you were with him, you were sure to laugh and feel good. He touched the lives of countless colleagues, students and friends. We miss him terribly but there are so many memories that we’ll cherish. We are all so very fortunate to have known and loved Professor Mark Wise.

Photo credit: Clara Murgui, Caltech

Doug NatelsonA few optics/metamaterials highlights from META 2026

This past week I attended META 2026, the 16th International Conference on Metamaterials, Photonic Crystals and Plasmonics, at Trinity College, Dublin.  This was the first time I've ever gone to this conference, which has grown from somewhat blurry beginnings to a ~ 900+ person annual event.  Here are a few scientific highlights:
  • Metasurfaces, built up from spatial arrays of dielectric (or sometimes semiconductor or plasmonic) resonators called "meta-atoms", have matured into very impressive, versatile tools.  In her plenary talk, Ruwen Peng from Nanjing showcased different approaches, combining angularly rotated meta-atoms ("Pencharatnam-Berry") and size-modulated meta-atoms.  The result can produce polarization-entangled photon beams, entangle photon spin and orbital angular momentum for quantum key distribution, and do full entanglement distribution over many channels.  Similarly, Federico Capasso gave a very impressive talk about the progress in the field, from visible wavelength flat optics ten years ago to compact platforms for sophisticated quantum tomography.
  • Nikolay Zheludev gave a great overview about combining measurements + machine learning estimators (e.g., here) to achieve effective optical resolution far better than conventional limits.  This can be used to make optics-base estimates of nanowire lateral displacements down to the 100 pm level, for example.  Rather than looking at the flow of energy in an optical imaging system, one can look at the flow of Fisher information regarding the object being imaged.
  • There were a series of talks throughout the meeting about chirality of optical scattering, what this means, and what it can lead to (including enantiomer-selective imaging and chemistry).  Note that it's important to distinguish between intrinsic chirality (e.g., the object scattering the light has a real structural handedness), extrinsic chirality (the object scattering the light is not chiral, but the experimental arrangement to do and measure the scattering introduces chirality into the measurement), and chirality in the fields themselves (think swirling Poynting vectors locally) that don't necessarily extend to the far field.  There are some neat probes of local effects, like this use of local polymerization.
  • Roman Quidant gave a talk about metalenses that are also optomechanical structures (e.g., use a pump beam to excite mechanical deformation of the metalens to steer the focus of a probe beam).  This lets you do some pretty neat things, like control the sign of optical forces by dynamically tuning the relative importance of momentum transfer (pushing objects with light by direct momentum kick from photons) and polarization forces (the classical optical tweezer situation where polarizable objects "seek" regions of high intensity).  This can enable feedback control to do optical cooling of trapped, levitated particles, potentially down to the quantum level.
  • Alessandra Boltasseva presented a variety of recent advances, including a look at how plasmonic ceramics like TiN and HfN have properties that can be dramatically tuned as their thickness gets down to the few-unit-cell level, a regime she and collaborators term "transdimensional" (to distinguish from atomically thin 2D van der Waals materials).  The possibility of Wigner crystallization in such systems is exciting, though disorder is a likely complication.  
  • A 4-channel wavelength division multiplexer
    made from etched Si3N4, from this paper.
    Jeremy Baumberg talked about building metamaterials out of molecularly-spaced nanoparticles, and how this has opened up real opportunities for chemical sensing based on surface-enhanced Raman and infrared absorption, as in this example.  Neat stuff.
  • There were multiple talks about metasurfaces for nonlinear optics, including one by Igal Brener on cool ways to use GaAs metasurfaces to produce entangled photon pairs via bound states in the continuum.  
  • Likewise, there were a number of presentations about inverse design, where computational tools are used to produce very funky looking structures which can act as, e.g., multichannel routers of optical signals.  Jelena Vuckovic presented an overview of this, showing how it can be done at scale to produce a chip that acts as a 1 TB/s optical router.  Structures produced this way always seem to me like some kind of eldritch geometry out of HP Lovecraft (see figure), but they work.  
As always, apologies to those whose work I didn't mention above; my note-taking was pretty uneven.

(I am trying to strike a balance between talking and educating people about science, which is basically the point of this blog, and keeping people informed/voicing some of my personal opinions about the crisis in the US research ecosystem (arguably the most consequential challenge facing US researchers today, with long-term implications that will be felt for many years).  There is still very exciting work being done in nanoscience, the physics of materials, etc. - we are just facing a future where if current trends continue the major advances may increasingly happen outside the US.)

July 17, 2026

Matt von HippelIn Defense of Reductionist Chauvinism

I don’t think people who argue about reductionism are really arguing about reductionism.

Reductionism is the idea that the behavior of big, complicated things (people, economies, ecosystems) boils down to the behavior of their smallest constituents (molecules, atoms, subatomic particles). It’s often contrasted with emergence, the idea that new rules emerge in those big, complicated systems that are more than just the rules that govern the smallest scales.

Emergence can be divided into two kinds: weak and strong. In strong emergence, the big, complicated things have their own causal powers that aren’t due to smaller things at all. This tends to get mystical, with ideas like “lifeforce” and “consciousness”. Weak emergence is much milder, and while the big complicated things are best described by their own laws, in weak emergence they are in principle still caused by laws on smaller scales.

In practice, basically no-one believes in strong emergence: it seems way too much like magic for most scientists. And basically everyone believes in weak emergence: it would be nuts to insist that economists and biologists aren’t discovering important rules that would be almost impossible to find for someone who just had physics and chemistry to work with.

So if everyone agrees, what do people argue about?

Arguments about reductionism are really arguments about attitudes. If you position yourself as a reductionist, or an emergentist, you’re defending a particular way of thinking about the world, one that privileges one scale over another. When people argue for emergence, what they really seem to be doing is opposing a kind of “reductionist chauvinism”, where people like physicists insist that their perspective is the most valuable one.

And that’s understandable, because physicists can definitely be jerks sometimes. Let it be known I am no fan of jerks.

But I think it’s worth defending reductionism, not as a philosophy, but as an attitude or perspective. Worth arguing not about whether things in practice reduce or not, but about whether reduction is a good goal, about whether science that reduces more successfully is healthier science.

Because I think it is. And the reason boils down to agreement.

Simpler systems are less controversial systems. When we write down the laws that govern subatomic particles, they’re more definite: less heuristic, more precisely specified, with fewer exceptions. That isn’t to say there’s zero controversy in these subjects, there’s even controversy between mathematicians. But the more thoroughly you can boil something down to simple rules, the easier time you have of convincing others you’re right.

In contrast, the laws of the largest scales, like psychology and biology, are deeply heuristic. They thrive on exceptions and guidelines, general tendencies without clearly defined limits. And the problem with such laws is that they can lead to intractable arguments. Different schools of thought in psychology may simply never be able to convince each other, and may just have to wait for one to die off, the aesthetic feel of one set of ideas falling out of fashion as people become preoccupied with different sorts of problems.

Every time you can reduce, you avoid those insoluble disagreements. The simpler a system you can invoke, the more you can cooperate and build off each other’s work, the less time you have to waste disagreeing, the more you can accept people with different aesthetic and philosophical preferences as just different perspectives on what are ultimately the same facts. Reductionism is a technology for peace, and one of the most powerful we have.

So yeah, I’m a bit of a chauvinist about reductionism. That’s typical of an ex-physicist, sure. But it comes from a place of concern. I want a peaceful world, where we can learn from each other and find ways to agree. And reductionism is how I get there.

Terence TaoTwo more apps: visualizing the zeta process and the motions of the heavens

I believe that the creation of visualization apps to illustrate mathematical or scientific concepts is a particularly favorable use case for modern coding agents, as many of the downside risks attached to other LLM use cases are limited:

  1. Not mission-critical. As such apps are not authorative sources of truth and only used for secondary purposes, a small positive error rate in the output can be acceptable.
  2. Stand-alone. As the applets are not destined to be incorporated into a larger codebase or literature, the technical debt incurred by delegating all the coding to an LLM agent is bounded.
  3. End product is deterministic (and sandboxed). As the applets run on a deterministic language (Javascript), are sandboxed against file or internet access, and do not make any LLM calls at run-time, security and privacy concerns are minimal, and the applet can be maintained without continued premium LLM access or resource-intensive compute.
  4. Not replacing primary skills. While deskilling is the tradeoff one accepts when relying on these tools to accelerate output, I am perfectly willing to forego the opportunity to keep my Javascript skills at a high level, as this is a tertiary skill for me at best in my chosen profession. (I continue to manually program in Lean and in Python to keep in practice with programming in general.)
  5. Not competing with humans. To my knowledge, there is no existing human effort that is being duplicated by these applets (the activity in this direction appears to have peaked two decades ago).

I would however caution against unrestricted LLM use when one or more of the above five favorable situations is not in effect.

With these points in mind, I have used such an agent to create two further apps. The first app illustrates the “zeta process” that was introduced in my recent paper with Alexeev, Barreto, Li, Lichtman, Price, Shah, and Tang, though it was first discovered by an AI. For each s > 1, the zeta distribution Z_s is a random natural number with distribution

\displaystyle {\bf P}( Z_s = k ) = \frac{1}{\zeta(s) k^s}.

It has long been known that this distribution has good number-theoretic properties: for instance, the number of times a given prime p divides Z_s has a geometric distribution of mean p^{-s}. However, the new observation is that these random variables Z_s can be chained together into a single stochastic process, which we call the “zeta process”, which is an infinite divisibility chain. I used an agent to create an app to visualize this process:

The underlying process is generated by several exponential random variables at each prime: in the above instantiation of the process, two such variables are visible at the prime p=2, and one variable at the primes p=3,5. At a given choice of s, Z_s is formed by collecting all the variables below this threshold (and for which all predecessors also lie below the threshold); in the above illustration, this amounts to one variable at each of the primes p=2,3,5, leading to Z_2 = 30 in this case. Additional visualizations in the app display the distribution of each Z_s, as well as the distribution of the hitting probability \nu_\Lambda, which among other things can be used to give a quick solution to Erdős problem #1196.

The second app is rather different in nature, and is a somewhat whimsical attempt to display the motion of the heavens, both at “human” scales of space and time, and at more “astronomical” scales (in which the motion of the planets in particular are more apparent). It is very loosely inspired by the game “Katamari Damacy“, in which one absorbs both terrestrial and celestial objects of many different scales. Here is how the app typically looks at a human scale:

And here is how it looks when one’s perspective leaves the Earth’s atmosphere:

(As I did not want to render an entire explorable world in this app, the observer in the app is only limited to changing his or her size, from a human to a creature of comparable size to the Earth itself; they cannot move horizontally on the planet.) At the largest scales of space and time, the classic orrery diagram appears:

After lengthy conversations with the agent, I was able to implement many astronomical phenomena, including phases of the Moon, the effect of Earth’s rotation against the fixed stars (though one can also stabilize one’s view against those stars to see the Earth’s rotation more directly), and so forth.

July 16, 2026

Tommaso DorigoDefining AI

Defining AI

Artificial Intelligence will probably be remembered as the least well-defined technological advancement of humankind.

Tommaso Dorigo

July 15, 2026

Terence TaoVisualizing the Gilbreath expectation sequence

One byproduct of learning how to use coding agents to create visualization apps is that it now becomes straightforward to convert any figure in one’s papers that had already been generated by code (e.g., in Python) into a more interactive, animated applet.

I can illustrate this with Figure 1 from my recent paper on the Gilbreath conjecture with Chase and Hunter, reproduced below:

This plot displays both exact and numerically simulated values of a certain poorly understood sequence c_n relating to the Gilbreath conjecture, which I will call the “Gilbreath expectation sequence” here for lack of a better name. The definition of the sequence is as follows. Consider a “Gilbreath array” which is an inverted pyramid, where the top entries are n+1 independent exponential random variables of mean 1, and all the other entries are the absolute values of the differences of the two entries immediately above it. Thanks to the visualizer app, I can quickly give an example (with n+1=6):

The left diagonal entries are then random variables; the sequence c_0,\dots,c_n are defined to be the expectation of these values. (The process is stationary, so in fact any entry on the i^{th} row will have expectation c_i.)

If one starts with the first n normalized prime gaps (which have expectation about \log n/2, and are conjecturally distributed asymptotically according to a geometric distribution), then standard conjectures (e.g., the prime tuples conjecture) predict that the k^{th} row entries should decay like c_k\log n/2, at least for small k. So the Gilbreath conjecture appears to be tied to how fast the sequence c_k decays with k.

One can in principle work out each value of c_k as an explicit rational number by performing a certain complicated multivariate integral, but in the paper we only did this for k \leq 3 (the orange line in the above figure); for the remaining k we performed a Monte Carlo simulation with 10^6 Gilbreath arrays to obtain a numerical approximation (in blue), which (as per the law of large numbers) agreed well with the theoretical values. A later calculation of Michael Ross extended the theoretical values to k \leq 6, maintaining the good fit:

The asymptotic behavior of the sequence c_n remains mysterious. Clearly, it is not monotonic; in fact we cannot even prove it is bounded. The best we could do in our paper was establish an inequality which, roughly speaking, showed that c_n cannot decay faster than 1/n.

In a recent preprint of Ross, these numerics were extended, and a rough empirical prediction

\displaystyle c_n \approx C \lambda^{s_2(n)} / n

was proposed for some constants C>0 and \lambda>1 (empirically \lambda \approx 1.17), where s_2(n) is the number of 1’s in the binary expansion of n; in particular, it is the fluctuation in this quantity s_2 that is intended to explain much of the non-monotonic behavior of c_n. These are now all displayed in the following companion applet, which was a routine matter to generate in about an hour by the coding agent (which by this point has extensive experience with creating such apps, encoded via a “skill” markdown file that it maintains):

The appearance of the quantity s_2(n) may initially appear mysterious, but it is related to Lucas’s theorem, Kummer’s theorem, and the Sierpinski gasket. Consider for instance a Gilbreath array where all the entries are zero except for a single “spike”. Then the following Sierpinski pattern emerges:

Here is what an n=64 version of this picture looks like (with the spike positioned at the 32th entry):

The number of 1s in the k^{th} row is then 2^{s_2(k)} (if we index the rows starting from zero), which is at least of the same shape as the empirical prediction, albeit with different constants. (This sequence is also known as Gould’s sequence.)

Numerically, we seem to observe fragments of Sierpinski gaskets being generated before decaying (often due to “collisions” with other gaskets):

However, it is not clear to me at all what the asymptotic probabilistic model should be, even heuristically; it does not resemble any random shape model that I am familiar with. But perhaps there are readers more expert in probability theory or statistical physics who may be able to suggest such an asymptotic limit?

Tommaso DorigoBlack Cats in Dark Rooms

Black Cats in Dark Rooms

[As a meta-text, this blog column contains a large number of posts written over the years, that describe new results from physics analyses, among other articles commenting on news in science and outside.

Tommaso Dorigo
Categories

July 13, 2026

Scott Aaronson Held Prize call for nominations (+ call for postdocs)

Here at the National Academy of Sciences, it seems that my first job is to serve on the selection committee for the prestigious Michael and Sheila Held Prize in combinatorial and discrete optimization and related areas. The committee chair, my former MIT colleague Madhu Sudan (now at Harvard), invited me to share the following message here on Shtetl-Optimized. (I’d add: put in the effort to nominate someone, and you can actually influence how things go!)

Dear Colleagues

I am writing to seek nominations for the 2027 Michael and Sheila Held Prize. The scope of the prize and nomination needs are described below. If you intend to nominate someone I would appreciate a heads up by email to madhu@cs.harvard.edu one month before the deadline (so email by Sept 8, 2026) to let me know your nomination is coming. (We may also reach out to you in response to coordinate multiple/overlapping nominations.)

The Held prize honors outstanding, innovative, creative, and influential research in the areas of combinatorial and discrete optimization, or related parts of computer science, such as the design and analysis of algorithms and complexity theory. This $100,000 prize is intended to recognize recent work (defined as published within the last eight years, i.e., on or after October 6, 2018).

All nominations must be submitted online by Monday, October 5, 2026 and include:

1. Nomination letter describing the candidate’s work and why he or she should be selected for the award. No more than three (3) pages.

2. Curriculum vitae. No more than two (2) pages.

3. Bibliography listing no more than twelve (12) of the nominee’s most significant publications.

4. Suggested citation. A 50-word summary stating why the nominee should be considered for this award.

5. Two letters of support. No more than one letter of support can be written by someone of the same primary work institution as the nominee.

The Held Prize is given to a person or a set of persons, as supported by a paper or a body of work. Unless otherwise stated, preference will be given to scientists who may be earlier in their careers or those whose work has not been recognized by other prizes or awards. Nomination restrictions can be found here. Joint nominations will only be considered when nominees have collaborated closely on the paper to be recognized by the award. If nominating multiple individuals for a paper with additional authors, please clearly explain the reason for nominating those chosen, as well as the reason for excluding other collaborators, if applicable. 

Please feel free to circulate this call further within your department

Best
Madhu Sudan, on behalf of The Michael and Sheila Held Prize Selection Committee


And while I have your attention, a second CS theory announcement: David Soloveichik, my wonderful friend and colleague in UT Austin’s Electrical and Computer Engineering Department, has funding for a postdoc for 1-2 years, to work on the thermodynamics of computation here at UT. This is a topic that I’ve been trying to learn more about as well, so I might get involved too! David writes, “the big picture is to think of thermodynamics (energy dissipation / entropy production) as CS complexity measures like time and space usage.” If you’re on the postdoc market and this sounds potentially up your alley, email David to learn more.

July 12, 2026

n-Category Café Octonions and the Standard Model (Part 13)

When Lee and Yang suggested that the laws of physics might not be invariant under spatial reflection — that there’s a fundamental difference between left and right — Pauli was skeptical. In a letter to Victor Weisskopf in January 1957, he wrote:

“Ich glaube aber nicht, daß der Herrgott ein schwacher Linkshänder ist.”

(I do not believe that the Lord is a weak left-hander.)

But just two days after Pauli wrote this letter, Chien-Shiung Wu’s experiment confirmed that Lee and Yang were correct. There’s an inherent asymmetry in nature.

We can trace this back to how the ‘left-handed’ fermions and antifermions live in a different representation of the Standard Model gauge group than the right-handed ones. And when we try to build grand unified theories that take this into account, we run into the fact that while we can fit the Standard Model gauge group into Spin(10)\text{Spin}(10) in various ways, not all these ways produce the required asymmetry. There’s a way where it fits into Spin(9)\text{Spin}(9), which is too symmetrical to work… and alas, this one has a nice octonionic description!

To keep things simple I’ll explain this by focusing, not on the whole Standard Model gauge group, but its subgroup SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). Here is a theorem proved by Will Sawin in response to a question of mine on MathOverflow:

Theorem 10. There are exactly two conjugacy classes of subgroups of Spin(10)\text{Spin}(10) that are isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). One of them has a representative that is a subgroup of Spin(9)Spin(10)\text{Spin}(9) \subset \text{Spin}(10), while the other does not.

I’ll describe representatives of these two subgroups; then I’ll say a bit about how they show up in physics, and then I’ll show you Sawin’s proof.

We can get both subgroups in a unified way! There’s always an inclusion

SO(m)×SO(n)SO(m+n) \text{SO}(m) \times \text{SO}(n) \to \text{SO}(m+n)

and taking double covers of each group we get a 2-1 homomorphism

Spin(m)×Spin(n)Spin(m+n) \text{Spin}(m) \times \text{Spin}(n) \to \text{Spin}(m+n)

In particular we have

Spin(4)×Spin(6)Spin(10) \text{Spin}(4) \times \text{Spin}(6) \to \text{Spin}(10)

so composing with the exceptional isomorphisms:

Spin(4)SU(2)×SU(2),Spin(6)SU(4) \text{Spin}(4) \cong \text{SU}(2) \times \text{SU}(2), \qquad \text{Spin}(6) \cong \text{SU}(4)

we get a 2-1 homomorphism

k:SU(2)×SU(2)×SU(4)Spin(10) k \colon \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \to \text{Spin}(10)

Now, there are three obvious ways to include SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) in SU(2)×SU(2)×SU(4)\text{SU}(2) \times \text{SU}(2) \times \text{SU}(4). There is an obvious inclusion

j:SU(3)SU(4) j \colon \text{SU}(3) \hookrightarrow \text{SU}(4)

but there are three obvious inclusions

,r,δ:SU(2)SU(2)×SU(2) \ell, r, \delta \colon \text{SU}(2) \hookrightarrow \text{SU}(2) \times \text{SU}(2)

namely the left one:

:SU(2) SU(2)×SU(2) g (g,1) \begin{array}{ccc} \ell \colon \text{SU}(2) &\to& \text{SU}(2) \times \text{SU}(2) \\ g & \mapsto & (g,1) \end{array}

the right one:

r:SU(2) SU(2)×SU(2) g (1,g) \begin{array}{ccc} r \colon \text{SU}(2) &\to& \text{SU}(2) \times \text{SU}(2) \\ g & \mapsto & (1,g) \end{array}

and the diagonal one:

δ:SU(2) SU(2)×SU(2) g (g,g) \begin{array}{ccc} \delta \colon \text{SU}(2) &\to& \text{SU}(2) \times \text{SU}(2) \\ g & \mapsto & (g,g) \end{array}

Combining these with our earlier maps, we actually get a one-to-one map from SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) to Spin(10)\text{Spin}(10). So we get three subgroups of Spin(10)\text{Spin}(10), all isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3):

  • There’s the left subgroup G G_\ell, which is the image of this composite homomorphism:

SU(2)×SU(3)×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\ell \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

  • There’s the diagonal subgroup G δG_\delta, which is the image of this:

SU(2)×SU(3)δ×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\delta \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

  • And there’s the right subgroup G rG_r, which is the image of this:

SU(2)×SU(3)r×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{r \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

The left and right subgroups are actually conjugate, but the diagonal one is truly different! We’ll prove this by taking a certain representation of Spin(10)\text{Spin}(10), called the Weyl spinor representation, and restricting it to those two subgroups. We’ll get inequivalent representations of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). This proves the two subgroups aren’t conjugate.

This argument is also interesting for physics. When restrict to the left subgroup, we get a representation of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) that matches what we actually see for one generation of fermions! This is the basis of the so-called SO(10)\text{SO}(10) grand unified theory, which should really be called the Spin(10)\text{Spin}(10) grand unified theory.

(In fact this works not only for SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) but for the whole Standard Model gauge group, which is larger. I’m focusing on SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) just because it makes the story simpler.)

When we restrict the Weyl spinor representation to the diagonal subgroup, we get a representation of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) that is not physically correct. Unfortunately, it’s the diagonal subgroup that shows up in several papers connecting the Standard Model gauge group to the octonions. I plan to say a lot more about this later.

The left subgroup

Let’s look at the left subgroup G G_\ell, the image of this composite:

SU(2)×SU(3)×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\ell \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

Spin(10)\text{Spin}(10) has a 32-dimensional unitary representation called the ‘Dirac spinor’ representation. This representation is really on the exterior algebra Λ 5\Lambda \mathbb{C}^5. It’s the direct sum of two irreducible parts, the even grades and the odd grades:

Λ 5Λ even 5Λ odd 5 \Lambda \mathbb{C}^5 \cong \Lambda^{\text{even}} \mathbb{C}^5 \oplus \Lambda^{\text{odd}} \mathbb{C}^5

Physicists call these two irreducible representations ‘right- and left-handed Weyl spinors’, and denote them as 16\mathbf{16} and 16*\mathbf{16}\ast since they’re 16-dimensional and one is the dual of the other.

Let’s restrict the 16\mathbf{16} to the left subgroup G G_\ell and see what we get.

To do this, first we can restrict the 16\mathbf{16} along kk and get

214124* \mathbf{2} \otimes \mathbf{1} \otimes \mathbf{4} \; \oplus \; \mathbf{1} \otimes \mathbf{2} \otimes \mathbf{4}\ast

Here 1\mathbf{1} is the trivial representation of SU(2)\text{SU}(2), 2\mathbf{2} is the tautologous representation of SU(2)\text{SU}(2), and 4\mathbf{4} is the tautologous rep of SU(4)\text{SU}(4).

Then let’s finish the job by restricting this representation along ×j\ell \times j. Restricting the 4\mathbf{4} of SU(4)\text{SU}(4) to SU(3)\text{SU}(3) gives 31\mathbf{3} \oplus \mathbf{1}: the sum of the tautologous representation of SU(3)\text{SU}(3) and the trivial representation. Restricting 21\mathbf{2} \otimes \mathbf{1} to the left copy of SU(2)\text{SU}(2) gives the tautologous representation 2\mathbf{2}, while restricting 12\mathbf{1} \otimes \mathbf{2} to this left copy gives 11\mathbf{1} \oplus \mathbf{1}: the sum of two copies of the trivial representation. All in all, we get this representation of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3):

2(31)(11)(3*1) \mathbf{2} \otimes (\mathbf{3} \oplus \mathbf{1}) \; \oplus \; (\mathbf{1} \oplus \mathbf{1}) \otimes (\mathbf{3}\ast \oplus \mathbf{1})

This is what we actually see for one generation of left-handed fermions and antifermions in the Standard Model! The representation 31\mathbf{3} \oplus \mathbf{1} describes how the left-handed fermions in one generation transform under SU(3)\text{SU}(3): 3 colors of quark and one ‘white’ lepton. The representation 3*1\mathbf{3}\ast \oplus \mathbf{1} does the same for the left-handed antifermions. The left-handed fermions form an isospin doublet, giving us the 2\mathbf{2}, while the left-handed antifermions have no isospin, giving us the 11\mathbf{1} \oplus \mathbf{1}.

This strange lopsidedness is a fundamental feature of the Standard Model.

The right subgroup would work the same way, up to switching the words ‘left-handed’ and ‘right-handed’. And by Theorem 10, the left and right subgroups must be conjugate in Spin(10)\text{Spin}(10), because now we’ll see one that’s not conjugate to either of these.

The diagonal subgroup

Consider the diagonal subgroup G δG_\delta, the image of this composite:

SU(2)×SU(3)δ×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\delta \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

Let’s restrict the 16\mathbf{16} to G δG_\delta.

To do this, first let’s restrict the 16\mathbf{16} along k:SU(2)×SU(2)×SU(4)Spin(10)k \colon \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \to \text{Spin}(10) and get

214124* \mathbf{2} \otimes \mathbf{1} \otimes \mathbf{4} \; \oplus \; \mathbf{1} \otimes \mathbf{2} \otimes \mathbf{4}\ast

as before. Then let’s restrict this representation along δ×j\delta \times j. The SU(3)\SU(3) part works as before, but what happens when we restrict 21\mathbf{2} \otimes \mathbf{1} or 12\mathbf{1} \otimes \mathbf{2} along the diagonal map δ:SU(2)SU(2)×SU(2)\delta \colon \text{SU}(2) \to \text{SU}(2) \times \text{SU}(2)? We get 2\mathbf{2}. So, this is the representation of G δG_\delta that we get:

2(31)2(3*1) \mathbf{2} \otimes (\mathbf{3} \oplus \mathbf{1}) \; \oplus \; \mathbf{2} \otimes (\mathbf{3}\ast \oplus \mathbf{1})

This is not good for the Standard Model. It describes a more symmetrical universe than ours, where both left-handed fermions and antifermions transform as doublets under SU(2)\text{SU}(2).

The fact that we got a different answer this time proves that G G_\ell and G δG_\delta are not conjugate in Spin(10)\text{Spin}(10). So to complete the proof of Theorem 10, we only need to prove

  1. Every subgroup of Spin(10)\text{Spin}(10) isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) is conjugate to G G_\ell or G δG_\delta.

  2. G δG_\delta is conjugate to a subgroup of Spin(9)Spin(10)\text{Spin}(9) \subset \text{Spin}(10), but G G_\ell is not.

I’ll prove 2, and then I’ll turn you over to Will Sawin to do the rest.

Why the diagonal subgroup fits in Spin(9)\text{Spin}(9)

Every rotation of n\mathbb{R}^n extends to a rotation of n+1\mathbb{R}^{n+1} that leaves the last coordinate fixed, so we get an inclusion SO(n)SO(n+1)\text{SO}(n) \hookrightarrow \text{SO}(n+1), which lifts to an inclusion of the double covers, Spin(n)Spin(n+1)\text{Spin}(n) \hookrightarrow \text{Spin}(n+1). Since we have exceptional isomorphisms

Spin(3)SU(2),Spin(4)SU(2)×SU(2) \text{Spin}(3) \cong \text{SU}(2), \qquad \text{Spin}(4) \cong \text{SU}(2) \times \text{SU}(2)

it’s natural to ask how the inclusion Spin(3)Spin(4)\text{Spin}(3) \hookrightarrow \text{Spin}(4) looks in these terms. And the answer is: it’s the diagonal map! In other words, we have a commutative diagram

SU(2) Spin(3) δ SU(2)×SU(2) Spin(4) \begin{array}{ccc} \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(3) \\ \delta \downarrow & & \downarrow \\ \text{SU}(2) \times \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(4) \end{array}

Now, we can easily fit this into a larger commutative diagram involving some natural maps Spin(m)×Spin(n)Spin(m+n)\text{Spin}(m) \times \text{Spin}(n) \to \text{Spin}(m+n) and Spin(n)Spin(n+1)\text{Spin}(n) \to \text{Spin}(n+1):

SU(2) Spin(3) Spin(3)×Spin(6) Spin(9) δ SU(2)×SU(2) Spin(4) Spin(4)×Spin(6) Spin(10) \begin{array}{ccccccc} \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(3) & \to & \text{Spin}(3) \times \text{Spin}(6) & \to & \text{Spin}(9) \\ \delta \downarrow & & \downarrow & & \downarrow & & \downarrow \\ \text{SU}(2) \times \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(4) & \to & \text{Spin}(4) \times \text{Spin}(6) & \to & \text{Spin}(10) \end{array}

We can simplify this diagram using the isomorphism Spin(6)SU(4)\text{Spin}(6) \cong \text{SU}(4):

SU(2)×SU(4) Spin(9) δ×1 SU(2)×SU(2)×SU(4) Spin(10) \begin{array}{ccccccc} \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(9) \\ \delta \times 1 \downarrow & & \downarrow \\ \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(10) \end{array}

and then we can use our friend the inclusion j:SU(3)SU(4)j \colon \text{SU}(3) \to \text{SU}(4):

SU(2)×SU(3) 1×j SU(2)×SU(4) Spin(9) δ×1 SU(2)×SU(2)×SU(4) Spin(10) \begin{array}{ccccccc} \text{SU}(2) \times \text{SU}(3) & \xrightarrow{1 \times j} & \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(9) \\ & & \delta \times 1 \downarrow & & \downarrow \\ & & \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(10) \end{array}

This shows that the diagonal subgroup G δG_\delta of Spin(10)\text{Spin}(10) is actually a subgroup of Spin(9)\text{Spin}(9)!

Why the left subgroup does not fit in Spin(9)\text{Spin}(9)

The three-fold way is a coarse classification of irreducible complex representations of compact Lie group. Every such representation is of one and only one of these three kinds:

1) not self-dual: not isomorphic to its dual,

2a) orthogonal: isomorphic to its dual via an invariant nondegenerate symmetric bilinear form, also called an orthogonal structure,

2b) symplectic: isomorphic to its dual via an invariant nondegenerate antisymmetric bilinear form, also called a symplectic structure.

I’ve written about how these three cases are related to the division algebras ,\mathbb{C}, \mathbb{R} and \mathbb{H}, respectively:

A complex representation is orthogonal iff it’s the complexification of a representation on a real vector space, and symplectic iff it’s the underlying complex representation of a representation on a quaternionic vector space.

But we don’t need most of this yet. For now we just need to know one fact: when nn is odd, every irreducible representation of Spin(n)\text{Spin}(n), and thus every representation of this Lie group, is self-dual: that is, isomorphic to its dual. In particular this is true of Spin(9)\text{Spin}(9).

Why does this matter? Assume the left subgroup G Spin(10)G_\ell \subset \text{Spin}(10) is a subgroup of Spin(9)\text{Spin}(9). When we restrict the Weyl spinor representation of Spin(10)\text{Spin}(10) to Spin(9)\text{Spin}(9) it will be self-dual, like every representation of Spin(9)\text{Spin}(9). Then when we restrict this representation further to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) it must still be self-dual, since the restriction of a self-dual representation is clearly self-dual.

However, we know this representation is

2(31)(11)(3*1) \mathbf{2} \otimes (\mathbf{3} \oplus \mathbf{1}) \; \oplus \; (\mathbf{1} \oplus \mathbf{1}) \otimes (\mathbf{3}\ast \oplus \mathbf{1})

and this is not self-dual, since 1*1\mathbf{1}\ast \cong \mathbf{1} and 2*2\mathbf{2}\ast \cong \mathbf{2} but 3*3\mathbf{3}\ast \ncong \mathbf{3}.

So, it must be that G G_\ell is not a subgroup of Spin(9)\text{Spin}(9).

Proof of Theorem 10

To complete the proof of Theorem 10 we just need to see why there are just two conjugacy classes of subgroups of Spin(10)\text{Spin}(10) isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). But in fact Will Sawin proved a stronger result! He was answering this question of mine:

Define the Standard Model gauge group to be S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)), the subgroup of SU(5)\text{SU}(5) consisting of block diagonal matrices with a 2×22 \times 2 block and then a 3×33 \times 3 block. (This is isomorphic to the quotient of U(1)×SU(2)×SU(3)\text{U}(1) \times \text{SU}(2) \times \text{SU}(3) by the subgroup of elements (α,α 3,α 2(\alpha, \alpha^{-3}, \alpha^2) where α\alpha is a 6th root of unity.)

Up to conjugacy, how many subgroups isomorphic to the Standard Model gauge group does Spin(10)\text{Spin}(10) have?

This question is relevant to grand unified theories of particle physics, as explained here:

This paper focuses on one particular copy of S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)) in Spin(10)\text{Spin}(10), given as follows. By definition we have an inclusion S(U(2)×U(3))SU(5)\text{S}(\text{U}(2) \times \text{U}(3)) \hookrightarrow \text{SU}(5), and we also have an inclusion SU(5)Spin(10)\text{SU}(5) \hookrightarrow \text{Spin}(10) because for any nn we have an inclusion SU(n)SO(2n)\text{SU}(n) \hookrightarrow \text{SO}(2n), and SU(n)\text{SU}(n) is simply connected so this gives a homomorphism SU(n)Spin(2n)\text{SU}(n) \hookrightarrow \text{Spin}(2n).

However I think there is also an inclusion S(U(2)×U(3))Spin(9)\text{S}(\text{U}(2) \times \text{U}(3)) \hookrightarrow \text{Spin}(9), studied by Krasnov:

Composing this with Spin(9)Spin(10)\text{Spin}(9) \hookrightarrow \text{Spin}(10), this should give another inclusion S(U(2)×U(3))Spin(10)\text{S}(\text{U}(2) \times \text{U}(3)) \hookrightarrow \text{Spin}(10), and I believe this one is ‘truly different from’ — i.e., not conjugate to — the first one I mentioned.

So I believe my current answer to my question is “at least two”. But that’s not good enough.

Sawin’s answer relies heavily on the 3-fold way — that’s why I told you that stuff about orthogonal and symplectic representations. When we embed the group SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) in Spin(10)\text{Spin}(10), we are automatically giving this group an orthogonal 10-dimensional representation, thanks to the map Spin(10)SO(10)\text{Spin}(10) \to \text{SO}(10). We can classify the possibilities.

He writes:

There are infinitely many embeddings. However, all but one of them is “essentially the same as” the one you studied as they become equal to the one you studied on restriction to SU(2)×SU(3)\text{SU}(2)\times \text{SU}(3). The remaining one is the one studied by Krasnov.

I follow the strategy suggested by Kenta Suzuki.

SU(3)\text{SU}(3) has irreducible representations of dimensions 1,3,3,6,8,6,10,101,3,3,6,8,6, 10, 10, and higher dimensions. The 1010-dimensional ones are dual to each other, as are the 66-dimensional ones, so they can’t appear. The 33-dimensional ones are dual to each other and can only appear together. So the only 1010-dimensional self-dual representations of SU(3)\text{SU}(3) decompose as irreducibles as 8+1+18+1+1, 3+3+1+1+1+13+3+1+1+1+1, or ten 11s. All of these are orthogonal because the 8-dimensional representation is orthogonal. However, the ten 11s cannot appear because then SU(3)\text{SU}(3) would act trivially.

A representation of SU(3)×SU(2)\text{SU}(3) \times \text{SU}(2) is a sum of tensor products of irreducible representations of SU(3)\text{SU}(3) and irreducible representations of SU(2)\text{SU}(2). Restricted to SU(3)\text{SU}(3), each tensor product splits into a sum of copies of the same irreducible representation. So SU(2)\text{SU}(2) can only act nontrivially when the same representation appears multiple times. Since the 3+33+3 is two different 33-dimensional representation, only the 11-dimensional representation can occur twice. Thus, our 10-dimensional orthogonal representation of SU(3)×SU(2)\text{SU}(3) \times \text{SU}(2) necessarily splits as either the 88-dimensional adjoint repsentation of SU(3)\text{SU}(3) plus a 22-dimensional orthogonal representation of SU(2)\text{SU}(2) or the 66-dimensional sum of standard and conjugate [i.e., dual] representations of SU(3)\text{SU}(3) plus a 44-dimensional orthogonal representation of SU(2)\text{SU}(2). However, SU(2)\text{SU}(2) has a unique nontrivial representation of dimension 22 and it isn’t orthgonal, so only the second case can appear. SU(2)\text{SU}(2) has representations of dimension 1,2,3,41,2,3,4 of which the 22 and 44-dimensional ones are symplectic and so must appear with even multiplicity in any orthogonal representation, so the only nontrivial 44-dimensional orthogonal ones are 2+22+2 or 3+13+1.

So there are two ten-dimensional orthogonal representations of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) that are nontrivial on both factors, those being the sum of two different 33-dimensional irreducible representations of SU(3)\text{SU}(3) with either two copies of the two-dimensional irreducible representation of SU(2)\text{SU}(2) or the three-dimensional and the one-dimensional irreducible representation of SU(2)\text{SU}(2). The orthogonal structure is unique up to isomorphisms, so these give two conjugacy classes of homomorphisms SU(2)×SU(3)SO(10)\text{SU}(2) \times \text{SU}(3) \to SO(10) and thus two conjugacy classes of homomorphisms SU(2)×SU(3)Spin(10)\text{SU}(2) \times \text{SU}(3) \to \text{Spin}(10). The first one corrresponds to the embedding you studied while only the second one restricts to Spin(9)\text{Spin}(9) so indeed these are different.

To understand how to extend these to S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)), I consider the centralizer of the representation within Spin(10)\text{Spin}(10). Since the group is connected, this is the same as the centralizer of its Lie algebra, which is therefore the inverse image of the centralizer in SO(10)\text{SO}(10). Now there is a distinction between the two examples because the example with irrep dimensions 3+3+2+23+3+2+2 has centralizer with identity component U(1)×SU(2)\text{U}(1) \times \text{SU}(2) while the example with irrep dimensions 3+3+3+13+3+3+1 has centralizer with identity component U(1)\text{U}(1). In the second case, the image of U(2)×U(3)\text{U}(2) \times \text{U}(3) must be the image of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) times the centralizer of the image of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3), so this gives a unique example, which must be the one considered by Krasnov.

In the first case, we can restrict attention to a torus U(1)×U(1)\text{U}(1) \times \text{U}(1) in SU(2)×SU(2)\text{SU}(2) \times \text{SU}(2). The center of S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)) maps to a one-dimensional subgroup of this torus, which can be described by a pair of integers. Explicitly, given a two-by-two-unitary matrix AA and a three-by-three unitary matrix BB with det(A)det(B)=1\det(A) \det(B) =1, we can map to U(5)\text{U}(5) by sending (A,B)(A,B) to Aγ aBγ bA \gamma^a \oplus B \gamma^b where γ=det(A)=det(B) 1\gamma = \det (A) = \det(B)^{-1}, and then map from U(5)\text{U}(5) to SO(10)SO(10). This lifts to the spin group if and only if the determinant in U(5)\text{U}(5) is a perfect square. The determinant is γ 1+2a1+3b=γ 2a+3b\gamma^{ 1 + 2a - 1 + 3b} = \gamma^{2a+3b} so a lift exists if and only if bb is even.

The only possible kernel of this embedding is the scalars. The scalar A=λ 3I 2,B=λ 2I 3A = \lambda^3 I_2, B = \lambda^{-2} I_3 maps to λ 3+6aI 2λ 2+6bI 3\lambda^{3+ 6a} I_2 \oplus \lambda^{-2 + 6b} I_3 and so the kernel is trivial if and only if gcd(3+6a,2+6b)=1\gcd(3+6a,-2 + 6b)=1.

However, there are infinitely many integer solutions a,ba,b to gcd(3+6a,2a+6b)=1\gcd(3+6a,-2a+6b)=1 with bb even (in fact, a random aa and even bb works with probability 9/π 29/\pi^2), so this gives infinitely many examples.


  • Part 1. How to define octonion multiplication using complex scalars and vectors, much as quaternion multiplication can be defined using real scalars and vectors. This description requires singling out a specific unit imaginary octonion, and it shows that octonion multiplication is invariant under SU(3)\mathrm{SU}(3).
  • Part 2. A more polished way to think about octonion multiplication in terms of complex scalars and vectors, and a similar-looking way to describe it using the cross product in 7 dimensions.
  • Part 3. How a lepton and a quark fit together into an octonion — at least if we only consider them as representations of SU(3)\mathrm{SU}(3), the gauge group of the strong force. Proof that the symmetries of the octonions fixing an imaginary octonion form precisely the group SU(3)\mathrm{SU}(3).
  • Part 4. Introducing the exceptional Jordan algebra 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}): the 3×33 \times 3 self-adjoint octonionic matrices. A result of Dubois-Violette and Todorov: the symmetries of the exceptional Jordan algebra preserving their splitting into complex scalar and vector parts and preserving a copy of the 2×22 \times 2 adjoint octonionic matrices form precisely the Standard Model gauge group.
  • Part 5. How to think of 2×22 \times 2 self-adjoint octonionic matrices as vectors in 10d Minkowski spacetime, and pairs of octonions as left- or right-handed spinors.
  • Part 6. The linear transformations of the exceptional Jordan algebra that preserve the determinant form the exceptional Lie group E 6\mathrm{E}_6. How to compute this determinant in terms of 10-dimensional spacetime geometry: that is, scalars, vectors and left-handed spinors in 10d Minkowski spacetime.
  • Part 7. How to describe the Lie group E 6\mathrm{E}_6 using 10-dimensional spacetime geometry. This group is built from the double cover of the Lorentz group, left-handed and right-handed spinors, and scalars in 10d Minkowski spacetime.
  • Part 8. A geometrical way to see how E 6\mathrm{E}_6 is connected to 10d spacetime, based on the octonionic projective plane.
  • Part 9. Duality in projective plane geometry, and how it lets us break the Lie group E 6\mathrm{E}_6 into the Lorentz group, left-handed and right-handed spinors, and scalars in 10d Minkowski spacetime.
  • Part 10. Jordan algebras, their symmetry groups, their invariant structures — and how they connect quantum mechanics, special relativity and projective geometry.
  • Part 11. Particle physics on the spacetime given by the exceptional Jordan algebra: a summary of work with Greg Egan and John Huerta.
  • Part 12. The bioctonionic projective plane and its connections to algebra, geometry and physics.
  • Part 13. Two ways to embed SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) in Spin(10)\text{Spin}(10), and their consequences for particle physics.

July 10, 2026

Scott Aaronson Announcing BQP Partners: my and my brother’s new angel-investing venture

As I’ve written before, these past couple years I’ve often felt like the last remaining person in either quantum computing or AI who lacked a stake in some startup company whose valuation is right now shooting into interstellar space. My academic colleagues, including the ones who seemed the most singleminded about quantum oracle separations and other gloriously useless pursuits? One by one, like in a zombie movie, I learn that they too have now launched startups, and invariably raised tens of millions of dollars, for the sorts of ideas we might’ve idly traded at coffee breaks back in the day, before getting back to our real work.

So why didn’t I join this rollicking party? Partly because of a lifelong fear that, the instant my self-worth became tied to how much money I made, I’d need to humble myself before people who bluster and bully and lie and hype and conceal … yet who nevertheless succeed at becoming orders of magnitude richer than me. I’ve been terrified of even starting down that road, of whether I’d still be myself at the end of it.

It’s also partly that I can’t stand failure, or regret, or being wrong. Of course, as an academic researcher I also fail, and regret things, and am wrong constantly—but there it feels tolerable, because normally I can tell myself that it’s all just down to my inborn limitations. After all, if I could’ve solved the major open problem that someone else solved, or written the brilliant book that someone else wrote, then presumably I would’ve done it!

Clearly, though, I could’ve mined bitcoin in 2010. I could’ve gotten an early stake in Amazon or Google. It’s not even like those ideas never crossed my mind. I just … didn’t act on them, for some reason. (But even if I had, I’d probably just be full of regret that I hadn’t done even more.) Thus, my only way to avoid paralyzing regrets, has been to tell myself constantly that I’m not in the forecasting or money-making businesseses in the first place.

It helped that, insofar as I’m shallow or covetous, insofar as I’ve desired things of this world rather than insight or eternal truth, it’s never really been money that I cared about, but just being respected and liked. Elon Musk is the richest man on earth, but also one of the most despised—which isn’t a bargain that I could imagine ever appealing to me.

Plus, when I actually meet billionaires, I don’t find myself envious of their mansions or cars or anything else that they have; I don’t feel like such things would make my life any happier. Maybe I slightly envy their ability to fund the causes they care about, or their professional staffs who relieve them of drudgery, but mostly I envy the way their wealth announces, to whatever extent it does: “I was right when others weren’t.” Again, though, I’ve never trusted the world to cause me to be right about the future valuations of companies or anything similar, so I’ve settled for having been right about PostBQP and algebrization and BosonSampling.

The bottom line is that I made a choice decades ago to forgo trying to get rich, no matter how many of my friends did the same, and to strive instead to discover and tell the truth—to be a professor, a blogger, a jokester, and an “objective” arbiter and commentator. “Then, surely, everyone will like me!” my internal monologue went. “Then, surely, they’ll be grateful for all the free service I’ve rendered them—for decades of blogging, without once so much as asking for a donation or running an ad!”


HAHAHAHAHAHA.

As any regular reader will know, my attempts to be loved as a blogger backfired pretty spectacularly. Or rather: they did lead to thousands of strangers liking me (and I’m grateful for every last one of you), but they also led to probably an order of magnitude more strangers hating me, and congregating on Reddit and Twitter and elsewhere to discuss how badly I suck. And of course, trying to shift that balance by writing what people want to hear, rather than what I actually believe, was never within my realistic option set.

In the startup context, it didn’t matter how carefully I avoided taking a direct stake for or against any of the companies I blogged about. People on Twitter simply assumed that I had a stake—for example, that I must’ve shorted D-Wave or IonQ, or invested in their competitors, or had equity in AI companies. For why else would anyone write what I wrote?

Amusingly, my attackers here typically did have precisely the conflicts-of-interest that they falsely accused me of having, but that was never at issue; only my imaginary conflicts-of-interest were. Even as the Scott-haters greedily filled their pockets (or tried to), I alone needed to keep turning my pockets out to prove that they were still empty.


So then, screw it! In partnership with my brother David Aaronson, who’s long done investing professionally, and on David’s guidance and encouragement, I’m hereby embarking on a new policy.

Namely: when I hear about a brand-new startup that sounds relevant to my interests—in quantum, AI, or anything else—and I like and trust the founders (ideally, because of their previous academic research work), David and I will often make a small seed investment if the founders are open to it. Or, of course, we might become advisors or get involved in some other way.

In fact, David and I are launching BQP Partners—the link goes to our AngelList, where you can read about how to invest with us if you’re interested. (See also whether you can spot any differences between David’s writing style and preoccupations and mine!)

So far, David and I are investing in:

I have little doubt that more potential investments will come our way very soon (some, probably, as a direct result of this post).

Crucially, I can handle my burden of regret—the “why didn’t I do this much earlier, if I was going to do it at all?” question—by telling myself that friends of mine were not founding companies left and right until very recently. I can also tell myself that I’m doing this less as a bet about the future (in which case … what if I’m wrong?), than simply as a way to support brilliant colleagues doing things that I genuinely admire.

When I blog about a company, I’ll always disclose if I have a financial position that presents a clear conflict of interest, so you can judge for yourself whether to listen to me. (Although, if that’s the sort of thing you’d demand, then you probably weren’t listening to me in the first place, were you?)

Having reflected on it a lot these past few months, I’m happy with my new policy and with my and David’s new venture, and I’m curious to see where it goes. I’m at peace with the possibility that we’ll lose our shirts, but I’m even at peace with a more disturbing possibility—that we’ll make millions and then people will scream at me online for being a sellout, a hack, and a shill. Those people, as I’ve learned, were going to scream at me anyway.

Matt von HippelAmplitudes 2026, Part II

This is a continuation of my conference coverage from last week. The same warnings apply: this is much more technical than my usual posts, readers beware!

Last week I covered the talks from Monday through Wednesday, so I’ll jump right in here with Thursday morning, where the first speaker was Francesco Riva, who after a brief ad for his new board game Tutti Quantum gave a review of positivity constraints, first-principles restrictions on quantum field theories based on their behavior at high energies. After covering some of the research program’s successes like arguments against Galileons and massive gravity, he talked about how new methods allow one to take into account the possibility of loops of massless particles, essentially by invoking a formal version of the idea that real experiments have finite size. He was followed by Grant Remmen, who talked about his work deriving string theory-like amplitudes from increasingly minimal assumptions. While one can always quibble with the assumptions they impose, I do find it encouraging that they are now managing to do this game with both gravity and gauge theory, and with five-particle amplitudes, not just four.

Paul Heslop then covered applications of bootstrap techniques to non-supersymmetric theories, where analytic superspace still finds a way to be useful. Tomasz Taylor covered progress in calculating Yang-Mills amplitudes in de Sitter space. Andrea Puhm and Nima Arkani-Hamed don’t have slides online yet: Puhm’s title suggests her talk was part of the celestial holography field, while Nima’s was likely similar to his talk at Lancefest, where he talked about calculations of amplitudes in the limit of a very large number of loops (represented in the field with a capital L) or very large numbers of particles (represented in the field with a lower-case n). Given the different context, I’m guessing he left out the self-effacing jokes where he was “little n” and Lance was “big L”, though I’m hoping he at least mentioned his students had been checking their results against what they called the “lanswer”.

The evening ended with a gong show, which for the non-initiated is a series of short student talks with a strict time limit (hence the gong). I’m not quite so intrepid as to read all of the slides for these, commenters who attended are welcome to highlight special examples.

Friday began with a talk by Agnese Bissi, who reviewed the connection between holographic correlators and amplitudes in AdS space. I liked her emphasis on this as a lab to find nice new representations of amplitudes, and her summary at the end of the current frontiers. Axel Kleinschmidt followed with a talk on one-loop string amplitudes, where he and his collaborators have gotten gradually more proficient at manipulating the rich structure of elliptic functions that make an appearance. Piotr Tourkine talked about his work using the S-matrix bootstrap to find scattering amplitudes in higher dimensions, a context where accounting for thresholds presented a new challenge. This is a problem people are approaching with genuine supercomputers, he quoted one calculation at 100,000 CPU hours. Lauren Williams is one of a small community of mathematicians who have been intrigued by the amplituhedron, her talk was a walk through a series of conjectures, some made by physicists, some by mathematicians, most with counterexamples found in the last few years.

Finally, Zvi Bern closed the conference with a talk on the frontier of amplitudes-based gravitational wave calculations, referring to it, probably to David Kosower’s annoyance, as 5PM. Some parts of this frontier have been calculated, but a few have proved hard going, bottlenecked by immensely challenging integrals, which go beyond the capabilities of publicly available codes. The new integrals have a variety of strange functions, including the elliptics and Calabi-Yaus I spent time on in my own career, as well as Heun integrals, which I imagine I will have to learn more about from the slides of Elliptics & beyond ’26, as I don’t remember people talking about them when I was in the field. New integration strategies have led to improvements in the public codes and seem to be making good progress, but Zvi highlighted that ideally they want to not need supercomputers at all, as they’ll need to go higher in loops to see effects from, for example, the deformability of neutron stars.

Amplitudes 2027 will be in Munich, with the summer school in Mainz under Stefan Weinzierl’s capable hands. I’m looking forward to seeing what the state of the art looks like then!

Peter Rohde Zinalrothorn (4,221m)

An album of GoPro headcam footage climbing Zinalrothorn (AD, 4,221m) in Switzerland.

Full album (49 videos): https://youtube.com/playlist?list=PLFMVEM4j3NZ0

July 09, 2026

Andrew JaffeWhat My Thirty-Year-Old Algorithm Taught an AI

As a scientist in the later stages of my career, the managerial and mentorship load has increased, leaving less time for math and programming. These technical activities were also why I wanted to be a scientist in the first place, and I regretted losing the time for this hands-on research. Formerly the core of my scientific work, these activities require sustained attention, hours at a time, a resource now in short supply.

Also like many scientists, I’ve watched the coming of artificial intelligence over the last few years with growing interest. The technical predecessors and underlying substructure of these large language models (LLMs) are neural networks, a computer technology that has already begun to revolutionise many scientific fields, including cosmology and astrophysics. I wrote about LLMs in my recent book, The Random Universe, but hadn’t really used them for my own work.

So I started using Anthropic’s Claude (no particular reason for this choice, except that some colleagues had had positive coding experiences with it), figured out how to wire it up at the command line, and started talking to it. (After, yes, paying a subscription fee.)

My first project with one of these newly-capable AIs was a minor reanalysis of our data from the Planck satellite, seeing how the cosmological inferences respond to small changes in the data, part of the work for a recent paper. I knew exactly what I needed to do, but it was a lot of plumbing: getting disparate bits of software written by other people to work together in a way different from their authors had intended. I figured I could do it in a few days of solid work.

Instead, I pointed Claude to the draft of the paper, along with publicly available Planck data and code repositories, and asked it to implement the paper’s algorithms with Planck’s data. A few hours later, there was code, alongside tables, figures, and lots of tests to make sure I could trust — and understand — the results. It wasn’t (we weren’t) just able to write and debug the code quickly, it was able to run it, again and again, making small tweaks to the inputs and the code itself, and to the figures it generated, now part of our recent paper.

Next was something more involved: colleagues and I have created a program called Almanac to analyse specific kinds of cosmological data. We wanted to apply Almanac to some new results, in a regime in which it hadn’t really been tested (a very small patch, around 1% of the total sphere of the sky). Almanac helps us measure a curve called the power spectrum, which I’ve written about before.

I pointed Claude to our code, our papers, and to the new data, and explained the problem. Even ensuring that the (poorly documented) data was in a form that our code could understand would have taken me a few hours, but Claude suggested and implemented a series of tests to ensure that everything was self-consistent.

Almanac is a Monte Carlo sampler: because we are trying to understand the probability distribution of matter in the Universe, using noisy and incomplete data, the answer to our questions can only be given as probability distributions. Almanac is essentially a very complicated random number generator, and you can do self-consistency checks to see whether it is producing random numbers with the right properties.

Almanac’s results failed these tests. Could we understand why? Could we fix it? Now, rather than just plumbing, I needed Claude to help me diagnose the problem. It took a while.

Was it a simple bug? We did a series of tests showing that Almanac does work, essentially perfectly, on simpler datasets covering much more of the sky. In fact, this gave me the opportunity to ask Claude to write some new software, based on a paper and related code that I first wrote, with Dick Bond and Lloyd Knox, about 30 years ago. This older algorithm (“BJK”, from our initials) answered the same statistical question as Almanac, using a very different technique. On large areas of sky, Almanac and this older algorithm got the same answer — the code works.

We went on a long rabbit-hole modifying the details of Almanac’s setup, making it more similar to other state of the art samplers. This change also didn’t solve our problem, although it seems to help on the margins. I made one suggestion that I thought would help, based on our long-ago experience with the BJK algorithm, bundling up some of the numbers we were trying to determine into “bands”. I don’t know if Claude would have come up with this idea on its own, and it took a while to get the details right. In fact, Claude would sometimes declare premature victory, admitting its mistakes only when I pointed them out.

It worked, eventually. After a lot of iterations, we transformed a problem unsolvable with the previous version of the code to one that was, well, easy.

But it only worked because my knowledge and experience — literally decades working on problems of this sort — meshed with Claude’s own “talents” — quick turnaround, patience, and encyclopaedic, if not always discriminating, knowledge of computing and of at least some aspects of the underlying science, statistics, and mathematics.

And it was fun! I thought that I liked programming, but I am very happy to have Claude do most of the grunt-work for me. The quick turnaround, and not having to sweat the details of writing and running re-writing and re-running program after program, was a delight.

In many way, working with Claude was like working with a junior colleague. But Claude is not a colleague, but a machine. And, as David Hogg has advocated, the point of doing astrophysics, a beautiful but useless field of science, is exactly the training and fulfilment of the people doing it. That has at least two implications. First, given how much my own experience was necessary to getting good results, that means we had better make sure that we are training humans, not just better LLMs. Second, no matter how delightful the interactions, they mustn’t replace training our students and collaborating with our colleagues.

(This post was written by me, not by Claude, though I did ask it to suggest a title — and this sentence.)

July 05, 2026

Scott Aaronson Happy 250th!

I’m at the New Jersey shore with family and friends, where we’ve spent this Fourth of July eating hot dogs, playing miniature golf, and wading into the ocean that my great-grandparents crossed to escape calamities they knew about and much greater calamities that they didn’t. Tonight we’ll see the fireworks, weather permitting.

And yes, on the crowded beach today you can find people sporting MAGA and “45-47” hats, and even a giant “Trump 2028” flag—a stark reminder of the millions who would redefine the meaning of our 250-year-old experiment to something dark and authoritarian, the opposite of what its founders intended. Of course, those forces find mirror images on the left end of the political spectrum, where one can find millions more who fully agree with MAGA about the failures of liberalism and the Enlightenment, differing only on the secondary question of which racist thugs should rule instead.

Despite everything, I don’t believe that both factions together constitute a majority. Even on the beach, the MAGA hats are vastly outnumbered by “250” banners and girls in stars-and-stripes bikinis, Americans who just want to celebrate.

Despite everything, I remain thoroughly American if I’ve ever been anything, and invested in the country’s future if I’ve ever been invested in anything.

JD Vance and his friends, who might rule the country after the predictable failure of “Trump 2028,” make a huge deal about “Heritage Americans.” Of course the point of such phrases is to exclude those like me, and recent immigrants, and even (incredibly) JD’s own wife. On reflection, though: could I, too, count as a Heritage American at this point? After all, my family has now been here for half the country’s history. My grandfather, who grew up in poverty, became a professional boxer in Philadelphia and Atlantic City during the Great Depression. He then joined the Army and ended up clearing German mines in North Africa and Italy in WWII. He was assigned to a company of Southerners, who had never met a Jew and were shocked that my grandfather didn’t have horns—but by the war’s end, my grandfather and the relatively few others in his company who remained alive had become best friends. My grandfather told me that he could understand the German POWs who they captured tolerably well, since German was similar enough to Yiddish, but who he could never understand was the British.

As for me, I grew up in the town of Washington Crossing, PA, maybe a mile’s walk from where this happened (and where it’s still reenacted every Christmas):

My earliest childhood hero (that I can remember) was Ben Franklin, whose institute in Philadelphia I visited often. I didn’t even recognize as unusual at the time how the founding of the country didn’t feel like a remote abstraction to me, but was all around me, as if it was yesterday.

That the founders of the United States created the model for all time of how to bend the arc of human history a little bit away from its usual horribleness, of how to overthrow a despotism without instituting an even worse despotism in its place, of how to found a new civilization on ideas and principles rather than raw power … is one of those things that seemed true to me as a child and that still seems true to me today.

May this greatest experiment continue for another 250 years. May it triumph against all those within and without who would see it destroyed.

July 01, 2026

Peter Rohde Aiguilles Crochues Traverse (2,840m)

An album of GoPro headcam footage from climbing the Aiguilles Crochues Traverse (PD, 2,840m) near Chamonix, France in 2022.

Full playlist (45 videos): https://youtube.com/playlist?list=PLM4i-DL0BZ0Q

June 30, 2026

Peter Rohde Frenchmans Cap (Sydney Route)

Footage from our climbing trip to Frenchmans Cap, Tasmania (Australian grade 17, 380m) in 2022.

Full playlist (47 videos): https://youtube.com/playlist?list=PLT0z6qQjCS3c

Peter Rohde Triglav, Slovenia (2,864m)

GoPro headcam footage from climbing Triglav (2,864m), highest mountain in Slovenia, via ferrata. Climbed in 2022.

Full playlist (40 videos): https://youtube.com/playlist?list=PLZgovD57Nsr4

June 29, 2026

John PreskillThe physicists of Florence

A scientist in Florence can’t avoid bumping into colleagues. 

When visiting the Renaissance’s birthplace last summer, I ran into a fellow physicist even on a Saturday morning. I was wandering around the Uffizi Gallery, a museum blessed with some of the greatest hits in western art. A familiar face arrested me on the first floor.

Another colleague cropped up outside the museum. (Some might classify him as an applied physicist or an engineer, but he exhibited a theoretical physicist’s overactive imagination.)

One colleague, I’d been looking forward to meeting for over four years. Jae Dong Noh is a professor of physics at the University of Seoul in South Korea. He’d conducted the first numerical tests (classical-computer simulations) of an idea I’d helped midwife, the non-Abelian eigenstate thermalization hypothesis (NAETH). An earlier blog post described this mouthful, which predicts how certain quantum many-particle systems thermalize, or experience the flow of time. These systems’ dynamics conserve properties, analogous to energy, that are incompatible: one can’t measure the properties simultaneously, as one can’t measure a quantum particle’s position and momentum simultaneously. Because incompatibility helps distinguish quantum from classical physics, such systems’ thermodynamics qualifies as particularly quantum.

Jae Dong modeled such a system and others numerically in a paper. I admired his computational techniques and his grasp of symmetries (for experts: how non-Abelian symmetries affect chaotic quantum systems’ energy-level statistics). My postdoc Aleks Lasek was planning a more thorough numerical test of the NAETH, so I reached out to Jae Dong, and a collaboration crystallized. 

Seoul operates thirteen hours ahead of Maryland, but we managed to Zoom because Jae Dong is a night owl and I’m an early bird.1 Zoom introduced me to a man perpetually dressed in a neat button-down shirt and sweater, silver overriding the black in his hair. The neatness extended to Jae Dong’s explanations: if Aleks and I didn’t understand one of his emails, he’d explain it quietly and calmly, untangling the confusion as though pulling a comb through wool.

The collaboration settled into a rhythm: I’d pose a question or propose a goal, Jae Dong would respond with an analytical calculation,2 I’d find holes in the calculation, Jae Dong would plug the holes, I’d re-check the argument’s logic, and we’d repeat the cycle. Had I been in Jae Dong’s shoes, I’d have swallowed the constant objections as I’ve swallowed grape-flavored cough medicine,3 but he always responded with equanimity—sometimes even good cheer—and a possible solution. Meanwhile, Aleks and then-undergraduate Jade LeSchack checked our analytical arguments numerically.

Florence flaunted a little steampunk during my visit.

So smoothly did the collaboration hum along that we coauthored two papers before ever meeting in person. One demonstrates numerically that two quantum many-body systems (for experts: nonintegrable Heisenberg models) obey the NAETH.4 In the other paper, we derive a symmetry relation from the NAETH. If the 17-syllable NAETH is a mouthful, the symmetry’s name is half a mouthful: a Kubo–Martin–Schwinger (KMS) relation. It’s important because (i) it enables us to calculate how rapidly a thermodynamic system responds to a stimulus, such as a weak magnetic field, and (ii) physicists go gaga over symmetries generally. 

The KMS relation constrains thermal states—essentially, systems that have temperatures. Your typical isolated many-particle quantum system looks thermal if you can observe just a small chunk of it at a time. Accordingly, Jae Dong and collaborators had proved that isolated many-particle quantum systems obey the KMS relation approximately. The larger the system, the more accurate the approximation. 

We extended his argument to systems whose dynamics conserve incompatible properties. Such an extension might sound simple, but its proof filled 24 pages of appendices. (For experts: Clebsch–Gordan coefficients are tricky blighters.) We discovered that, under certain conditions, incompatible conserved quantities can reduce the extent to which a quantum system obeys the KMS relation. Quantum incompatibility can augment deviations from conventional thermodynamics.

Italy’s architecture impressed me.

Jae Dong planned to present about our work at StatPhys, an international statistical-physics conference, which Florence was hosting in 2024. Throughout the two-and-a-half months before the conference, the KMS relation consumed our team. (For experts: Clebsch–Gordan coefficients are very tricky blighters.) I even hid in my hotel room, working and reworking our proofs, during another conference during that time. 

The toil paid off. We submitted our KMS manuscript for public scrutiny the day I flew to Florence—because not only Jae Dong would be representing our team at StatPhys. I was looking forward to meeting him there for the first time.

A corner of the hall where the StatPhys opening ceremony took place.

The StatPhys committee outdid itself. The opening ceremony unfolded in Florence’s Palazzo Vecchio, where members of the Medici dynasty once lived. Giorgio Parisi, who won a Nobel Prize for statistical physics in 2021, lectured at the ceremony.

Giorgio Parisi, with another history maker.

The meat of the conference took place in two other palaces, the Palazzo dei Congressi and the Palazzo degli Affari. In one of them, I met Jae Dong. Although we’d shown that quantum incompatibility can defy thermodynamic predictions, he met my expectations.

We discovered another thermodynamic phenomenon challenged by incompatible conserved quantities, so stay tuned for another paper and blog post. Some colleagues, one can’t avoid; others are worth engaging with again and again.

1 Aleks has confessed to night-owl habits, but physics motivates him to adapt. Some days, he’s emailed me results before even I’ve woken up. Who needs coffee when the thrill of discovery electrifies one minutes after one hops out of bed?

2 An exact calculation written out on paper, as opposed to a numerical, or approximate, calculation performed by a silicon-based classical computer.

3 Does anyone like the grape flavor? Why do companies bother producing it?

4 Rohit Patil and Marcos Rigol, too, have checked numerically that a system obeys the NAETH.

June 16, 2026

n-Category Café Octonions and the Standard Model (Part 14)

Paul Schwahn and I have come out with a new paper about octonions and the Standard Model:

It builds on things I’ve discussed here, but it goes further. Let me explain a bit.

A bit is just a binary alternative: 1 or 0, true or false. That’s how it works in classical logic. We could also have a ‘trit’, meaning 3 alternatives.

In quantum physics we instead have qubits and qutrits.

Qubits and qutrits are usually described using complex numbers. The algebra of observables of a qubit is the Jordan algebra 𝔥 2()\mathfrak{h}_2(\mathbb{C}), consisting of 2×22 \times 2 self-adjoint complex matrices. Similarly, the algebra of observables of an qutrit is the Jordan algebra 𝔥 3()\mathfrak{h}_3(\mathbb{C}), consisting of 3×33 \times 3 self-adjoint complex matrices.

We can also study systems with more than 3 alternative ways to be. They work the same way, using the Jordan algebras 𝔥 n()\mathfrak{h}_n(\mathbb{C}) with n>3.n \gt 3.

But we can also do quantum mechanics using other number systems! The options have been mapped out, and the largest allowed number system for this purpose is the algebra of octonions.

A weird thing is that Jordan algebras built using octonions can describe qutrits, but not quantum systems with more than 3 alternative ways to be. The algebra of observables of an octonionic qutrit is the so-called ‘exceptional’ Jordan algebra 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}), consisting of 3×33 \times 3 self-adjoint octonion matrices. What makes it exceptional is that 𝔥 n(𝕆)\mathfrak{h}_n(\mathbb{O}) is not a Jordan algebra when nn is bigger than 3.

So, there’s something special about octonionic qutrits — and it turns out that every symmetry in the gauge group of the Standard Model is a symmetry of an octonionic qutrit!

Not every symmetry of an octonionic qutrit is a symmetry of the Standard Model. But those that do have a simple description. They are those that restrict to give symmetries of an ordinary qutrit sitting inside the octonionic qutrit… and an ordinary qubit sitting inside that!

That sounds exciting, but also vague, so let me make it precise.

While lots of people say the gauge group of the Standard Model of particle physics is U(1)×SU(2)×SU(3)\text{U}(1) \times \text{SU}(2) \times \text{SU}(3), in fact a certain subgroup of this acts trivially on all known particles. If we mod out by that, we’re left with

S(U(2)×U(3)) = {xSU(5):x=(* * 0 0 0 * * 0 0 0 0 0 * * * 0 0 * * * 0 0 * * *)}. \begin{array}{ccl} \text{S}(\text{U}(2) \times \text{U}(3)) &= & \Big\{ x \in \text{SU}(5) : x = \left( \begin{array}{c c c c c} \ast & \ast & 0 & 0 & 0 \\ \ast & \ast & 0 & 0 & 0 \\ 0 & 0 & \ast & \ast & \ast \\ 0 & 0 & \ast & \ast & \ast \\ 0 & 0 & \ast & \ast & \ast \end{array} \right) \; \Big\}. \end{array}

and this is the group I’m talking about.

We proved two theorems describing this group in terms of the symmetries of an octonionic qutrit. The group of automorphisms of the exceptional Jordan algebra 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) is a 52-dimensional Lie group known affectionately as F 4\text{F}_4 — so that’s what I mean by the symmetries of an octonionic qutrit.

Here’s our main result:

Theorem 1. Suppose X,BX,B are Jordan subalgebras of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) such that

X𝔥 2(),B𝔥 3(),XB. X \cong \mathfrak{h}_2(\mathbb{C}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; X \subset B.

Then

Stab(X)Stab(B) 0S(U(2)×U(3)). \text{Stab}(X) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Here Stab(X)\text{Stab}(X) is the stabilizer of XX — that is, the subgroup of F 4\text{F}_4 consisting of elements that map XX to itself — while Stab(B) 0\text{Stab}(B)_0 is the identity component of the stabilizer of BB.

This ‘identity component’ business is rather sneaky, but it turns out that guys in Stab(B) 0\text{Stab}(B)_0 are symmetries of an ordinary qutrit that can be described as unitary operators on \mathbb{C}, while Stab(B)\text{Stab}(B) also contains those symmetries that are described by antiunitary operators. The CPT symmetry of the Standard Model is antiunitary, for example.

Theorem 1 emerged from a related result, which grew out of the work of Todorov and Dubois-Violette:

Theorem 2. Suppose A,BA,B are Jordan subalgebras of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) such that

A𝔥 2(𝕆),B𝔥 3(),AB𝔥 2(). A \cong \mathfrak{h}_2(\mathbb{O}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; A \cap B \cong \mathfrak{h}_2(\mathbb{C}).

Then

Stab(A)Stab(B) 0S(U(2)×U(3)). \text{Stab}(A) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Todorov and Dubois–Violette proved this for a certain standard choice of subalgebras AA and BB. Thus, the challenge in proving Theorem 2 was to show that every other choice can be mapped to this standard choice using the action of F 4\text{F}_4. This shows that the theorem is not an artifact of a specific choice, but rather a general fact.

How do we prove these results?

We start by constructing the octonion product from SU(3)\text{SU}(3)-invariant operations on \mathbb{C} and 3\mathbb{C}^3. We then use this description to reprove Todorov and Dubois–Violette’s special case of Theorem 2. Then we show that F 4\text{F}_4 acts transitively on the set of subalgebras of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) that are isomorphic to 𝔥 3()\mathfrak{h}_3(\mathbb{C}). We also show every Jordan subalgebra of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) isomorphic to 𝔥 2()\mathfrak{h}_2(\mathbb{C}) is contained in a unique Jordan subalgebra isomorphic to 𝔥 2(𝕆)\mathfrak{h}_2(\mathbb{O}). This lets us prove that F 4\text{F}_4 acts transitively on the set of pairs of Jordan subalgebra A,B𝔥 3(𝕆)A, B \subset \mathfrak{h}_3(\mathbb{O}) with A𝔥 2(𝕆)A \cong \mathfrak{h}_2(\mathbb{O}), B𝔥 3()B \cong \mathfrak{h}_3(\mathbb{C}) and AB𝔥 3()A \cap B \cong \mathfrak{h}_3(\mathbb{C}). Theorem 2 then follows from Todorov and Dubois-Violette’s special case. We conclude by using these results to prove Theorem 1.

However, if you want to get into the details of the physics, the interesting part is how the strong force gauge group SU(3)\text{SU}(3) and the electroweak S(U(1)×U(2))\text{S}(\text{U}(1) \times \text{U}(2)) show up from the relation between octonionic qutrits, complex qutrits and complex qubits. You’ll see that in the proof of Lemma 4.

And if you want to get into the details of the math, the main interesting thing here is the use of Jordan algebra technology like ‘Peirce decompositions’ and ‘Jordan frames’ to figure out what it must be like when you have a Jordan algebra 𝔥 2(𝕃)\mathfrak{h}_2(\mathbb{L}) or 𝔥 3(𝕃)\mathfrak{h}_3(\mathbb{L}) sitting inside 𝔥 3(𝕂)\mathfrak{h}_3(\mathbb{K}), where 𝕃\mathbb{L} is some normed division algebra contained in a bigger normed division algebra 𝕂\mathbb{K}.

What it all ‘really means’, if anything, is a question for later. It could be just a coincidence. Of course I hope not.

John BaezOctonions and the Standard Model

Paul Schwahn and I have come out with a new paper about octonions and the Standard Model:

The Standard Model gauge group from the exceptional Jordan algebra

It builds on things I’ve discussed here, but it goes further. Let me explain a bit.

A bit is just a binary alternative: 1 or 0, true or false. That’s how it works in classical logic. We could also have a ‘trit’, meaning 3 alternatives.

In quantum physics we instead have qubits and qutrits.

Qubits and qutrits are usually described using complex numbers. The algebra of observables of a qubit is the Jordan algebra \mathfrak{h}_2(\mathbb{C}), consisting of 2 \times 2 self-adjoint complex matrices. Similarly, the algebra of observables of an qutrit is the Jordan algebra \mathfrak{h}_3(\mathbb{C}), consisting of 3 \times 3 self-adjoint complex matrices.

We can also study systems with more than 3 alternative ways to be. They work the same way, using the Jordan algebras \mathfrak{h}_n(\mathbb{C}) with n > 3.

But we can also do quantum mechanics using other number systems! The options have been mapped out, and the largest allowed number system for this purpose is the algebra of octonions.

A weird thing is that Jordan algebras built using octonions can describe qutrits, but not quantum systems with more than 3 alternative ways to be. The algebra of observables of an octonionic qutrit is the so-called ‘exceptional’ Jordan algebra \mathfrak{h}_3(\mathbb{O}), consisting of 3 \times 3 self-adjoint octonion matrices. What makes it exceptional is that \mathfrak{h}_n(\mathbb{O}) is not a Jordan algebra when n is bigger than 3.

So, there’s something special about octonionic qutrits—and it turns out that every symmetry in the gauge group of the Standard Model is a symmetry of an octonionic qutrit!

Not every symmetry of an octonionic qutrit is a symmetry of the Standard Model. But those that do have a simple description. They are those that restrict to give symmetries of an ordinary qutrit sitting inside the octonionic qutrit… and an ordinary qubit sitting inside that!

That sounds exciting, but also vague, so let me make it precise.

While lots of people say the gauge group of the Standard Model of particle physics is \text{U}(1) \times \text{SU}(2) \times \text{SU}(3), in fact a certain subgroup of this acts trivially on all known particles. If we mod out by that, we’re left with a group called \text{S}(\text{U}(2) \times \text{U}(3)), which is

\Big\{ x \in \text{SU}(5) : x =   \left(   \begin{array}{c c c c c}  \ast & \ast & 0 & 0 & 0 \\  \ast & \ast & 0 & 0 & 0 \\  0 & 0 & \ast & \ast & \ast \\  0 & 0 & \ast & \ast & \ast \\  0 & 0 & \ast & \ast & \ast   \end{array}  \right) \; \Big\}.

and this is the group I’m talking about.

We proved two theorems describing this group in terms of the symmetries of an octonionic qutrit. The group of automorphisms of the exceptional Jordan algebra \mathfrak{h}_3(\mathbb{O}) is a 52-dimensional Lie group known affectionately as \text{F}_4—so that’s what I mean by the symmetries of an octonionic qutrit.

Here’s our main result:

Theorem 1. Suppose X,B are Jordan subalgebras of \mathfrak{h}_3(\mathbb{O}) such that

X \cong \mathfrak{h}_2(\mathbb{C}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; X \subset B.

Then

\text{Stab}(X) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Here \text{Stab}(X) is the stabilizer of X—that is, the subgroup of \text{F}_4 consisting of elements that map X to itself—while \text{Stab}(B)_0 is the identity component of the stabilizer of B.

This ‘identity component’ business is rather sneaky, but it turns out that guys in \text{Stab}(B)_0 are symmetries of an ordinary qutrit that can be described as unitary operators on \mathbb{C}, while \text{Stab}(B) also contains those symmetries that are described by antiunitary operators. The CPT symmetry of the Standard Model is antiunitary, for example.

Theorem 1 emerged from a related result, which grew out of the work of Todorov and Dubois-Violette:

Theorem 2. Suppose A,B are Jordan subalgebras of \mathfrak{h}_3(\mathbb{O}) such that

A \cong \mathfrak{h}_2(\mathbb{O}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; A \cap B \cong \mathfrak{h}_2(\mathbb{C}).

Then

\text{Stab}(A) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Todorov and Dubois–Violette proved this for a certain standard choice of subalgebras A and B. Thus, the challenge in proving Theorem 2 was to show that every other choice can be mapped to this standard choice using the action of \text{F}_4. This shows that the theorem is not an artifact of a specific choice, but rather a general fact.

How do we prove these results?

We start by constructing the octonion product from \text{SU}(3)-invariant operations on \mathbb{C} and \mathbb{C}^3. We then use this description to reprove Todorov and Dubois–Violette’s special case of Theorem 2. Then we show that \text{F}_4 acts transitively on the set of subalgebras of \mathfrak{h}_3(\mathbb{O}) that are isomorphic to \mathfrak{h}_3(\mathbb{C}). We also show every Jordan subalgebra of \mathfrak{h}_3(\mathbb{O}) isomorphic to \mathfrak{h}_2(\mathbb{C}) is contained in a unique Jordan subalgebra isomorphic to \mathfrak{h}_2(\mathbb{O}). This lets us prove that \text{F}_4 acts transitively on the set of pairs of Jordan subalgebra A, B \subset \mathfrak{h}_3(\mathbb{O}) with A \cong \mathfrak{h}_2(\mathbb{O}), B \cong \mathfrak{h}_3(\mathbb{C}) and A \cap B \cong \mathfrak{h}_3(\mathbb{C}). Theorem 2 then follows from Todorov and Dubois-Violette’s special case. We conclude by using these results to prove Theorem 1.

However, if you want to get into the details of the physics, the interesting part is how the strong force gauge group \text{SU}(3) and the electroweak \text{S}(\text{U}(1) \times \text{U}(2)) show up from the relation between octonionic qutrits, complex qutrits and complex qubits. You’ll see that in the proof of Lemma 4.

And if you want to get into the details of the math, the main interesting thing here is the use of Jordan algebra technology like ‘Peirce decompositions’ and ‘Jordan frames’ to figure out what it must be like when you have a Jordan algebra \mathfrak{h}_2(\mathbb{L}) or \mathfrak{h}_3(\mathbb{L}) sitting inside \mathfrak{h}_3(\mathbb{K}), where \mathbb{L} is some normed division algebra contained in a bigger normed division algebra \mathbb{K}.

What it all ‘really means’, if anything, is a question for later. It could be just a coincidence. Of course I hope not.

June 06, 2026

n-Category Café A New Blog

Readers may have noticed that I haven’t been very active here for a while. That isn’t because I haven’t felt the “blogging urge”, but because I felt that the things I want to blog about right now wouldn’t be very interesting to much of the n-Category Cafe audience: they’re mostly fairly technical details about implementing proof assistants (because that’s what I’m mostly working on right now).

Accordingly, I’ve started a new blog! It’s at https://gwaithimirdain.github.io/blog/. (Gwaith-i-Mírdain is the github organization for development of Narya, the experimental proof assistant for Higher Observational Type Theory – and now Multimodal Type Theory as well – that I’ve been spending most of my time on, and will primarily be blogging about.) And I already wrote three posts (mostly about implementing multimodal type theory, with several survey questions for the reader), so you can check it out right now and see whether it’s likely to be your cup of tea.

Never fear, I’ll still come back here when I have more category-theoretic things to write about.

June 02, 2026

John BaezInterview with Micah Zarin

I’m not completely happy with this interview with Micah Zarin. It was nothing he did, it was me. I forgot to say that current-day AI wastes a lot of energy, and companies hope to use it to lay off people, and oligarchs are using it to extract lots of money from everyone. While obvious, these things are tremendously important and I should have emphasized them.

I was distracted by Micah’s fear that AI would make a career in math pointless, which really surprised me. So instead of giving my general thoughts on AI, I focused on putting myself in his place and imagining what to do in that situation. I suggested doing math with the help of AI as a way to overcome his fear and go ahead doing math while keeping abreast of new developments. If AI overtakes humans in math in his lifetime, which is far from certain, this could be a way to keep productively participating in math throughout this process. But I warned him to be very critical of what LLMs say, to lessen the danger of getting caught up in the ‘AI vortex’ that is turning many people into crackpots.

Mathematics, in case you haven’t been paying attention, is different from some other subjects because it’s an area where LLMs have shown some truly impressive problem-solving ability: read the various mathematicians’ comments in Remarks on the disproof of the unit distance conjecture where they grapple with this. But nobody really knows where this is going. So far LLMs have not shown much ability to invent new theories of mathematics, so it would be jumping to conclusions to assume AI will soon overtake humans in that realm. It would also be jumping to conclusions to assume it won’t.

Whatever happens, the real danger is not that AI will become too good, but that it will become too evil—most likely because of the oligarchs, corporations and governments behind it. I wish I had emphasized that point, which is always on my mind.

I think I succeeded in making another point, which is that life will not become pointless simply because some other entity gets better than humans at something and knocks us off our throne. To think that the meaning of life resides in our superiority is a childish attitude.

May 28, 2026

John PreskillUnleashing the Advantage of Quantum AI

As experimental capabilities advance rapidly, the quantum computing community faces a critical elephant in the room: What will these quantum machines eventually be useful for? Will they deliver the promised broad societal impact, or will they remain highly specialized devices for exotic tasks known only to the experts?

The elephant in the room

Despite decades of effort, conclusive evidence of large quantum advantage in real-world applications remains confined to a few niche domains, such as simulating quantum materials and cryptanalysis. These problems are either inherently quantum to begin with, or they possess specialized mathematical structure that quantum algorithms can easily exploit. But it seems unlikely that such structures appear broadly in everyday life.

Indeed, most applications of modern computation hinge on the processing of massive, noisy classical data, generated at an unprecedented pace across society. That is the driving force behind the overwhelming success of machine learning and AI. Since the data originates from the macroscopic classical world, there is no obvious reason it should exhibit the delicate, specialized structures that quantum computers require. To playfully adapt Richard Feynman’s famous quote: We live in an effectively classical world, dammit, and maybe classical computers and AI already suffice for most of our problems. (For those unfamiliar, Feynman originally quipped: “Nature isn’t classical, dammit, and if you want to make a simulation of nature, you’d better make it quantum mechanical.”)

The central challenge

To truly unlock the power of a quantum computer, quantum algorithms typically need to access data in quantum superposition, processing many different samples simultaneously in different branches of the quantum multiverse. To use technical jargon, this is called querying a quantum oracle. But in reality, the classical data samples that we want to process are generated from everyday activities in a classical world, and we can only access them one at a time.

Think of the movie reviews you scroll through on a streaming platform. How would you read the plain-text reviews from a million different users all at once in a quantum superposition? This bottleneck—the challenge of efficiently accessing the classical world in quantum superposition—is known as the data loading problem. It has arguably been one of the main obstacles to achieving broadly applicable quantum advantage.

Sketching a quantum oracle

In this new work [1], we provide a solution to this seemingly impossible challenge. We develop a framework, called quantum oracle sketching, that enables us to access the classical world in quantum superposition in an optimal way. Importantly, it automatically handles the noise and correlations in the data, and natively supports flexible data structures like vectors and matrices that enable machine learning applications.

The core mechanism relies on processing data as a continuous stream. For each classical data sample we observe, we apply a carefully designed, small quantum rotation to our system. By sequentially accumulating these quantum rotations, we incrementally build up an accurate approximation of the target quantum oracle, which can then be used in any quantum algorithm for data processing. Because every data sample is processed once and immediately discarded, we completely eliminate the massive memory overhead typically required to store the dataset. The fundamental price to pay for assembling quantum queries from classical data lies in the sample complexity: our algorithm consumes a number of samples that scales quadratically with the number of quantum queries we need to make. We show that this rate is optimal and fundamentally arises from the quadratic relationship between quantum amplitudes and classical probabilities governed by the Born rule.

With the data successfully loaded into the quantum computer, the final challenge is to efficiently read out classical results. To address this, we develop an efficient measurement protocol called interferometric classical shadow. Combined with quantum oracle sketching, it allows us to circumvent the data loading and readout bottleneck to construct exponentially compact classical models from massive classical data with quantum technology.

Exponential quantum advantage in machine learning

Using this new approach, we are finally able to find exponential quantum advantage in processing classical data and machine learning. We rigorously prove that a small quantum computer can perform large-scale classification and dimensionality reduction on massive classical data by processing samples on the fly. In contrast, any classical machine achieving the same prediction performance requires exponentially larger size. When the classical machine does not have the required exponentially large memory size, it needs super-polynomially more samples and time relative to our protocol running on a quantum device. Remarkably, this illustrates that quantum technology enables us to construct compact and accurate classical models out of classical data, which is impossible with classical machines alone unless given exponentially larger memory.

The true scale of this exponential memory advantage is staggering. A quantum processor with 300 logical qubits can outperform a classical machine built from every atom in the observable universe. Of course, to actually see such a comical contrast, we would also need universe-scale datasets and processing time.

To contextualize these results in realistic scenarios, consider a large-scale scientific experiment, like a large particle collider. Each experimental run generates a colossal volume of data. With a quantum computer, we can keep squeezing all the data into this tiny quantum chip to perform downstream machine learning tasks such as classification and dimensionality reduction. But if we only have classical machines, we would need to build massive, energy-consuming data centers to store the raw data to match the performance. Without this massive memory overhead, classical machines simply couldn’t extract the same clear signals from a single run, forcing us to repeat the massive, expensive experiment many more times to compensate. To put this into perspective, the Large Hadron Collider (LHC) at CERN generates petabytes (millions of gigabytes) of data per hour, but the data storage bottlenecks force researchers to discard all but a tiny fraction—retaining perhaps only one in a hundred thousand events.

We validated these quantum advantages on real-world datasets, including movie review sentiment analysis and single-cell RNA sequencing. In these public datasets, we demonstrate four to six orders of magnitude (ten thousand to a million times) reduction in memory size with fewer than 60 logical qubits. Given the rapid advancements in high-rate quantum error correction codes and experimental techniques, quantum computers capable of demonstrating such applications are foreseeable in the near future. Crucially, the quantum advantage we propose likely carries a clearer positive impact for society and likely arrives sooner than the applications in cryptanalysis, where the current best estimate requires a thousand logical qubits.

Towards Quantum AI

Our results provide strong evidence that the utility of quantum computers extends far beyond specialized tasks, opening a path for quantum computers to be broadly useful in our everyday life. Rather than fearing that classical AI will “eat quantum computing’s lunch,” we now have rigorous evidence pointing towards a much more exciting prospect: quantum-enhanced AI overpowering classical AI.

Of course, there is still a long way to go towards the dream of quantum intelligence. Our current results establish the provable supremacy of quantum machines in foundational machine learning tasks, such as high-dimensional linear classification and dimensionality reduction. They do not yet imply immediate utility for modern generative AI such as large language models.

That said, our results give me a strong feeling that we are living in an age strikingly reminiscent of the traditional machine learning era—an age dominated by support vector machines and random forests; an age when we relied on rigorous statistical analysis because we lacked the computational resources for large-scale heuristic exploration; an age that ultimately heralded the birth of deep learning and the AI revolution. Today, quantum AI seems to sit at a similar historical position. I cannot wait to see what quantum AI will become once we are capable of unconstrained heuristic exploration on large-scale fault-tolerant quantum computers.

To accelerate this dawn of quantum AI, we invite physicists, computer scientists, developers, and machine learning practitioners to join our efforts and help us push the boundaries of what quantum AI can achieve. To bridge the gap between abstract quantum theory and hands-on machine learning practice, we are open-sourcing our core framework. Our numerical implementation of quantum oracle sketching is built in JAX, natively supporting GPU/TPU acceleration and automatic differentiation to integrate nicely with modern machine learning pipelines. Check out the code, run the simulations, and help us shape the future of quantum AI at github.com/haimengzhao/quantum-oracle-sketching!


References

[1]. Haimeng Zhao, Alexander Zlokapa, Hartmut Neven, Ryan Babbush, John Preskill, Jarrod R. McClean, and Hsin-Yuan Huang. Exponential quantum advantage in processing massive classical data, arXiv:2604.07639, 2026.

May 16, 2026

Jacques Distler Code

I’ve been playing around with Claude Code (Claude Opus 4.7 (1M context)) because, well, who hasn’t?

I am, so far, only moderately impressed by its abilities in physics. The most impressive bit so far was when, in response to a question about nilpotent orbits, it responded

I don’t know. I could waste your time by guessing, but …

I had never had an LLM tell me that it doesn’t know something, much less that it didn’t want to waste my time by making stuff up. So this was positively shocking to read.

Claude Code is, however, unfathomably good at generating code, so I set it the task of modernizing Instiki and Heterotic Beast, my forum software. Both are Rails applications and both have extensive test suites. So they use a software framework Claude is familiar with and have an objective standard for whether the changes made are correct.

When there is no test, however, things can go wildly off the rails (pun intended). For instance, consider the following snippet of Ruby code


def foo(text)
  ...
  con = text
  ...
  (now mutate con)
  ...
  con
end

If text is a frozen string, this will generate an error, as you can’t mutate a frozen string. Obviously, what you should do is write


def foo(text)
  ...
  con = text.dup
  ...
  con
end

which copies the caller’s string to a new unfrozen string which you can mutate to your heart’s content.

What did Claude do?


def foo(text)
  ...
  con = text.encode
  ...
  con
end

which also produces a new unfrozen string, transcoded from the caller’s encoding to Encoding.default_internal (which turns out to be nil). This is both (a) nondeterministic and (b) blows up spectacularly when text contains astral plane characters, like “𝔸”. I had to tell Claude not to do that, and to write some tests to check that astral plane characters are handled correctly.

Still …

I would set Claude the task of rewriting this blogging software, but alas I don’t have a test suite to compare with.

What I really should do, though, is find some physics I would trust it to work on.

May 09, 2026

Tim GowersA recent experience with ChatGPT 5.5 Pro

We are all having to keep revising upwards our assessments of the mathematical capabilities of large language models. I have just made a fairly large revision as a result of ChatGPT 5.5 Pro, to which I am fortunate to have been given access, producing a piece of PhD-level research in an hour or so, with no serious mathematical input from me.

The background is that, as has been widely reported, LLMs are now capable of solving research-level problems, and have managed to solve several of the Erdős problems listed on Thomas Bloom’s wonderful website. Initially it was possible to laugh this off: many of the “solutions” consisted in the LLM noticing that the problem had an answer sitting there in the literature already, or could be very easily deduced from known results. But little by little the laughter has become quieter. The message I am getting from what other mathematicians more involved in this enterprise have been saying is that LLMs have got to the point where if a problem has an easy argument that for one reason or another human mathematicians have missed (that reason sometimes, but not always, being that the problem has not received all that much attention), then there is a good chance that the LLMs will spot it. Conversely, for problems where one’s initial reaction is to be impressed that an LLM has come up with a clever argument, it often turns out on closer inspection that there are precedents for those arguments, so it is still just about possible to comfort oneself that LLMs are merely putting together existing knowledge rather than having truly original ideas. How much of a comfort that is I will not discuss here, other than to note that quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques.

I decided to try something a little bit different. At least in combinatorics, there are quite a lot of papers that investigate some relatively new combinatorial parameter that leads naturally to several questions. Because of the sheer number of questions one can ask, the authors of such papers will not necessarily have the time to spend a week or two thinking about each one, so there is a decent probability that at least some of them will not be all that hard. This makes such papers very valuable as sources of problems for mathematicians who are doing research for the first time and who will be hugely encouraged by solving a problem that was officially open. Or rather, it used to make them valuable in that way, but it looks as though the bar has just been raised. It is no longer enough that somebody asks a problem: it needs to be hard enough for an LLM not to be able to solve it.

In any case, a little over a week ago I decided to see how ChatGPT 5.5 Pro would fare with a selection of problems asked by Mel Nathanson in a paper entitled Diversity, Equity and Inclusion for Problems in Additive Number Theory. Nathanson has a remarkable record of being interested in problems and theorems that have later become extremely fashionable, which has led him to write a series of extremely well timed and therefore highly influential textbooks. In this paper, he argues for the interest of several other problems, some of which I will now briefly describe.

If A is a set of integers, then its sumset A+A is defined to be \{a+b:a,b\in A\}. For a positive integer h, the hfold sumset, denoted hA, is defined to be \{a_1+\dots+a_h: a_1,\dots,a_h\in A\}. Nathanson is interested in the possible sizes of hA given the size of A. To that end one can define a set \mathcal R(h,k) to be the set of all t such that there exists a set A with |A|=k and |hA|=t.

An obvious first question to ask is simply “What is \mathcal R(h,k)?” When h=2, the answer is the set of all integers between 2k-1 and \binom{k+1}2. It is an easy exercise to show that if |A|=k, then 2k-1\leq|A+A|\leq\binom{k+1}2, so this result is saying that all sizes in between can be realized. However, it is not true in general that hA can take every size between its minimum and maximum possibilities, and we do not currently have a complete description of \mathcal R(h,k).

Another natural question one can ask, and this is where ChatGPT came in, is how large a diameter you need if you want a set A with A and hA having prescribed sizes. (Of course, the size of hA must belong to \mathcal R(h,k).) Nathanson showed that for every t\in[2k-1,\binom{k+1}2] there is a subset A of \{0,1,2,\dots,2^k-1\} with |A|=k and |A+A|=t, and asked whether the bound 2^k-1 could be improved. ChatGPT 5.5 Pro thought for 17 minutes and 5 seconds before providing a construction that yielded a quadratic upper bound, which is clearly best possible. It wrote up its argument in a slightly rambling LLM-ish style, so I asked if it could write the argument up as a LaTeX file in the style of a typical mathematical preprint. After two minutes and 23 seconds it gave me that, after which I spent some time convincing myself that the argument was correct.

The basic idea behind both Nathanson’s argument and ChatGPT’s was that in order to obtain a set of a given size with a sumset of a given size, it is useful to build it out of a Sidon set, which means a set with sumset of maximal size (that is not quite the usual definition but it is the simplest to use in this discussion), and an arithmetic progression. Also, for a bit of fine tuning one can take an additional point near the arithmetic progression. Then if one plays around with the various parameters, one finds that one can obtain sets of all the sizes one wants. Nathanson doesn’t express his argument this way (it is Theorem 5 of this paper), instead giving an inductive argument, but I think, without having checked too carefully, that if one unravels his argument, one finds that effectively that is what he ends up with, and the Sidon set in question consists of powers of 2. ChatGPT obtained its improvement by simply using a more efficient Sidon set — it is well known that one can find Sidon sets of quadratic diameter. (One might ask why Nathanson didn’t do that in the first place: I think it is because the obvious idea of using a more efficient Sidon set becomes obvious only after one has redescribed his inductive construction. Is that what ChatGPT did? It is very hard to say.)

Next, I asked ChatGPT to see whether it could do the same for a closely related question, where instead of looking at the size of the sumset, one looks at the size of the restricted sumset, which is defined to be \{a+b:a,b\in A, a\ne b\}. Unsurprisingly, it was able to do that with no trouble at all. I got it to write both results up in a single note, to avoid a certain amount of duplication. If you are curious, you can see the note here.

I then asked what it could do for general h. I was much less optimistic that it would manage to do anything interesting, because the proof for h=2 makes fundamental use of the fact (due to Erdős and Szemerédi) that we know exactly which sizes we need to create. If we don’t know what the set \mathcal R(h,k) is, then it seems that we are forced to start with a hypothetical set A with |A|=k and |hA|=t and build out of it a set of small diameter with the same property. As it happens, I still don’t know how to get round that difficulty (I’m mentioning that just to demonstrate that my mathematical input was zero, and I didn’t even do anything clever with the prompts), but Nathanson mentioned in his paper a remarkable paper of Isaac Rajagopal, a student at MIT, who must have got round the difficulty somehow, because he had managed to prove an exponential dependence of \mathcal R(h,k) on k for each fixed h.

I’ll leave the previous paragraph there, but Isaac has subsequently explained to me that that isn’t really the difficulty. His argument gives a complete description of \mathcal R(h,k) when k is sufficiently large, and if one wants to prove a polynomial dependence for fixed h, then assuming that k is sufficiently large is clearly permitted. The real difficulty is that constructing the sets with given sumset sizes was significantly more complicated, and necessarily so because the degree of the polynomial grows with h, and one therefore needs more and more parameters to define the sets.

In any case, the task faced by ChatGPT was not to solve the problem from scratch, but to see whether it was possible to tighten up Isaac Rajagopal’s argument. Here’s what happened.

  1. After 16 minutes and 41 seconds, it came back with an argument that claimed to have improved the upper bound from exponential in k to exponential in k^\alpha for any \alpha>1/2.
  2. I asked it to write that in preprint form too, which took it a further 47 minutes and 39 seconds.
  3. That preprint would have been hard for me to read, as that would have meant carefully reading Rajagopal’s paper first, but I sent it to Nathanson, who forwarded it to Rajagopal, who said he thought it looked correct.
  4. Both ChatGPT and Rajagopal speculated a little on what might need to be done to push things further and get a polynomial bound, so I got greedy and asked ChatGPT to give that a go.
  5. After 13 minutes and 33 seconds it told me it felt optimistic about the existence of such an argument but there were a couple of technical statements that needed checking.
  6. I asked it to check them.
  7. After 9 minutes and 12 seconds it got back to me with the check having been done, so I asked for this too to be written in preprint form.
  8. After 31 minutes and 40 seconds the “preprint” was ready. Here it is.
  9. Isaac Rajagopal looked at it and declared it to be almost certainly correct. It was clear that he meant this not just at a line-by-line level but at the level of ideas.

Isaac made some very interesting remarks about the nature of what the additional ideas were that ChatGPT contributed. Since, as I have already said, my mathematical input was zero, I invited him to write a guest section to this post. Just before we get to that, I want to raise a question (that will undoubtedly have been raised by others as well), which is simple: what should we do with this kind of content? Had the result been produced by a human mathematician, it would definitely have been publishable, so I think it would be wrong to describe it as AI slop. On the other hand, it seems pointless even to think about putting it in a journal, since it can be made freely available, and nobody needs “credit” for it (except that Isaac deserves plenty of credit for creating the framework on which ChatGPT could build). I understand that arXiv has a policy against accepting AI-written content, which makes good sense to me. So maybe there should be a different repository where AI-produced results can live. But various decisions would need to be made about how it was organized. I myself think that one would probably want to have some kind of moderation process, so that results would be included only if a human mathematician was prepared to certify that they were correct — or, better still, that they had been formalized by a proof assistant — and perhaps also that they answered a question that had been asked in a human-written paper. On the other hand, I wouldn’t want a moderation process that created vast amounts of work (unless the work was itself done by AI, but there are obvious dangers in going down that route). Anyway, until these questions are answered, this result is available from the link above, and perhaps, now that LLMs are so good at literature search, that will be enough to make it findable by anyone who wants to know whether Nathanson’s problem has been solved.

Isaac’s evaluation of what ChatGPT achieved

With just a few prompts, ChatGPT was able to improve the upper bound on N(h,k) (which I will define very soon) from exponential in k to polynomial in k. While its first improvement of the bound, from exponential in k to exponential in k^{\frac{1}{2} + \varepsilon}, was a routine modification of my work, the improvement to polynomial in k is quite impressive. To do this, ChatGPT came up with an idea which is original and clever. It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour to find and prove, using similar methods to those in my own proof. My goal is to explain that idea, in a manner that will be digestible to my friends who are computer science majors as well as my math major friends.

The problem of bounding N(h,k) is closely related to a problem I worked on at the Duluth REU (Research Experience for Undergrads) program, of determining \mathcal{R}(h,k). In particular, \mathcal{R}(h,k) is the set of possible h-fold sumset sizes |hA|, where A can be chosen to be any set of k integers. N(h,k) is the minimal N such that we can achieve all of the values of \mathcal{R}(h,k) using k-element sets A \subset \{0,1,2,\ldots,N\}. I spent last summer explicitly characterizing the set \mathcal{R}(h,k) for large k, by constructing sets A such that |hA| achieves all sizes which I could not rule out as impossible. So, N(h,k) can be upper-bounded by optimizing my constructions.

I constructed these sets A by combining smaller component sets which are simpler to analyze. Some of these components are the geometric series

\displaystyle S = \{0,1,m,m^2,\ldots,m^{\ell-2}\} \quad \hbox{and} \quad T = \{1,m,m^2,\ldots,m^{\ell-1}\} \qquad (1)

for various values of 2 \leq m \leq h and 2 \leq \ell \leq k. Unfortunately, the elements of S and T are exponentially large in terms of k. So, I asked ChatGPT (through Tim) whether there exist sets of \ell elements which have similar sumset sizes to these geometric series, but contain only numbers of polynomial size in \ell: I had no idea if this was possible, or how to begin constructing such sets. ChatGPT came back with an answer, constructing sets G and H which behave like “half a geometric series squeezed into a polynomial interval,” which is counterintuitive. Before I discuss the construction of G and H, I will explain the important properties of the sumset sizes of S and T which they recreate.

For h > 0, a set A is called a B_h set if the only solutions to

\displaystyle x_1+\cdots+x_h = y_1+\cdots+y_h

with x_i,y_i in A are the “trivial” solutions, by which I mean that one side of the equation is a reordering of the other side. If A is a B_h set of size \ell, then elements of hA correspond exactly to choices of h elements of A, with repetition allowed. Using “stars and bars,” one can see that |hA| = \binom{h+\ell - 1}{h} and this is the maximum possible value of |hA| among sets of size \ell. So, another definition is that A is a B_h set if |hA| = \binom{h+|A| - 1}{h}. Sidon sets, which Tim discussed, are exactly B_2 sets.

To make things more concrete, let us assume that m = 4 in (1). Then, S is a B_3 set, but it is not a B_4 set because of the relations

\displaystyle 4^{a} + 4^a + 4^a + 4^a = 4^{a+1} + 0 + 0 + 0 \qquad (2)

for any choice of a in \{0,1,2,\ldots, \ell-3\}. In particular, \binom{\ell+3}{4} - |4S| = \ell-2, as these \ell-2 relations are the only ones preventing S from being a B_4 set. T lacks the relations in (2) because 0 is not in T. So, T is a B_4 set, but it is not a B_5 set because of the relations

\displaystyle 4^{a} + 4^a + 4^a + 4^a + 4^{b+1} = 4^{a+1} + 4^b + 4^b + 4^b + 4^b \qquad (3)

for any choices of a \neq b in \{0,1,2,\ldots, \ell-2\}. This gives \binom{\ell-1}{2} relations, and one can check that \binom{\ell+4}{5} - |5T| = \binom{\ell-1}{2}. To summarize, we have seen that

(a) S is a B_{m-1} set.

(b) \binom{m+\ell-1}{m} - |mS| = \ell -2 is a linear function of \ell.

(c) T is a B_{m} set.

(d) \binom{m+\ell}{m+1} - |(m+1)T| = \binom{\ell-1}{2} is a quadratic function of \ell.

    ChatGPT was able to find sets G and H of \ell elements which satisfy (a)-(d), but whose elements all have polynomial size in \ell. The construction of G and H uses h^2-dissociated sets, which are sets A where the only solutions to

    \displaystyle x_1+\cdots+x_s = y_1+\cdots+y_{s'} \qquad (4)

    with s,s' \leq h^2 and x_i,y_i in A are the “trivial” solutions, i.e. s = s' and one side of the equation is a reordering of the other side. For r > 0, it is possible to construct an h^2-dissociated set U = \{u_1,\ldots,u_r\} \subseteq \{0,1,2,\ldots,N\}, where N is approximately r^{h^2}, and in particular polynomial in r. Constructions of such a U using finite fields date back to Singer (1938) and Bose–Chowla (1963) and are described in Appendix 1. Define

    \displaystyle G = \{0, u_1,u_2,\ldots,u_r,mu_1,mu_2,\ldots, mu_r\}

    and

    \displaystyle H= \{u_1,u_2,\ldots,u_r,mu_1,mu_2,\ldots, mu_r\}. \qquad (5)

    In hindsight, I have good intuition for the construction of G and H. All of the relations in (2) and (3) are formed by combining one or two relations of the form 4x = y. There are approximately \ell relations of the form mx = y in S and T, and approximately \ell/2 such relations in G and H. There are few other low-order relations in S and T, and similarly in G and H because U is h^2-dissociated. So, G and H manage to contain half as many mx = y-relations as their geometric series counterparts, while also containing few low-order relations.

    We now see why (a)-(d) hold with S and T replaced by G and H, respectively. For concreteness, we assume that m = 4 and h>4, so U contains no nontrivial relations as in (4) with s,s' \leq 25 \leq h^2. Then, G is a B_3 set, but it is not a B_4 set because of the relations

    \displaystyle u_i + u_i + u_i + u_i = 4u_i + 0 + 0 + 0

    for any choice of i in \{1,2,\ldots, r\}. If we let \ell = |G| = 2r+1, we can check that \binom{\ell + 3}{4} - |4G| = r = \frac{\ell-1}{2} is linear in \ell. In particular, (a) and (b) hold with S replaced by G, and the linear function \ell-2 replaced by \frac{\ell-1}{2}. We can also see that H is a B_4 set, but it is not a B_5 set because of the relations

    \displaystyle u_i + u_i + u_i + u_i + 4u_j = 4u_i + u_j + u_j +u_j + u_j

    for any i\neq j in \{1,2,\ldots, r\}. If we let \ell = |H| = 2r, we can check that \binom{\ell + 4}{5} - |5H| = \binom{r}{2} = \binom{\ell/2}{2} is quadratic in \ell. In a similar manner, (c) and (d) hold with T replaced by H, and the quadratic function \binom{\ell-1}{2} replaced by \binom{\ell/2}{2}.

    Even though I can motivate it in retrospect, ChatGPT’s idea to use h^2-dissociated sets to control relations of order at most h feels quite ingenious. As far as I can tell, this idea is completely original.

    ChatGPT’s proof that its construction produces the desired values of |hA| is very similar to my proof that the sets A which I construct achieve all possible values of |hA|, after replacing S and T by G and H, respectively. Properties (a)-(d) capture many of the important properties of S and T (or G and H) which are used in this proof. The final constructions involve combining the sets G and H (or S and T in my paper) for each value of m between 2 and h with another set which is the union of an arithmetic progression and a point. Intuitively, G and H (or S and T) have large sumsets, while arithmetic progressions have small sumsets, so it is plausible that one could get sets which achieve all the medium-sized sumsets by combining them. However, the proof of this is quite involved, and it occupies Section 4 of my paper and the entirety of the ChatGPT preprint. In Appendix 2, I work out the details of the ChatGPT construction to show that for k sufficiently large,

    \displaystyle N(h,k) \leq O\left(k^{10h^3}\right).

    For comparison, it is easy to see that N(h,k) is at least on the order of k^{h}, and it is unknown what the real value is. In Appendix 3, I give details of the correspondence between my paper and the ChatGPT preprint, which will be helpful for those who want to read either.

    Finally, I want to express my deep gratitude to Tim for allowing me to contribute to this blog. I am still stunned by the coincidence that the problem he chose to put into ChatGPT 5.5 Pro led him to my paper on the arXiv.

    Tim on what this means for mathematical research

    I would judge the level of the result that ChatGPT found in under two hours to be that of a perfectly reasonable chapter in a combinatorics PhD. It wouldn’t be considered an amazing result, since it leant very heavily on Isaac’s ideas, but it was definitely a non-trivial extension of those ideas, and for a PhD student to find that extension it would be necessary to invest quite a bit of time digesting Isaac’s paper, looking for places where it might not be optimal, familiarizing oneself with various algebraic techniques that he used, and so on.

    It seems to me that training beginning PhD students to do research, which has always been hard (unless one is lucky enough, as I have often been, to have a student who just seems to get it and therefore doesn’t need in any sense to be trained), has just got harder, since one obvious way to help somebody get started is to give them a problem that looks as though it might be a relatively gentle one. If LLMs are at the point where they can solve “gentle problems”, then that is no longer an option. The lower bound for contributing to mathematics will now be to prove something that LLMs can’t prove, rather than simply to prove something that nobody has proved up to now and that at least somebody finds interesting.

    I would qualify that statement in two ways though. First, there is the obvious point that a beginning PhD student has the option of using LLMs. So the task is potentially easier than proving something that LLMs can’t prove: it is proving something in collaboration with LLMs that LLMs cannot manage on their own. I have done quite a lot of such collaboration recently and found that LLMs have made useful contributions without (yet) having game-changing ideas.

    A second point is that I don’t know how much of what I have said generalizes to other areas of mathematics. Combinatorics tends to be quite focused on problems: you start with a question and you reason back from the question or if you reason forwards you do so very much with the question in mind. In other areas there can be much more of an emphasis on forwards reasoning: you start with a circle of ideas and see where it leads. To do it successfully, you need to have some way of discriminating between interesting observations and uninteresting ones, and it isn’t obvious to me what LLMs would be like at that.

    Of course, everything I am saying concerns LLMs as they are right now. But they are developing so fast that it seems almost certain that my comments will go out of date in a matter of months. It is also almost certain that these developments will have a profoundly disruptive effect on how we go about mathematical research, and especially on how we introduce newcomers to it. Somebody starting a PhD next academic year will be finishing it in 2029 at the earliest, and my guess is that by then what it means to undertake research in mathematics will have changed out of all recognition.

    I sometimes get emails from people who are interested in doing mathematical research but are not sure whether that makes sense any more as an aspiration. I have a view on that question, but it may very well change in response to further developments. That view is that there is still a great deal of value in struggling with a mathematics problem, but that the era where you could enjoy the thrill of having your name forever associated with a particular theorem or definition may well be close to its end. So if your aim in doing mathematics is to achieve some kind of immortality, so to speak, then you should understand that that won’t necessarily be possible for much longer — not just for you, but for anybody. Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would.

    So what is the point of struggling with a difficult mathematics problem? One answer is that it can be very satisfying to solve a problem even if the answer is already known, but I don’t think that is a sufficient reason to spend several years of your life on this peculiar activity. A better answer is that by solving hard problems you get an insight into the problem-solving process itself, at least in your area of expertise, in a way that you simply don’t if all you do is read other people’s solutions. One consequence of this is that people who have themselves solved difficult problems are likely to be significantly better at using solving problems with the help of AI, just as very good coders are better at vibe coding than not such good coders, or people who have a solid grasp of how to do basic arithmetic are likely to be more skilled at using calculators (and especially at noticing when an answer feels off). Mathematics is a highly transferable skill, and that applies to research-level mathematics as well. By doing research in mathematics, you may not get the same rewards as your equivalents a generation ago, but there is a good chance that you will be equipping yourself very well for the world we are about to experience.

    Appendix 1 (Isaac)

    We will construct an h-dissociated set U = \{u_1,\ldots,u_r\} \subseteq \{0,1,2,\ldots,N\}, where N is approximately r^{h}. This construction is a very minor modification of Bose–Chowla (1963)’s construction of a B_h set, which I learned about from this paper. For whatever reason, the GPT preprint (Lemma 3.1) uses a different, less efficient construction using moment curves.

    Let p > r be a prime, let N = p^{h+1}-2, let K be the finite field with p^{h+1} elements and fix a generator \theta of K^\times, so that K^\times is equal to \{\theta^0,\theta^1,\ldots, \theta^N\}. Define a set of p elements

    \displaystyle U = \{a \in \{0,1,2,\ldots,N\}: \theta^a - \theta \in \mathbb{F}_p\}.

    Then, each element a \in U corresponds to a unique value of \tilde{a} \in \mathbb{F}_p, by taking \tilde{a} = \theta^a - \theta. Now an additive relation of the form in (4) with s,s' \leq h can be reframed by taking powers of \theta as

    \displaystyle (\theta + \tilde{x_1})(\theta + \tilde{x_2})\cdots (\theta + \tilde{x_s}) = (\theta + \tilde{y_1})(\theta + \tilde{y_2})\cdots (\theta + \tilde{y_{s'}}). \qquad (6)

    As K is a degree-h+1 extension of \mathbb{F}\sb{p} and \theta is a generator of K as an \mathbb{F}\sb{p}-extension, this means that \theta does not satisfy any nonzero polynomials in \mathbb{F}\sb{p}[x] of degree \leq h. So, both sides of (6) are identical as polynomials in \mathbb{F}_{p}[\theta] and thus the additive relation in (4) is trivial. So, U is h-dissociated, and of course one can prune a few elements to reduce U to size r.

    Appendix 2 (Isaac)

    Fix constants \alpha,\beta,\gamma such that 0.5 < \beta\gamma < \beta < \alpha < 1 (in my paper I arbitrarily chose (\alpha,\beta,\gamma) = (0.9,0.8,0.7)). Let the two sets in (5) be called G_{m,r} and H_{m,r}. Let [a,b] denote the set of integers x satisfying a \leq x \leq b. Similarly to my paper, the constructions of A such that hA achieves the desired sizes will combine sets of the following four types:

    • B_{j,b} := [0,b-2] \cup \{b-2+j\} with choices of b \in [3, k-k^\gamma] and j \in [1,hb].
    • G_{m,r_m} for each value of m \in [3, h], with choices of r_m \in [0, (k-b)^\alpha].
    • H_{m,u_m} for each value of m \in [2,h-1], with choices of u_m \in [0, (k-b)^\beta].
    • A B_h set of the correct size so that |A| = k.

    One reason that this construction needs to be complicated is that we need to create at least \Omega(k^h) many sets. To do this, we vary 2h-4 parameters r_m and u_m in the domain [0,k^\alpha] and 2 parameters b and j in the domain [1,hk]. We can choose \alpha to be slightly bigger than 1/2, and then the above construction gives us O(k^{\alpha(2h-4)+ 2})=O(k^{h + \delta}) different sets where \delta >0 can be made arbitrarily small. So, if we were to remove any of the above parameters from the construction, and not change the others, this construction would no longer create \Omega(k^h) many sets. In comparison, Nathanson’s construction when h=2 only needs to create \Omega(k^2) sets. He does this by combining a Sidon set, an arithmetic progression, and one extra value, and varying the size of the arithmetic progression and the extra value in ranges of size O(k).

    We want to combine q = 2h-2 sets A_1,\ldots,A_q, which are given by B_{j,b}, G_{m,r_m} for the h-2 values of m \in [3,h], H_{m,u_m} for the h-2 values of m \in [2,h-1], and a B_h set. By Appendix 1, for all r \leq k, there exists a h^2-dissociated set {u_{1},\ldots,u_{r}} of diameter M \leq r^{2h^2} \leq k^{2h^2}. By the constructions of G_{m,r_m} and H_{m,u_m}, we can take each A_i \subseteq [0,M], where M \leq hk^{2h^2}. Let \mathbb{Z}^{2q} have basis vectors e_1,\ldots,e_{2q}. To combine A_1,\ldots,A_q, we can define A \subseteq \mathbb{Z}^{2q} as

    \displaystyle A = \bigcup_{i=1}^q (A_i e_i + e_{q+i}) \subseteq \{0,1,2,\ldots,M\}^{2q} \subseteq \mathbb{Z}^{2q}.

    Similarly to my Lemma 4.9, this construction ensures that the generating function product \mathcal{F}_{A}(z) = \prod_{i=1}^q \mathcal{F}_{A_i}(z) holds, which is the identity that both my paper and the GPT preprint use (see either paper for a definition of these generating functions). By (the standard) Lemma 2.3 of the GPT preprint, A is Freiman-isomorphic of order h to a subset of [0,2qM(2hM)^{2q-1}]. Therefore, for k sufficiently large (the whole construction relies on this for the same reasons as in my paper),

    \displaystyle N(h,k) \leq 2qM(2hM)^{2q-1} \leq 2\left(2h^2k^{2h^2}\right)^{2(2h-2)} \leq k^{10 h^3}.

    Appendix 3 (Isaac)

    In Section 4.2 of my paper, I use a different, simpler construction to construct sets A achieving the values in \mathcal{R}(h,k) which have |hA| < \varepsilon k^h, for some small \varepsilon. These sets A are subsets of {0,1,2,\ldots,k^h}, meaning that all elements have polynomial size in k. This is observed in Section 5 of the GPT preprint.

    Section 4.3 of my paper carries out the construction which combines many components including S and T. This corresponds to Sections 2, 3, 4, and 6 of the GPT preprint. This section has a lot of moving parts; I give an outline in Section 4.3.1.

    In Section 4.3.2, I describe how the different components will be combined, using a construction which I call the disjoint union, and introduce generating functions \mathcal{F}_A(z) as a bookkeeping tool to keep track of the sumset sizes of a set A. This corresponds to Section 2 and Section 4 of the GPT preprint.

    In Section 4.3.3, I compute the generating function of each of the component sets, including \mathcal{F}_S(z) (Lemma 4.15) and \mathcal{F}_T(z) (Lemma 4.17). This corresponds to Section 3 and Section 6.1 of the GPT preprint. In particular, \mathcal{F}_{G}(z) is computed in Lemma 3.3 and \mathcal{F}_{H}(z) is computed in Lemma 3.4. Once these generating functions have been computed, the remainder of the proof is almost identical in my paper and in the GPT preprint.

    In Section 4.3.4, I put all the pieces together to show that as we range over the sets A which I have constructed, the values of |hA| will assume all of the elements of {\lceil\varepsilon k^h\rceil, \lceil\varepsilon k^h\rceil+1,\ldots ,\binom{h+k-1}{h} }. The key idea is to show that the set of all values of |hA| forms an interval, and contains numbers both smaller than \varepsilon k^h and equal to \binom{h+k-1}{h}.

May 02, 2026

n-Category Café Quantum Mechanics of the Inverse Cube Force Law

In the last episode of my column in Notices of the American Mathematical Society, we looked at a particle moving in an attractive central force whose strength is proportional to the inverse cube of the distance from the origin. Among other things, we saw that a particle moving in such a force can spiral in to the origin in a finite time. But that was classical mechanics. What about quantum mechanics?

Here things get more tricky. The uncertainty principle tends to prevent the particle from falling in to the origin. But when the attractive force is strong enough, the particle can still fall in. We can make up a theory where the particle shoots back out, but there are choices involved: we need to say how the particle changes phase when shoots back out. So there is not just a single theory, but many!

Why does the particle come back out? There are theories where it does not. In these theories, at least those studied so far, time evolution is nonunitary: that is, the probability of finding the particle somewhere or other does not stay equal to 11, because the particle simply disappears when it hits the origin. Here we focus on theories where time evolution is unitary and the particle comes back out. Many people have written about these, running into ‘paradoxes’ when they weren’t careful enough. Only rather recently have things been straightened out.

Let us dig into the details. In quantum mechanics, the Hilbert space of states of a particle in 3\mathbb{R}^3 is L 2( 3)L^2(\mathbb{R}^3). In a central force whose strength is proportional to 1/r 31/r^3, such a particle has a Hamiltonian of this form:

H= 2+cr 2 H = -\nabla^2 + c r^{-2}

The first term describes the particle’s kinetic energy, while the second describes its potential energy: remember, taking the gradient of an inverse square potential gives an inverse cube force. I have set some constants to 11 to remove irrelevant clutter, but we need the constant cc to say how strong the force is. When c<0c \, &lt; \, 0, the force is attractive.

In this game, analysis is paramount. We should interpret HH as a densely defined linear operator on L 2( 3)L^2(\mathbb{R}^3). For this, we choose a dense linear subspace DL 2( 3)D \subset L^2(\mathbb{R}^3) and treat HH as a linear map from DD to L 2( 3)L^2(\mathbb{R}^3). Different choices of DD correspond to different physical assumptions: for example, assumptions about what happens when the particle falls into the origin.

To get unitary time evolution in quantum mechanics, we need the Hamiltonian to be self-adjoint. But adjoints of densely defined operators are tricky. Let us briefly recall how they work. Given a Hilbert space \mathcal{H} and a linear operator AA from a dense linear subspace D(A)D(A) \subseteq \mathcal{H} to \mathcal{H}, we define D(A *)D(A^*) to be the set of all ψ\psi \in \mathcal{H} for which there exist ψ\psi' \in \mathcal{H} such that

ψ,ϕ=ψ,Aϕ for all ϕD(A). \langle \psi' , \phi \rangle = \langle \psi, A \phi \rangle \; \text{ for all } \; \phi \in D(A).

If such a vector ψ\psi' exists, it is unique, and it depends linearly on ψ\psi. Thus, for ψD(A*)\psi \in D(A\ast) we define A*ψA\ast \psi to be the vector ψ\psi' with the above property. The adjoint of AA is then the linear operator A*:D(A*) A\ast \colon D(A\ast) \to \mathcal{H}. We say AA is self-adjoint if A=A*A = A\ast. We say that AA is essentially self-adjoint if it has a unique extension to a self-adjoint operator. If it does, this extension must be A*A\ast.

All this raises the question of whether the Hamiltonian HH for the inverse cube force law can be made self-adjoint with a suitable choice of domain. It turns out we can always do it, but sometimes in more than one way. There are three regimes:

  • c34c \ge \tfrac{3}{4}. In this case we can start with the domain C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}) consisting of smooth functions that are compactly supported on 3\mathbb{R}^3 minus the origin. The operator HH is unambiguously defined on this domain, and it is essentially self-adjoint.

  • 14c<34-\tfrac{1}{4} \le c \, &lt; \, \tfrac{3}{4}. In this case HH is still well-defined on the domain C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}), but it is not essentially self-adjoint. In fact, it admits more than one self-adjoint extension! However, HH is bounded below: there is a constant E 0E_0 such that ψ,HψE 0ψ,ψ \langle \psi, H \psi \rangle \ge E_0 \langle \psi, \psi \rangle for all ψC 0 ( 3{0})\psi \in C_0^\infty(\mathbb{R}^3 - \{0\}). Physically, this means that the particle’s energy is bounded below by E 0E_0. Mathematically, this implies that HH has a canonical choice of self-adjoint extension called the ‘Friedrichs extension’, with the smallest possible domain. But there is another canonical choice, the ‘Krein extension’, with the largest possible domain.

  • c<14c \, &lt; \, -\tfrac{1}{4}. In this case HH is well-defined on the domain C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}), and it has more than one self-adjoint extension, but it is not bounded below.

These strange results demand explanation. For example, what is special about c=14c =-\tfrac{1}{4}? In classical mechanics, the energy of a particle in the inverse cube force ceases to be bounded below as soon as c<0c \, &lt; \,0. Quantum mechanics is different. To get a lot of negative potential energy, the particle’s wavefunction must be peaked near the origin, but that gives it kinetic energy. The tradeoff is captured by Hardy’s inequality. This says that for any ψC 0 ( 3)\psi\in C_0^\infty(\mathbb{R}^3) we have

ψ,( 214r 2)ψ0. \langle \psi, (-\nabla^2 - \tfrac{1}{4} r^{-2}) \psi \rangle \ge 0 .

This is why HH is bounded below when c14c \ge -\tfrac{1}{4}.

On the other hand, the constant 14\tfrac{1}{4} in Hardy’s inequality cannot be improved, so if c<14c \, &lt; \, \tfrac{1}{4} we can find ψ\psi with ψ,Hψ<0 \langle \psi, H \psi \rangle \, &lt; \, 0. Then we can use a remarkable property of the r 2r^{-2} potential to show that HH is not bounded below. Namely, HH has a kind of symmetry under dilations. You can guess this by noting that both the Laplacian and r 2r^{-2} have units of 1/length 2{}^2. Indeed, if you take any smooth function ψ\psi, dilate it by a factor of α\alpha, and then apply HH, you get α 2\alpha^{-2} times what you get if you do these operations in the other order. This implies that if

ψ,Hψ=Eψ,ψ, \langle \psi, H \psi \rangle = E \langle \psi, \psi \rangle ,

we can dilate ψ\psi and get a function obeying the same equation with EE replaced by α 2E\alpha^{-2} E. Thus, as soon as EE can be negative, it can be made arbitrarily large and negative by choosing α\alpha to be very small. Thus HH is not bounded below.

Next, what is special about c=34c = \tfrac{3}{4}? This is more subtle. For any value of cc \in \mathbb{R} we can find spherically symmetric solutions of ( 2+cr 2)ψ=iψ ( -\nabla^2 + c r^{-2})\psi = i \psi on 3{0}\mathbb{R}^3 - \{0\} that are nonzero and smooth. When c<34c \, &lt; \, \tfrac{3}{4}, and only in this case, some of these solutions ψ\psi lie in L 2( 3)L^2(\mathbb{R}^3). This dooms the chance of HH being essentially self-adjoint, because it implies H*ψ=iψH\ast \psi = i \psi. If HH were essentially self-adjoint H*H\ast would be self-adjoint, and it is easy to see that a self-adjoint operator cannot have ii as an eigenvalue.

When c<34c \, &lt; \, \frac{3}{4} the operator HH has more than one self-adjoint extension from C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}) to some larger domain. To classify these we can use separation of variables, writing 2\nabla^2 as a sum of a radial part and an angular part, assuming the angular dependence of ψ\psi is given by a spherical harmonic Y mY_{\ell m}, and doing a change of variables u=ψ/ru = \psi/r to reduce HH to the ordinary differential operator

d 2dr 2+(c+(+1))1r 2 - \frac{d^2}{d r^2} + \left(c + \ell(\ell+1)\right) \frac{1}{r^2}

on the half-line (0,)(0,\infty). We can completely classify self-adjoint extensions of this differential operators from C 0 (0,)C_0^\infty(0,\infty) to larger domains; the answer depends on cc and \ell. A choice of self-adjoint extension is a choice of boundary conditions at r=0r = 0, and this says how the phase of an incoming wave changes as it reflects off the origin and bounces back. Finally, we can assemble the results for different spherical harmonics to classify self-adjoint extensions of HH.

There exist many self-adjoint extensions of HH that respect the rotational symmetry of the inverse cube force law, but for c<14c \, &lt; \, -\tfrac{1}{4} the extension must break the dilation symmetry discussed above. This is what physicists call an ‘anomaly’: a symmetry of a classical system that fails to be a symmetry of the corresponding quantum system. Intriguingly, for some even lower values of cc one can choose a self-adjoint extension that is symmetrical under a discrete subgroup of dilations. Determining precisely which values these are seems to be an open problem.

To explore this topic thoroughly, I recommend first this:

then this:

and finally this:

The first is an excellent overview of problems associated to singular potentials, including the inverse cube force. The second delves into self-adjoint extensions of the ordinary differential operators mentioned above, and the third works them out with exquisite thoroughness.

April 16, 2026

Clifford JohnsonComputing Correlators

[A more technical post follows]

My most recent paper, out on the arXiv today, is very exciting to me because it seems to be a genuinely new way of computing some important quantities and it is devilishly simple. So simple that I worried for months that it is all super-obvious to everyone. But another voice within me said to myself: Well if it is so obvious, why has nobody published it? Another (paranoid) voice within said: Maybe someone has published this method, and I just can't find it in the literature...

Well, I decided that the best way to find out for sure is to put it on the arXiv and within a short time someone will email to say that I missed their important work. So, while I wait for that email (as I start writing it's only been 30 minutes since it has been "out there", so there's time), let me say a few things about why I like the many results in the paper.

I was already pleased enough with the core part of the paper that I was going to write a swift four-pager about it back in February. The core point being that I figured out how to build on work I'd done in a paper back in 2024 (expanded on with followup work I did with Wasif Ahmed and Krishan Saraswat, a student and postoc). Back in 2024, I found (here) a really nice way (almost miraculous in how it worked) of writing all the corrections to the spectral density of a class of models in terms of one function [latex]u_0[/latex] and its derivatives. It was obtainable from one simple ordinary differential equation (ODE) called the Gel'fand-Dikii equation, which takes in the function [latex]u_0(x)[/latex] as input. The ODE is for a special quantity called the diagonal resolvent [latex]{\widehat R}(x,E)[/latex]. You integrate that quantity [latex]\widehat R(x,E)[/latex] with respect to [latex]x[/latex] and you're more or less home. In general, it is a messy quantity that does not integrate to anything nice. But just when the function [latex]u(x)[/latex] obeys the "string equation" it is supposed to (as dictated by the governing model's physics), then [latex]{\widehat R}(x,E)[/latex] is a total derivative (a seeming miracle-see later), and the corrections it gives to the density become of just the right form!

Those corrections can be called [latex]W_{g,1}(E)[/latex] where the [latex]g[/latex] is the order in perturbation theory. [latex]g=0[/latex] is leading order, [latex]g=1[/latex] is the torus, [latex]g=2[/latex] the double torus, etc. Indeed [latex]g[/latex] is the number of handles or "genus" of an associated Riemann surface. The one subscript on the other hand, corresponds to the one energy entry available when just discussing the density [latex]\rho(E)[/latex]. All the [latex]W_{g,1}[/latex] end up being written nicely in terms of a function [latex]u_0(x)[/latex] and its derivatives, evaluated at a special point.

An already nice feature (among many) of the construction was that this one ODE, recursively solved, gave rise to the [latex]W_{g,1}[/latex] of many different problems across a range, including certain random matrix models, gravity problems, intersection theory and topology, and so on. All you need to do is change the function [latex]u_0(x)[/latex]. Moreover, for this (wide) class of problems, you can compute the desired results faster and with way less machninery than other methods, such as topological recursion, which was an interesting observation. This includes very famous problems like the Weil-Petersson volumes (of the compactified moduli space [latex]\overline{\cal M}_{g,1}[/latex] of Riemann surfaces with genus [latex]g[/latex] and [latex]n=1[/latex] boundaries) and generalisations. Another nice feature is that you also get non-perturbative data beyond the genus expansion, an aspect I explored recently (in this paper) with student Joao Rodrigues, and expert in resurgence techniques.

The core breakthrough of the new paper is this: For some time, I've wondered how to compute correlators for more energies (amounting to multi-point correlators of [latex]\rho[/latex]) in this same way: [...] Click to continue reading this post

The post Computing Correlators appeared first on Asymptotia.