Overview
Bhargava and Shankar proved that the average size of the -Selmer group of elliptic curves over , ordered by naive height, is exactly (Bhargava & Shankar, 2015). Consequently the average -Selmer rank, and hence the average Mordell—Weil rank, is at most . This was the first unconditional proof that the average rank of all elliptic curves over is bounded.
The proof is not primarily an analytic estimate for -functions. It is an arithmetic orbit-counting theorem. Nonidentity elements of are represented by locally soluble integral binary quartic forms with prescribed invariants. The paper counts such forms by geometry of numbers, proves the necessary congruence-uniformity estimates, and evaluates the local masses. The final identity
is the Tamagawa number computation for in disguise.
Elliptic Curves Ordered by Height
Every elliptic curve over is isomorphic to a unique curve
with satisfying the minimality condition
for every prime (Bhargava & Shankar, 2015). The discriminant is
and nonsingularity is the condition
The paper orders curves by the naive height
This height is adapted to the weights of the Weierstrass coefficients: has weight and has weight under the change of variables
Thus a box
has the rough shape
and contains on the order of curves.
The exponent is the background scale of the whole paper. The same exponent appears when one counts binary quartic forms by their invariants. The arithmetic content is not the exponent itself, but the exact average after imposing local solubility and dividing by the number of elliptic curves.
Selmer Groups and Rank Bounds
Let be an elliptic curve. The Kummer sequence for multiplication by gives a finite subgroup
cut out by local solubility conditions at every completion of . It fits into the exact sequence
Here is the Tate—Shafarevich group. The group is an elementary abelian -group, so
for an integer , the -Selmer rank.
The rank connection is immediate from the exact sequence. Since
one has
Therefore
and in particular
Bhargava and Shankar prove the sharper average statement
Since
for every nonnegative integer , this gives
The same upper bound holds for the average Mordell—Weil rank. The result is compatible with the Poonen—Rains heuristic model, which predicts average size for -Selmer groups .
The important point is that the theorem counts Selmer elements, not rational points directly. A Selmer element is a locally soluble covering. It may or may not have a rational point. This is exactly the difference between the computable upper bound and the actual rank.
Binary Quartic Forms
A binary quartic form is a homogeneous degree-four polynomial
Let denote the lattice of such forms. The group acts by linear substitution:
For the action, the invariant ring is generated by two classical invariants
and
The discriminant is
Under the full action these become relative invariants:
The height used for binary quartics is
It is homogeneous of degree in the coefficients of :
The elliptic curve attached to invariants is
The constants are not cosmetic. They align the invariant theory of quartics with the standard Weierstrass normalization used in the -descent correspondence.
The eligibility problem is already nontrivial. Not every pair occurs as the invariants of an integral binary quartic. Bhargava and Shankar prove that a pair is eligible if and only if it satisfies one of the following congruence conditions:
or
or
or
This finite congruence description is one reason the counting problem can be normalized cleanly.
Counting Orbits with Bounded Invariants
Let be the set of integral binary quartic forms of nonzero discriminant with pairs of complex conjugate roots, equivalently with real roots in . Let
denote the number of -orbits of irreducible forms in with .
The first major theorem of the paper is the asymptotic count
and
for every (Bhargava & Shankar, 2015).
This is a geometry-of-numbers theorem for an arithmetic group acting on a representation. The ambient real representation is
and one wants to count lattice points in a fundamental domain for
The main complication is that the fundamental domain is not compact. It has cuspidal regions where some coefficients become small and others become large.
The paper handles this cusp by separating reducible and irreducible behavior. In the cusp, most lattice points are reducible. Since irreducible quartics are the objects relevant to nontrivial Selmer elements, this allows the main term to be recovered from the noncuspidal region, where lattice-point counting is closer to volume computation.
The method has four moving parts:
- construct fundamental domains using reduction theory for ,
- average over compact sets in the group to make boundaries negligible,
- prove that the cusp contributes only a lower-order number of irreducible forms,
- extend the count uniformly to congruence-defined subsets of .
The last item is essential for Selmer groups. Local solubility is not a single congruence condition modulo one integer. It is an infinite family of conditions, one for each completion and . The proof therefore needs congruence counting together with a sieve that remains uniform as more local conditions are imposed.
The Selmer Correspondence
The bridge from quartics to elliptic curves is the classical interpretation of -coverings. A -covering of is a genus-one curve together with a diagram which becomes isomorphic over to multiplication by on . Soluble -coverings correspond to elements of , while locally soluble -coverings correspond to .
Binary quartic forms enter because a locally soluble -covering has a degree-two map to , so it can be represented as
This is the Birch—Swinnerton-Dyer binary quartic representation of -descent (Birch & Swinnerton-Dyer, 1963). It remains central in explicit algorithms for elliptic curves (Cremona, 1997; Cremona & Stoll, 2002; Cremona & Fisher, 2009).
Over a field of characteristic not or , let
There is a bijection between and -orbits of -soluble binary quartic forms with invariants and (Bhargava & Shankar, 2015). Explicitly, a point maps to the orbit of
The identity element corresponds to the orbit of quartics with a linear factor over . The stabilizer in of a nondegenerate quartic with invariants is isomorphic to .
For the global Selmer problem, Bhargava and Shankar use the integral version:
Integral Parametrization of -Selmer Elements
Let be an elliptic curve over . The elements of are in one-to-one correspondence with -equivalence classes of locally soluble integral binary quartic forms with invariants and . The class of forms with a rational linear factor corresponds to the identity element.
The powers of are normalization artifacts of passing between the invariant-theoretic model and an integral Weierstrass model. They do not affect the final average, but they do matter in the local mass calculation, especially at .
Local Solubility and Integral Models
For a field , a binary quartic form is -soluble if
has a solution with and . A rational form is locally soluble if it is soluble over and over for every prime .
The Selmer group imposes local solubility at every place. If one could count all binary quartics with prescribed invariants and then simply multiply by a fixed density of soluble forms, the proof would be much shorter. The actual difficulty is that local solubility depends on the orbit, the invariants, the prime, and the integral model.
The paper handles this by defining weighted sets of integral forms attached to large families of elliptic curves. The weighting compensates for stabilizers:
Counting orbits with weight is the natural stack-theoretic count. It is also the count compatible with the local mass formula, because stabilizers appear in the denominator of orbit volumes.
A large family is, roughly, a family of elliptic curves specified by acceptable local conditions. It includes all elliptic curves, families defined by finitely many congruence conditions on and , and natural families such as semistable curves (Bhargava & Shankar, 2015). The result is therefore stable under many arithmetic restrictions, not just for the undifferentiated set of all curves.
Uniformity enters through the following kind of estimate. The sieve must show that forms failing the desired local condition at some large prime contribute negligibly on average. This requires controlling forms whose discriminants have large square factors. Without such a bound, the infinite product of local densities would not be justified.
This is structurally similar to many arithmetic counting arguments: first count objects satisfying finitely many congruence conditions, then prove that the contribution from bad behavior at large primes is negligible. In this paper, the noncompact cusp makes this substantially harder than in a bounded lattice-point problem.
The Mass Formula
The decisive computation is not merely that an average exists, but that its value is exactly . The count of nonidentity Selmer elements is
because the identity class is represented by quartics with a rational linear factor and is treated separately.
For a large family , Bhargava and Shankar prove a formula of the form
where and are the corresponding real and -adic masses (Bhargava & Shankar, 2015).
The local ratios are then evaluated. At finite primes one uses the local identity
for , and
for . The real ratio contributes
After these local evaluations, the expression collapses to
With the paper’s normalizations this equals
Thus
and therefore
This is the conceptual center of the proof. The number does not arise from a crude inequality or a numerical accident. It comes from a global volume computation, with the Euler product matching the Tamagawa number of .
Consequences for Average Rank
The exact average size gives a rank bound by convexity at the level of powers of two. If
then
Averaging and using gives
Since
the limsup average rank is also at most .
This statement is weaker than the conjectural average rank , but it is unconditional and over the full two-parameter family of elliptic curves. Earlier finiteness results for all elliptic curves depended on conjectures such as GRH and Birch—Swinnerton-Dyer , while computations suggested subtle biases in finite-height data (Bektemirov et al., 2007).
The theorem also controls average -torsion in from above. The exact sequence gives
Since rational -torsion occurs in density zero in the full family, the same Selmer-rank bound gives an average upper bound for the -torsion dimension of .
Scope and Limitations
The result is a theorem about -Selmer groups, not a direct theorem about the distribution of Mordell—Weil ranks. A curve can have a nontrivial Selmer element that is not represented by a rational point; such elements measure -torsion in . Therefore the bound
is generally not an equality.
The proof also depends on the special structure of a coregular representation. Binary quartic forms have a polynomial invariant ring generated by and , and their orbits parametrize the relevant descent objects. This representation-theoretic clarity is not available for arbitrary arithmetic moduli problems.
The local solubility sieve is another delicate point. Counting all integral orbits with bounded invariants is not enough. One must count the locally soluble orbits with uniform control over infinitely many local conditions. The geometry-of-numbers theorem, the cusp analysis, and the squarefree-discriminant estimates are coupled parts of one argument.
Finally, the paper gives an upper bound for average rank, not the conjectural distribution. The Poonen—Rains model predicts much more detailed distributions for Selmer groups (Poonen & Rains, 2012), but the Annals result proves the exact first moment for -Selmer groups.
Transferable Mechanisms
The first transferable mechanism is the replacement of a difficult arithmetic object by integral orbits in a representation with explicit invariants. The Selmer element is hard to count directly; the binary quartic representative is countable by reduction theory. This same philosophy appears throughout arithmetic invariant theory, where orbit spaces encode algebraic structures and local conditions select arithmetic subclasses (Bhargava & Shankar, 2015; Borel & Harish-Chandra, 1962).
The second mechanism is counting with stabilizer weights. The natural quantity is not always the number of integral points, but the mass
This weighting is forced by the local-to-global product formula and by the fact that different orbits can have different automorphism groups. It is the same conceptual correction that appears whenever one counts moduli objects rather than rigid labeled objects (Bhargava & Shankar, 2015).
The third mechanism is finite congruence counting plus a uniform sieve for infinitely many local conditions. The paper first proves asymptotics for congruence-defined subsets of binary quartic forms, then proves that large-prime failures are negligible. This is the arithmetic analogue of isolating the bad region after a main equidistribution theorem, and it is structurally close to sieve arguments where finite local information is upgraded to an Euler product by uniform error estimates .
The fourth mechanism is the interpretation of an exact average as a volume identity. The number is explained by a product of real and -adic masses, and the nonidentity contribution is the Tamagawa number of . This is a useful template: when an arithmetic average is predicted to be a small integer, look for a mass formula whose Euler factors encode the local classification problem (Poonen & Rains, 2012; Bhargava & Shankar, 2015).
See Also
on orbit closures in moduli space — Both papers convert a global classification problem into controlled orbit behavior. Eskin—Mirzakhani—Mohammadi classify and isolate orbit closures in moduli space, while Bhargava—Shankar count arithmetic orbits in a representation using reduction theory and local masses.
on bounded gaps between primes — The local-to-global structure is different, but both arguments require a finite-level counting theorem together with uniform estimates strong enough to pass through a sieve over many congruence conditions.