unconditional-mertens-at-large-base.md

61.5 kB · markdown

Write the integers in a large base q and keep only those that avoid one chosen digit: they form a thin set S_F with A_F(x) = x^(alpha + o(1)) members up to x, alpha = log(q-1)/log q just under 1, and no multiplicative structure at all. This paper proves that the Mobius function cancels on that set with no hypothesis: |M_F(x)| <= C A_F(x) exp(-c sqrt(log x)) at every x >= 2, where M_F(x) sums mu(n) over the set up to x, at every base q >= 584 and every avoided digit by a proof, and at every q >= 115 on certificates computed base by base and digit by digit, with C and c effective and depending on q and the digit alone; the same holds at m avoided digits under a condition (W) on a proved one-step constant. The proof expands the indicator of the set in additive characters modulo q^n and cuts the frequencies by Dirichlet approximation into four regions. Far from every fraction of small denominator, the minor-arc bound of Basak, Robles and Zaharescu (2023) for mu pays the l^1 mass of the digit transform against x^(4/5). At middle denominators a hybrid l^1 bound, proved here from the one-step constant alone, pays the decay of that bound in the denominator. Near a fraction whose denominator has a prime outside the base the transform is itself small. Near a fraction whose denominator divides a power of the base the sum becomes mu in progressions to moduli built from the primes of q, where every possible exceptional zero belongs to a finite family of characters fixed by q, so Siegel's theorem is never used. The first region alone sets the base, through an l^1 bound on the digit transform at every shift: a digit-uniform chain of Dirichlet kernels proves it below the bar 1/5 from 584, window certificates verify it from 301 and per-digit certificates from 115, the least base this route reaches, while a one-step constant alone stops at 39363 and can never go below 33. The shape is in print for the primes: Maynard (2019) remarks the prime asymptotic for one avoided digit at every q >= 12 with an unquantified o(1), and Maynard (2022) proves it at q > 2 * 10^6 with a saving of any power of log x, remarks that q > 2500 is reachable by the same method, and remarks that his error terms could be made effective. What is new here is the Mobius function with an effective constant, and the wall 584, below both of those bases; the same dissection gives the prime count with an effective error of the same shape from the same base.

Introduction

Four concentric rings, each the whole circle of the 1000 frequencies a/1000 of base 10 at level 3, every frequency inked on exactly one ring by the region of its Dirichlet fraction at Q = 1000^(3/5) and Z = 8: the outer ring yellow on 40 frequencies near 0, 1/5, 1/4, 2/5, 1/2 and their mirrors, whose denominators divide a power of 10; the next ring orange on 26 frequencies near thirds, sixths and sevenths; the next ring dim on 202 frequencies, the flanks of those arcs and the fractions of denominator 8 to 15; the inner ring blue on the 732 minor-arc frequencies.
Four concentric rings, each the whole circle of the 1000 frequencies a/1000 of base 10 at level 3, every frequency inked on exactly one ring by the region of its Dirichlet fraction at Q = 1000^(3/5) and Z = 8: the outer ring yellow on 40 frequencies near 0, 1/5, 1/4, 2/5, 1/2 and their mirrors, whose denominators divide a power of 10; the next ring orange on 26 frequencies near thirds, sixths and sevenths; the next ring dim on 202 frequencies, the flanks of those arcs and the fractions of denominator 8 to 15; the inner ring blue on the 732 minor-arc frequencies.

Every integer below y = q^n has n digits, and whether it avoids the digit e_0 can be read off its additive characters: the indicator of the digit strings is y^(-1) sum_(a mod y) hat F_n(a/y) e(-ua/y), with hat F_n the digit transform of Definition 2.1. So a Mobius sum over the set becomes a weighted sum, over the y frequencies a/y, of the Mobius exponential sums sum mu(u) e(ua/y), and everything turns on how each frequency sits against the rationals of small denominator. The figure makes that cut at base 10, level 3: each ring is the whole circle of 1000 frequencies read clockwise from the top, and every frequency is inked on exactly one ring. The cut never sees the digits; only the weights |hat F_n(a/y)| it is paid against do. The outer ring holds the frequencies closest to a fraction whose denominator divides a power of 10, the next those closest to a fraction whose denominator has another prime, the third the flanks of those arcs and the middle denominators, and the inner ring everything else, the minor arcs. Base 10 is far below the range of the theorem; the picture is the cut, not the bound.

The cut at other bases, levels and Z, with the weights |hat F_n(a/y)| it is paid against, the set's own M_F(x) and prime count, and the chain's alpha_1 against the bar, is live in the dissection demo.

Theorem 1.1. Let q >= 3, let E be a set of m >= 1 digits, F = {0, ..., q-1} less E and k = q - m. Suppose F contains two consecutive digits and carries a shifted-grid l^1 certificate below 1/5, constants C_F >= 1 and alpha_1 < 1/5 with

sum_(a < q^i) |hat F_i(s + a/q^i)| <= C_F k^i q^(i alpha_1)     at every i >= 0 and every real s .

Then there are C > 0 and c > 0, depending on q and F alone and effectively computable, such that for every x >= 2

|M_F(x)| = |sum_(n in S_F, n <= x) mu(n)| <= C A_F(x) exp(-c sqrt(log x)) .

Corollary 1.2 (the walls). At one avoided digit the hypotheses of Theorem 1.1 hold at every base q >= 584 (Proposition 8.5), and they are verified at every 301 <= q <= 583 by the digit-uniform windows of Fact 8.6 and at every 115 <= q <= 300 by the per-digit certificates of Fact 8.7, so the bound holds at every one-avoided-digit set from base 115 on, proved from 584 and verified below it. Base 114 missing 56 carries no certificate below 1/5, so 115 is the floor of this route at one avoided digit. The wall condition

(W)   P_q(m) < k q^(-4/5) ,   or   m = 1, E = {e_0}, q >= 36 and P'_q(e_0) < (q - 1) q^(-4/5) ,

with the proved one-step constants P_q(m) and P'_q(e_0) of Definition 2.4, is a certificate with C_F = 1 that forces two consecutive digits. At one avoided digit it holds exactly from q = 39363, and from 28352 at the end digits e_0 in {0, q-1}, walls Verified by a float scan (Fact 8.3); at m avoided digits it is the certificate used here, it forces m < q^(2/5), and at q = 10^7 it admits every m <= 176.

The saving is of classical zero-free-region shape, not a power, and c is tiny: it is at most 1/40 of sqrt(|log rho|/24), where 1 - rho is about pi^2/(16 k q^2) (Section 7). The theorem says nothing at a set without a certificate, so nothing at any set of bounded fill; Proposition 10.2 shows the method itself needs k > q^(3/4).

The proof, after the blocks of Section 3 carry x off the powers of the base, bounds each block sum by y^(-1) sum_a |hat F_n(a/y)| |S_P(a/y)|, S_P the Mobius exponential sum over the block, region by region. Region A, the minor arcs, pays the whole l^1 mass of the transform against the uniform x^(4/5 + eps) of the minor-arc bound, and closes exactly when that mass grows like k^n y^(alpha_1) with alpha_1 < 1/5: that is the certificate at shift 0, and it is the only place the base is spent. Region B pays a hybrid l^1 mass, Lemma 5.1, against the d^(-1/2) decay of the same bound, and asks less, alpha_1 < 1/4. Region C1 needs no l^1 mass at all, because the transform at such a frequency is exponentially small in n, Lemma 6.1. Region C2 is where the arithmetic lives: there the block sum is mu in residue classes to moduli dividing a power of q, and Lemma 6.5 bounds it with explicit zero-free regions, every exceptional character being real of conductor dividing 8 rad(q).

The literature, as read here and detailed in Section 9. Maynard (2022), Theorem 1.1, proves sum_(n < q^j) Lambda(n) 1_(S_F)(n) asymptotic at one avoided digit for q > 2 * 10^6, from a minor-arc bound for Lambda with the same exponent 4/5 and four Fourier norms of the set, and remarks that q > 2500 suffices after a more involved calculation and that Siegel zeros play no role, so its error terms could be made effective of exp(-c sqrt(log x)) shape; Maynard (2019) already remarks the asymptotic itself, with an unquantified o(1), at every q >= 12. So for the primes what the large-base argument adds is the effective error shape. Its two inputs have standard Mobius analogues, so a Mobius version at q > 2 * 10^6 is a routine adaptation; it is not written there or in any source read here, and none of those sources carries a Mobius or Mertens sum over a missing-digit set. What this paper adds is that object written out, mu over the set with an effective constant, at the proved wall 584, below both the written 2 * 10^6 and the remarked 2500, the gain lying wholly in the l^1 input. It is first in its object and never in its shape.

Definitions and the one-step constant

Definition 2.1 (the set and its transform). Fix a base q >= 3, a set E of m >= 1 digits and F = {0, ..., q-1} less E, with k = q - m >= 2; q is the base and k the fill, and the two letters stand for them throughout. S_F is the set of positive integers whose every base-q digit lies in F, alpha = log k/log q its dimension, A_F(x) = #{n in S_F : n <= x} and M_F(x) = sum_(n in S_F, n <= x) mu(n), with mu(0) = 0 wherever 0 is summed. For n >= 0, D_n is the set of the k^n integers 0 <= u < q^n whose n padded digits lie in F. With e(t) = exp(2 pi i t), the digit transform is hat F(t) = sum_(a in F) e(at) and its level form is hat F_n(t) = prod_(i < n) hat F(q^i t) = sum_(u in D_n) e(ut), unnormalised, so |hat F_n| <= k^n; D_q(t) = sum_(a < q) e(at), with |D_q(t)| = |sin(pi q t)/sin(pi t)|, is the full kernel, and hat E(t) = sum_(a in E) e(at), so hat F = D_q - hat E.

Definition 2.2 (the one-step constant). B = B_q(F) = sup_t sum_(r mod q) |hat F((t + r)/q)|, the largest l^1 mass of one digit over a shifted grid of q points.

Lemma 2.3 (the peel and the floor). (i) For n >= 1, V = q^n and every real s, sum_(a < V) |hat F_n(s + a/V)| <= B^n, and the same sum with the factor of any one position replaced by 1 is at most q B^(n-1). (ii) int_0^1 |hat F_n| <= q^(-n) B^n. (iii) B >= q for every digit set, and B >= 2(q - 1) at one avoided digit.

Proof. (i) Write a = a' + q^(n-1) r with a' < q^(n-1), r < q, and t = s + a/V. The factors at positions 1, ..., n-1 form hat F_(n-1)(qt), and qt = qs + a'/q^(n-1) + r is qs + a'/q^(n-1) modulo 1, free of r; the factor at position 0 is hat F((u + r)/q) with u = q(s + a'/V). Summing over r first gives at most B times the same sum at level n - 1 and shift qs, or exactly q times it when the position-0 factor is 1; induct, peeling the lowest position each time. (ii) int_0^1 |hat F_n| = int_0^(1/V) sum_(a < V) |hat F_n(s + a/V)| ds. (iii) Parseval on Z/q gives sum_(r mod q) |hat F((t+r)/q)|^2 = qk at every t, and each term is at most k, so the l^1 sum is at least qk/k = q. At one avoided digit and t = 0, |hat F(0)| = q - 1 and |hat F(r/q)| = |0 - e(e_0 r/q)| = 1 at every r != 0. □

Definition 2.4 (the proved constants). Put H(n) = log n + gamma + 1/(2n), gamma Euler's constant, and

Phi_q = (4/pi) q + (2q/pi) H(ceil((q-2)/2)) + (1 - 2/pi)(q - 2) + 0.727 ,   P_q(m) = sqrt(m) + Phi_q/q .

At one avoided digit e_0 put c = e_0 - (q-1)/2, p = floor(q/2) and

Psi'_q = (q/pi)(2 H(p-1) - 1 + 1/p) + (1 - 2/pi) q/2                       at even q ,
Psi'_q = (q/pi)(2 H(p-1) - 1 + 2/p) + (1 - 2/pi)(q/2 + 1/(2q))              at odd q ,
P'_q(e_0) = ((4/pi) q + Psi'_q + q/2 - sec(pi c/q)/2)/q .

Write P_W for whichever constant the condition (W) of Corollary 1.2 reads, P_q(m) or P'_q(e_0), and put alpha_1 = log_q(q P_W/k). By Lemmas 2.5 and 2.6, B <= q P_W, so by Lemma 2.3(i) sum_(a < q^i) |hat F_i(s + a/q^i)| <= B^i <= k^i q^(i alpha_1) at every shift: (W) is a certificate in the sense of Theorem 1.1 with C_F = 1, and it is below 1/5 exactly when (W) holds. Any certificate bounds the l^1 exponent of the first-base paper, a limit over many digits, from above; the one-step certificate is the simplest and, by Proposition 8.8, never below 33.

Two elementary inequalities carry both constants: sin(pi v) <= 4v(1 - v) on [0, 1], and 1/sin z <= 1/z + 1 - 2/pi on (0, pi/2]. For the first, f(v) = 4v(1-v) - sin(pi v) vanishes at 0 and 1/2 and is symmetric about 1/2; f'' = pi^2 sin(pi v) - 8 changes sign once on [0, 1/2], so f is concave and then convex there, with f'(1/2) = 0: the convex piece decreases to f(1/2) = 0 and the concave piece lies above its chord. For the second, 1/sin z - 1/z increases on (0, pi/2] and equals 1 - 2/pi at the end. The harmonic number obeys H_n <= H(n), since H_n - log n - 1/(2n) is nondecreasing by the trapezoid rule on the convex 1/z and tends to gamma.

Lemma 2.5 (the kernel constant). For every q >= 3 and every set of m >= 1 avoided digits, B_q(F) <= q P_q(m).

Proof. Since hat F = D_q - hat E, B <= sup_t G(t) + sup_t sum_(r mod q) |hat E((t+r)/q)| with G(t) = sum_(r mod q) |D_q((t+r)/q)|. The avoided digits are distinct modulo q, so Parseval on Z/q gives sum_(r mod q) |hat E((t+r)/q)|^2 = qm and Cauchy-Schwarz gives sum_(r mod q) |hat E((t+r)/q)| <= q sqrt(m). For the kernel fix t in [0, 1) and let d_r be the distance from u_r = (t+r)/q to the nearest integer; |sin(pi q u_r)| = sin(pi t), so |D_q(u_r)| = sin(pi t)/sin(pi d_r). The two points nearest an integer sit at t/q and (1-t)/q and give at most sin(pi t)(q/(pi t) + q/(pi(1-t)) + 2(1 - 2/pi)) <= (4/pi) q + 0.727, by the two inequalities and 2(1 - 2/pi) < 0.727. Every other point has d_r >= min(r, q - 1 - r)/q, and each value j >= 1 of that minimum occurs at most twice, so those q - 2 points give at most sum (q/(pi j) + 1 - 2/pi) <= (2q/pi) H(ceil((q-2)/2)) + (1 - 2/pi)(q - 2). So G(t) <= Phi_q and B <= Phi_q + q sqrt(m) = q P_q(m). □

Lemma 2.6 (the chord constant). At one avoided digit e_0 and every q >= 36, B_q(F) <= q P'_q(e_0).

Proof. Four steps. First, the shifted grid is exact. For t in (0, 1) and u_r = (t+r)/q, D_q(u_r) = e((q-1) u_r/2) A_r with A_r = (-1)^r s/sin(pi u_r) and s = sin(pi t), so |hat F(u_r)| = |A_r - e(c u_r)|. For real A, |A - e(psi)|^2 = (|A| + 1)^2 - 2|A|(1 + sign(A) cos(2 pi psi)), and sqrt(1 - X) <= 1 - X/2 gives |hat F(u_r)| <= |A_r| + 1 - w_r (1 + sign(A_r) cos(2 pi c u_r)) with w_r = |A_r|/(|A_r| + 1) >= s/(1 + s) >= s/2, since |A_r| >= s. Second, the phases sum in closed form: sign(A_r) = (-1)^r, and since (-e(c/q))^q = -1, the geometric sum gives sum_(r mod q) (1 + (-1)^r cos(2 pi c u_r)) = q + cos(2 pi c (t - 1/2)/q) sec(pi c/q), every term of the left side being nonnegative. Hence, with G(t) = sum_r |A_r| as in Lemma 2.5,

sum_(r mod q) |hat F(u_r)| <= G(t) + q - (s/2)(q + cos(2 pi c (t - 1/2)/q) sec(pi c/q)) .

Third, the kernel under its chord: G(t) <= (4/pi) q + s Psi'_q. The Taylor series of csc z - 1/z has positive coefficients, so it is convex on (0, pi/2] and lies under its chord, csc z <= 1/z + (2/pi)(1 - 2/pi) z. Take t in (0, 1/2]; G(1 - t) = G(t) covers the rest. Pairing r with q - 1 - r writes G(t) = s sum_(r < p) (csc(pi (t+r)/q) + csc(pi (1-t+r)/q)), plus s csc(pi (t+p)/q) at odd q, every argument in (0, pi/2]. The 1/z half of the pair r = 0 is s q/(pi t (1-t)) <= (4/pi) q; that of a pair r >= 1 is at most (sq/pi)(1/r + 1/(r+1)), convex and symmetric in t, and these sum to (sq/pi)(2 H_(p-1) - 1 + 1/p), the odd middle term adding at most sq/(pi p). The chord halves add s (2/q)(1 - 2/pi) times the sum of the arguments over pi/q, which is p^2 = q^2/4 at even q and p(p+1) + t <= (q^2 + 1)/4 at odd q. With H_(p-1) <= H(p-1) that is (4/pi) q + s Psi'_q. Fourth, the maximum sits at t = 1/2. Write t = 1/2 + tau, so s = cos(pi tau) and the bound reads (4/pi) q + q + lambda(tau) with lambda(tau) = cos(pi tau)(Psi'_q - q/2 - cos(2 pi c tau/q) sec(pi c/q)/2), even in tau. Since |c| <= (q-1)/2, sec(pi c/q) <= 1/sin(pi/(2q)) <= q; with sin(pi tau) >= 2 tau and c sin(2 pi c tau/q) <= 2 pi c^2 tau/q on [0, 1/2], lambda'(tau) <= 2 pi tau (q + pi q/4 - Psi'_q), which is <= 0 once Psi'_q >= (1 + pi/4) q, and so once Psi'_q >= (1 + pi) q/2. That holds at q = 36 and q = 37, and Psi'_q/q increases along the even and along the odd bases, because 2(H(p) - H(p-1)) > 2/p - 1/(p(p-1)) outweighs the fall of the terms in 1/p and 1/q^2; so it holds at every q >= 36 (lab/py/mrly-pairing, verb onestep). So lambda(tau) <= lambda(0) = Psi'_q - q/2 - sec(pi c/q)/2, and B <= (4/pi) q + Psi'_q + q/2 - sec(pi c/q)/2 = q P'_q(e_0), the endpoint t = 0 by continuity. □

The bound falls as |c| grows: the end digits e_0 in {0, q-1} carry the smallest constant, sec(pi c/q) = 1/sin(pi/(2q)), and the digit nearest the middle the largest. Two comparisons are used in Section 8. First, P'_q(e_0) < P_q(1) at every q >= 36 and every e_0: with Psi_q = (2q/pi) H(ceil((q-2)/2)) + (1 - 2/pi) q, q P_q(1) - q P'_q(e_0) = (Psi_q - Psi'_q) + q/2 + sec(pi c/q)/2 + 0.727 - 2(1 - 2/pi), and Psi_q - Psi'_q is q/2 - 2/pi at even q and, at odd q = 2p + 1, (2q/pi)(H(p) - H(p-1)) + (q/pi)(1 - 2/p) + (1 - 2/pi)(q/2 - 1/(2q)), positive because H(p) - H(p-1) = log(p/(p-1)) - 1/(2p(p-1)) > 1/p - 1/(2p(p-1)) > 0. Second, no one-step constant beats the floor of Lemma 2.3(iii), the subject of Proposition 8.8.

The blocks

The expansion lives on a grid of q^n frequencies, so it sees digit strings of a fixed length; a general x is cut into such strings first.

Lemma 3.1 (the block expansion). For n >= 0, y = q^n and an integer P >= 0 put Sigma(P, n) = sum_(u in D_n) mu(Py + u) and S_P(theta) = sum_(Py <= v < (P+1) y) mu(v) e(v theta). Then

Sigma(P, n) = y^(-1) sum_(a mod y) hat F_n(a/y) S_P(-a/y) ,   so   |Sigma(P, n)| <= y^(-1) sum_(a mod y) |hat F_n(a/y)| |S_P(a/y)| .

Proof. Completeness of the additive characters modulo y gives 1_(D_n)(u) = y^(-1) sum_(a mod y) hat F_n(a/y) e(-ua/y) for 0 <= u < y. Put v = Py + u: e(-ua/y) = e(-va/y) e(Pa) and e(Pa) = 1. Since mu is real, |S_P(-theta)| = |S_P(theta)|. □

Lemma 3.2 (the split). Let x >= q have L digits. Then S_F below x is the disjoint union of the point x, when every digit of x lies in F, and of the sets Py + D_n at scales n < L, the element 0 discarded where it occurs, with at most k + 1 of them at each scale; and A_F(x) >= k^(L-1) - 1.

Proof. An element with L digits is x or differs from x first at some position j, counted from the bottom, where its digit f is below the digit of x. Above j it copies the digits of x, which must then all lie in F; f lies in F, with f >= 1 at the top position; below j its digits are free in F. So it lies in P q^j + D_j with P the digits of x above j followed by f, at most k such sets at each scale. The shorter elements are D_(L-1) less {0} when 0 is in F, one set at P = 0, and the sets D_l, 1 <= l < L, whose members have exactly l digits, when it is not. Either way the shorter elements hold at least k^(L-1) - 1 members of S_F below q^(L-1) <= x. □

Proposition 3.3 (the reduction). Fix kappa > 0 and put K = ceil(kappa sqrt(log x)). Suppose there are C' > 0, c' > 0 and x_0, depending on q and F alone and effective, with |Sigma(P, n)| <= C' k^n exp(-c' sqrt(log x)) for every x >= x_0 and every set Py + D_n of Lemma 3.2 with x/y < q^K. Then Theorem 1.1 holds with c = min(c', kappa log k)/2.

Proof. The sets at scales n < L - K hold at most (k + 1) sum_(n < L-K) k^n <= ((k+1)/(k-1)) k^(L-K) <= 6 k^(1-K) A_F(x) elements together, since k^(L-1) <= 2 A_F(x), and k^(-K) <= exp(-kappa log k sqrt(log x)). Every other set has y = q^n >= q^(L-K) > x q^(-K), and there are at most k + 1 at each scale, so by hypothesis they contribute at most C' exp(-c' sqrt(log x)) (k+1) sum_(n < L) k^n <= 6k C' A_F(x) exp(-c' sqrt(log x)). The point x contributes at most 1, and A_F(x) grows like x^alpha. Below any fixed x_1, |M_F(x)| <= A_F(x) is absorbed into C. □

So everything reduces to one block Py + D_n whose length y is within the factor q^K = x^(o(1)) of x. From here to Section 7, x is large, the block is fixed with x/y < q^K, and log y >= (log x)/2.

The dissection

Definition 4.1 (the four regions). Put Q = y^(3/5) and Z = exp(C_0 sqrt(log x)), with 0 < C_0 <= 1/2 fixed in Section 7. By Dirichlet's theorem every residue a mod y has a reduced fraction l/d with 1 <= d <= Q and |a/y - l/d| <= 1/(dQ); fix one, and call h = |ad - ly| its height, so h <= y/Q = y^(2/5). Then

A  :  d >= y^(2/5) ;
B  :  d < y^(2/5)  and  max(d, h) >= Z ;
C1 :  d < Z  and  h < Z ,  d with a prime factor not dividing q ;
C2 :  d < Z  and  h < Z ,  d dividing a power of q .

Since Z < y^(2/5) for x large, the four regions partition the residues. The figure draws them at q = 10, n = 3, Z = 8, the fraction taken as the last continued-fraction convergent of a/y with denominator at most Q: 732, 202, 26 and 40 frequencies, asserted in its binary.

Lemma 4.2 (the minor-arc input, quoted). For X >= 2, real theta, (l, d) = 1 with |theta - l/d| <= 1/d^2 and fixed eps > 0,

|S_mu(X, theta)| = |sum_(v <= X) mu(v) e(v theta)| <<_eps X^(4/5 + eps) + X d^(-1/2) (log X)^3 + (X d)^(1/2) (log X)^3 ,

with an effective implied constant depending on eps alone. This is Theorem 1.4 of Basak, Robles and Zaharescu (2023), read at source. Its proof is Vaughan's identity at U = V = min(X^(2/5), d, X/d) against the Type I and Type II estimates of Koukoulopoulos (2019), Theorems 23.5 and 23.6, whose constants are absolute, together with the divisor bound; the one Siegel-Walfisz input of that paper sits in its major-arc section and is not used here. It is the one analytic input of regions A and B not derived in this paper.

Lemma 4.3 (two approximations). Fix eps > 0. On region A, |S_P(a/y)| <<_eps x^(4/5 + 2 eps). On regions B, C1 and C2, |S_P(a/y)| <<_eps x^(4/5 + 2 eps) + x (log x)^3 max(d, h)^(-1/2).

Proof. S_P is the difference of two sums S_mu(X, a/y) with X < (P+1) y <= 2x, and y >= x q^(-K) = x^(1 - o(1)). Every fraction has d <= Q, so |a/y - l/d| <= 1/(dQ) <= 1/d^2 and Lemma 4.2 applies at l/d. On region A, y^(2/5) <= d <= y^(3/5) makes X d^(-1/2) <= 2x y^(-1/5) and (Xd)^(1/2) <= (2x y^(3/5))^(1/2), both x^(4/5 + o(1)); the powers of log x go into the second eps. On the other regions d < y^(2/5), and at l/d the last term is at most (2x y^(2/5))^(1/2) <= x^(7/10 + o(1)), which leaves x (log x)^3 d^(-1/2). When h >= 1 there is a second approximation. Dirichlet at level 2y/h gives a reduced l'/d' with d' <= 2y/h and |a/y - l'/d'| <= h/(2 d' y) <= 1/d'^2. It is not l/d, which sits at distance h/(dy) > h/(2dy), so 1/(d d') <= |l/d - l'/d'| <= h/(dy) + h/(2 d' y), that is y <= h d' + h d/2; as dh <= y^(4/5) <= y, d' >= y/(2h). At l'/d' the last two terms are X d'^(-1/2) <= 2x (2h/y)^(1/2) <= x^(7/10 + o(1)), since h <= y^(2/5), and (X d')^(1/2) <= (4xy/h)^(1/2) <= 2x h^(-1/2). The better of the two approximations gives max(d, h)^(-1/2). □

Proposition 4.4 (region A). Under a certificate (C_F, alpha_1), y^(-1) sum_(a in A) |hat F_n(a/y)| |S_P(a/y)| <<_eps C_F k^n y^(alpha_1 - 1/5 + 3 eps), a power saving when alpha_1 < 1/5, at eps = (1/5 - alpha_1)/4.

Proof. The certificate at s = 0 bounds the whole unshifted mass, sum_(a mod y) |hat F_n(a/y)| <= C_F k^n y^(alpha_1), and by Lemma 4.3 every residue of A carries at most x^(4/5 + 2 eps) <= y^(4/5 + 3 eps) for x large. □

Region A is the only region that pays the whole l^1 mass against x^(4/5), and so the only one that sets the wall, alpha_1 < 1/5; it reads the transform through the certificate at s = 0 and nothing else. Every other region either pays a smaller l^1 mass against a decaying bound or pays no l^1 mass at all.

The hybrid l^1 lemma

Below y^(2/5) the minor-arc bound decays like max(d, h)^(-1/2), and what it is paid against is the l^1 mass of the transform on the residues near fractions of one size of denominator and one size of height. That mass is bounded from the certificate alone; no large sieve is quoted.

Lemma 5.1 (the hybrid l^1 bound). Let D >= 1 and H >= 0 with 16 q^2 D (D + H) <= y, and let R(D, H) be the residues whose fraction has D <= d < 2D and h < 2H, or h = 0 when H = 0. Let V_1 = q^(i_1) be the least power of q at least 4D^2 and V_2 = q^(i_2) the least at least 4H/D + 1. Then V_1 V_2 < 16 q^2 D (D + H) <= y and

sum_(a in R(D, H)) |hat F_n(a/y)| <= C_F^2 (1 + pi q^2 C_F/k) k^n (V_1 V_2)^(alpha_1) ,

under a certificate (C_F, alpha_1), and at most (1 + pi q) k^n (V_1 V_2)^(alpha_1) under (W).

Proof. V_1 < 4q D^2 and V_2 < q (4H/D + 1) give the first claim, so i_1 + i_2 <= n. Split the positions into the lowest i_1, the middle and the top i_2: hat F_n(t) = hat F_(i_1)(t) hat F_(n - i_1 - i_2)(V_1 t) hat F_(i_2)(yt/V_2), and bound the middle factor by k^(n - i_1 - i_2). At t = a/y the top factor is hat F_(i_2)(a/V_2). The residues attached to one fraction l/d satisfy |a - ly/d| < 2H/d <= 2H/D, so they are at most 4H/D + 1 <= V_2 consecutive integers, distinct modulo V_2, and by the certificate at s = 0 they pay at most the grid sum C_F k^(i_2) V_2^(alpha_1), the fraction 0/1 read with 1/1 so that its residues are again consecutive modulo y and, V_2 dividing y, distinct modulo V_2. The bottom factor at each such residue is at most sup_J |f| with f = hat F_(i_1) and J = J_(l,d) = [l/d - 1/(8D^2), l/d + 1/(8D^2)], which holds every residue of l/d because |a/y - l/d| = h/(dy) < 2H/(Dy) <= 1/(8D^2) by 16 D H <= y. Distinct fractions of denominator below 2D differ by more than 1/(4D^2), so the intervals J_(l,d) are disjoint modulo 1. On each, |f(t)| <= |J|^(-1) int_J |f| + int_J |f'|, by averaging |f(t)| <= |f(u)| + |int_u^t f'| over u in J; so the suprema sum to at most 4D^2 ||f||_1 + ||f'||_1, norms on [0, 1]. The certificate at the shift u/V_1, integrated over u, gives ||f||_1 <= V_1^(-1) C_F k^(i_1) V_1^(alpha_1), so 4D^2 ||f||_1 <= C_F k^(i_1) V_1^(alpha_1). Differentiating the product one factor at a time costs |hat F'| <= 2 pi sum_(a in F) a <= pi q (q - 1) at a position j, and the product with that factor removed is |hat F_j(t)| |hat F_(i_1 - j - 1)(q^(j+1) t)|; its integral over [0, 1] splits at the period of the second factor into q shifted grids of the first and is at most C_F^2 (k q^(alpha_1 - 1))^(i_1 - 1). With sum_(j < i_1) q^j < V_1/(q - 1) and q^(alpha_1) >= 1 that gives ||f'||_1 <= (pi q^2/k) C_F^2 k^(i_1) V_1^(alpha_1), so the fractions pay at most C_F (1 + pi q^2 C_F/k) k^(i_1) V_1^(alpha_1), and the three factors multiply to the claim. Under (W) Lemma 2.3 does better: ||f||_1 <= V_1^(-1) B^(i_1) by (ii), and (i) with the factor at j replaced by 1 gives ||f'||_1 <= sum_(j < i_1) q^j pi q (q-1) V_1^(-1) q B^(i_1 - 1) <= pi q^2 B^(i_1 - 1) <= pi q B^(i_1) by the floor B >= q, whence 1 + pi q with C_F = 1. □

Proposition 5.2 (region B). If alpha_1 < 1/4, then

y^(-1) sum_(a in B) |hat F_n(a/y)| |S_P(a/y)| <<_eps k^n ( y^((4/5)(alpha_1 - 1/4) + 3 eps) (log y)^2 + q^K (log x)^5 Z^(-(1/2 - 2 alpha_1)) ) .

Proof. Group the residues of B into classes by powers of two, D <= d < 2D and H <= h < 2H, or h = 0, with D < y^(2/5) and H <= y^(2/5): at most (log_2 y + 2)^2 classes, each inside R(D, H), with 16 q^2 D (D + H) <= 32 q^2 y^(4/5) <= y for y large. On a class max(d, h) >= max(D, H) > Z/2, since max(d, h) >= Z, d < 2D and h < 2H. By Lemmas 4.3 and 5.1, with V_1 V_2 < 32 q^2 max(D, H)^2, a class weighs <<_eps y^(-1) (x^(4/5 + 2 eps) + x (log x)^3 max(D, H)^(-1/2)) k^n max(D, H)^(2 alpha_1). With max(D, H) <= y^(2/5) and x <= y^(1 + eps) the first term is << k^n y^((4/5)(alpha_1 - 1/4) + 3 eps); with x/y < q^K and 2 alpha_1 - 1/2 < 0 the second is << k^n q^K (log x)^3 (Z/2)^(-(1/2 - 2 alpha_1)). □

So region B asks only alpha_1 < 1/4, together with a block depth kappa log q < C_0 (1/2 - 2 alpha_1). With the one-step constant that is P_W < k q^(-3/4), whose chord form holds from q = 1499 on at one avoided digit (lab/py/mrly-pairing, verb onestep). Under a certificate below 1/5, 1/2 - 2 alpha_1 > 1/10.

The characters

Near a fraction of small denominator the minor-arc bound says nothing, and the two regions there are paid by the digits and by the arithmetic of mu respectively. Write ||z|| for the distance from z to the nearest integer.

Lemma 6.1 (the contraction at a grid point). Let F contain two consecutive digits. Let d = e d_0 with e dividing a power of q, (d_0, q) = 1 and d_0 > 1, let (l, d) = 1, and let |eta| < q^(-2n/3)/(4q(q-1)). With m_d = max(1, floor(log_q(d/2)) + 1) and rho = 1 - (2/k)(1 - cos(pi/(4q))),

|hat F_n(l/d + eta)| <= k^n rho^(floor(2n/(3 m_d))) .

Proof. Let v, v + 1 lie in F. Then |hat F(z)| <= k - 2 + |e(vz) + e((v+1)z)| = k - 2 + 2 cos(pi ||z||). Put y_i = ||q^i l/d||. Since d_0 > 1 is prime to q and to l, q^i l/d is never an integer, so y_i >= 1/d; and y_i < 1/(2q) forces y_(i+1) = q y_i. If m_d consecutive positions all had y < 1/(2q), the last would be at least q^(m_d - 1)/d > 1/(2q), since q^(m_d) > d/2; so every window of m_d consecutive positions holds a position with y_i >= 1/(2q). At a position i < 2n/3, q^i |eta| < 1/(4q(q-1)) <= 1/(4q), so at such a position ||q^i (l/d + eta)|| > 1/(4q) and the factor is at most k - 2 + 2 cos(pi/(4q)) = k rho. The positions below 2n/3 hold at least floor(2n/(3 m_d)) disjoint windows, and every other factor is at most k. □

This is the perturbed form of Lemma A' on the coprimality page, at dimension one and stated at the grid point itself, so no passage from the exact fraction is needed. The same decay is Lemma 8.2 of Maynard (2019) at base 10 and Lemma 5.4 of Maynard (2022) at large base, each with a constant whose size is not given; here rho is explicit, and |log rho| is about pi^2/(16 k q^2).

Lemma 6.2 (two consecutive digits). Under (W), F contains two consecutive digits, and at m >= 2 avoided digits m < q^(2/5).

Proof. At m = 1 it is clear. At m >= 2 the condition is P_q(m) < k q^(-4/5) < q^(1/5), and P_q(m) >= sqrt(m) + 4/pi > 2 since Phi_q >= (4/pi) q; so q^(1/5) > 2 and sqrt(m) < q^(1/5), whence m < q^(2/5) < (q-1)/2 at every q >= 33. A digit set with no two consecutive digits has at most ceil(q/2) digits, so it avoids at least floor(q/2) >= (q-1)/2. □

Proposition 6.3 (region C1). If F has two consecutive digits, then for x large y^(-1) sum_(a in C1) |hat F_n(a/y)| |S_P(a/y)| <= 3 rho^(-1) k^n exp((2 C_0 - |log rho|/(6 C_0)) sqrt(log x)).

Proof. A residue of C1 is a/y = l/d + eta with |eta| = h/(dy) < Z/y, which is below y^(-2/3)/(4q(q-1)) once 4 q^2 Z <= y^(1/3); its denominator is d = e d_0 as in Lemma 6.1. So |hat F_n(a/y)| <= k^n rho^(floor(2n/(3 m_d))) with m_d <= log_q Z + 1. With n = log y/log q >= log x/(2 log q) and log_q Z = C_0 sqrt(log x)/log q, 2n/(3 m_d) >= log x/(3 (C_0 sqrt(log x) + log q)) >= sqrt(log x)/(6 C_0) once C_0 sqrt(log x) >= log q. Each fraction l/d holds the residues with |ad - ly| < Z, at most 2Z/d + 1 of them, so C1 holds at most sum_(d < Z) phi(d)(2Z/d + 1) <= 3 Z^2 residues, and |S_P| <= y trivially. □

Region C1 saves once C_0^2 < |log rho|/12. Region C2 is the arithmetic of mu itself: at a fraction whose denominator divides y, the block sum is a sum of mu over residue classes, and the classes are paid by zero-free regions. Two lemmas prepare it.

Lemma 6.4 (real characters of base-smooth modulus). A real primitive character whose modulus divides a power of q has conductor dividing 8 rad(q).

Proof. A primitive character factors into primitive characters modulo the prime powers of its conductor, each real when the character is. For an odd prime p, the kernel of reduction from (Z/p^j)^* to (Z/p)^* has odd order p^(j-1), so a character of order at most 2 is trivial on it and has conductor dividing p. For p = 2, every n = 1 mod 8 is a square modulo 2^j, so a real character modulo 2^j factors through (Z/8)^*. □

Lemma 6.5 (Mobius in progressions to base-smooth moduli). There are effective C_1, c_1 > 0, depending on q alone, such that for every X >= 2, every u <= X, every d dividing a power of q with d <= exp(sqrt(log X)/2) and every residue b,

|M(u; d, b)| = |sum_(v <= u, v = b mod d) mu(v)| <= C_1 X exp(-c_1 sqrt(log X)) .

Proof. If u <= X exp(-sqrt(log X)) it is trivial. Otherwise put g = (b, d). Every v = b mod d has (v, d) = g, so M = 0 unless g is squarefree; then mu(v) = mu(g) mu(v') for v = g v' with (v', g) = 1, and v' = b/g mod d/g. Let g_1 be the product of the primes of g not dividing d/g; by the Chinese remainder theorem the admissible v' fill at most phi(g_1) < rad(q) classes coprime to d_1 = (d/g) g_1, a modulus dividing a power of q with d_1 <= d rad(q), and orthogonality bounds each class by max_(chi mod d_1) |M(u/g, chi)|, with M(w, chi) = sum_(v <= w) mu(v) chi(v). Let chi be induced by the primitive chi* modulo q* | d_1. Then 1/L(s, chi) = L(s, chi*)^(-1) prod_(p | d_1) (1 - chi*(p) p^(-s))^(-1), so M(w, chi) = sum_t chi*(t) M(w/t, chi*) over the t <= w built from the primes of d_1, at most (1 + log_2 w)^(omega(q)) of them. The terms with w/t < sqrt(w) are trivially < sqrt(w); every other has W = w/t >= sqrt(w) >= X^(1/4) for X large, as w = u/g >= X exp(-(3/2) sqrt(log X)), and conductor q* <= rad(q) exp(sqrt(log X)/2) <= rad(q) exp(sqrt(log W)). It remains to bound M(W, chi*) by W exp(-c' sqrt(log W)) with c' > 0 effective and depending on q alone; summing back over t and the classes then gives the lemma with c_1 = min(1, c'/3) and C_1 >= 1, which also covers the trivial range.

At q* = 1 this is the Mertens function, M(W) << W exp(-c_2 sqrt(log W)), Exercise 8.4 of Koukoulopoulos (2019), read at source with its hint. At q* >= 3, the truncated Perron formula, Theorem 7.2 there, at sigma_0 = 1 + 1/log W and T = exp(sqrt(log W)) gives M(W, chi*) = (2 pi i)^(-1) int_(sigma_0 - iT)^(sigma_0 + iT) W^s ds/(s L(s, chi*)) + O(W (log W)/T + 1). By Lemma 2.1 of Chang and Martin (2019), read at source, L(s, chi*) is zero-free in sigma >= 1 - c_3/log(q* (|tau| + 4)), c_3 effective and absolute, except possibly for one real zero of one character, and for every other chi* it has |log L(s, chi*)| <= log log(q* (|tau| + 4)) + O(1) in half that width, so |1/L| <= exp |log L| << log(q* (|tau| + 4)) by their Lemma 2.4 at z = -1. Moving the segment to sigma_1 = 1 - (c_3/2)/log(q* (T + 4)) costs << W^(sigma_1) (log W)^4 + W (log W)^4/T, and log(q* (T + 4)) <= 3 sqrt(log W) makes it << W exp(-(c_3/7) sqrt(log W)). If chi* has the exceptional zero, it is real, because a complex chi* would share the zero with its conjugate and the lemma allows one character; so by Lemma 6.4 its conductor divides 8 rad(q), one of a finite family fixed by q. For such a character Proposition 2.3 of the same paper makes L(s, chi*) zero-free in sigma >= 1 - c_0/(1 + log^+ |tau|), with c_0 = eta_0/(q*^(1/2) log^2 q*) and eta_0 effective, and bounds |log L| <= (1/2) log q* + 3 log log(q* (|tau| + 4)) + O(1) there; moving the segment to sigma_1 = 1 - c_0/(1 + log T) gives <<_q W exp(-(c_0/3) sqrt(log W)), and c_0 is bounded below by q alone. So c' = min(c_2, c_3/7, c_0/3) over the finite family. □

Siegel's theorem is never used: the only exceptional zeros in play belong to finitely many real characters of conductor dividing 8 rad(q), fixed with the base, and their zeros are bounded away from 1 effectively.

Proposition 6.6 (region C2). With N_Z <= (1 + log_2 Z)^(omega(q)) the number of d < Z dividing a power of q, for x large,

y^(-1) sum_(a in C2) |hat F_n(a/y)| |S_P(a/y)| <= 96 C_1 k^n q^K Z^2 N_Z exp(-c_1 sqrt(log x)) .

Proof. Let d < Z divide a power of q. Each prime power p^j exactly dividing d has j < log_2 Z < n, so d divides y, and a = ly/d + j' with j' = (ad - ly)/d an integer, |j'| = h/d < Z/d. Then S_P(a/y) = sum_v mu(v) e(vl/d) e(v j'/y) over the block, and partial summation against e(v j'/y), whose total variation over the block is at most 2 pi |j'|, gives |S_P(a/y)| <= (1 + 2 pi |j'|) max_u |sum_(Py <= v <= u) mu(v) e(vl/d)|. The inner sum is sum_(b mod d) e(bl/d) (M(u; d, b) - M(Py - 1; d, b)), so |S_P(a/y)| <= 2 (d + 2 pi Z) max |M(u; d, b)| <= 16 Z max |M(u; d, b)|, over u <= 2x. Lemma 6.5 at X = 2x applies, since d < Z <= exp(sqrt(log x)/2), and bounds that maximum by 2 C_1 x exp(-c_1 sqrt(log x)). C2 holds at most N_Z (2Z + Z) = 3 Z N_Z residues, |hat F_n| <= k^n on them, and x/y < q^K. □

Assembly

Proof of Theorem 1.1. Assume a certificate (C_F, alpha_1) with alpha_1 < 1/5, and two consecutive digits in F. Take

C_0 = min(1/2, sqrt(|log rho|/24), c_1/4) ,   kappa = C_0/(20 log q) ,   eps = (1/5 - alpha_1)/4 ,

so that q^K <= q exp((C_0/20) sqrt(log x)). Fix a block with x/y < q^K, x large. By Lemma 3.1 its sum is at most the sum of the four region bounds, each k^n times a saving. Region A saves the power y^(-(1/5 - alpha_1)/4) by Proposition 4.4. Region B saves a power in its first term and, since 1/2 - 2 alpha_1 > 1/10, q^K Z^(-1/10) <= q exp(-(C_0/20) sqrt(log x)) in its second, by Proposition 5.2. Region C1 saves exp(-(|log rho|/(12 C_0)) sqrt(log x)) by Proposition 6.3, because C_0^2 <= |log rho|/24 makes 2 C_0 <= |log rho|/(12 C_0). Region C2 saves exp(-(c_1 - 2 C_0 - C_0/20) sqrt(log x)) by Proposition 6.6, and c_1 >= 4 C_0 makes that exponent at least (2 - 1/20) C_0. Every saving is at least exp(-(C_0/20) sqrt(log x)) up to powers of log x and constants depending on q and F, so the block bound of Proposition 3.3 holds with c' = C_0/21, and Theorem 1.1 follows with c = min(C_0/21, kappa log k)/2. Every constant is effective: the certificate's C_F, the implied constant of Lemma 4.2 at this eps, the constants c_2, c_3, eta_0 and C_1 of Lemma 6.5, and every threshold on x used above, each of which is explicit in q. □

No step used m = 1: region A reads the certificate at s = 0, region B asks alpha_1 < 1/4 and reads the certificate at every shift through Lemma 5.1, the split of Lemma 3.2 counts at most k + 1 sets per scale at every m, region C2 never sees F, and region C1 asks for two consecutive digits. (W) supplies both hypotheses at once, by Lemma 6.2. The saving is tiny: c <= kappa log k/2 <= C_0/40 <= sqrt(|log rho|/24)/40, and |log rho| <= pi^2/(15 k q^2) because 1 - cos z <= z^2/2, so at one avoided digit the c of this proof is below q^(-3/2)/200.

The wall

The proof uses its hypotheses at three points only: the certificate at shift 0 in region A, which asks alpha_1 < 1/5; the certificate at every shift in Lemma 5.1, where region B asks only alpha_1 < 1/4; and two consecutive digits in region C1. The base is spent at the first, through the exponent 4/5 of Lemma 4.2 and the certificate, and this section supplies certificates at one avoided digit and nothing else: the one-step constants of Section 2, which give (W) from 39363; a digit-uniform chain over all digits of the transform, which proves the wall 584; digit-uniform window certificates, which reach 301; and per-digit window certificates, which reach 115. A sharper certificate replaces this section alone.

Fact 8.1 (the wall scan). Over 36 <= q <= 2 * 10^5 and every avoided digit, the chord form of (W), P'_q(e_0) < (q-1) q^(-4/5), holds at every digit exactly from q = 39363 on and at the two end digits e_0 in {0, q-1} exactly from q = 28352 on: each set of bases is an up-set of the scan, holding at 160638 = 2 * 10^5 - 39363 + 1 and at 171649 bases. The worst digit is the middle one at odd q and one with c = 1/2 at even q. The least float gap (q-1) q^(-4/5) - P'_q(e_0) above each wall sits at the wall itself; re-read at 40 digits it is <= -2.3195 * 10^(-6) at 39362 against >= 2.3677 * 10^(-5) at 39363, and <= -1.3539 * 10^(-5) at 28351 against >= 1.8837 * 10^(-5) at 28352, far above float error at every base of the scan, though no base of it is interval-certified. With the kernel constant P_q(1) in place of the chord the same scan closes at 92317 at every digit, an up-set of 17 <= q <= 2 * 10^5, the gap <= -5.4242 * 10^(-6) at 92316 against >= 2.1054 * 10^(-6) at 92317. At the walls 1/5 - alpha_1 >= 2.6965 * 10^(-7) at 39363, 2.3643 * 10^(-7) at 28352 and 1.8712 * 10^(-8) at 92317, printed from the 40-digit gap as log(1 + gap/P_W)/log q. Generator: lab/py/mobius-dissection, verb wall.

Lemma 8.2 (the gap climbs). Let g(q) = (q-1) q^(-4/5) - P_q(1). Then g(q+1) > g(q) at every q >= 11221.

Proof. P_q(1) = 1 + 4/pi + (2/pi) H(ceil((q-2)/2)) + (1 - 2/pi)(1 - 2/q) + 0.727/q. From q to q + 1 the harmonic term moves at most once, by H(j+1) - H(j) <= 1/j <= 2/(q-2) with j = ceil((q-2)/2); the term (1 - 2/pi)(1 - 2/q) rises by 2(1 - 2/pi)/(q(q+1)) < 0.017/(q-2) for q >= 40; the last term falls. So P_(q+1)(1) - P_q(1) < (4/pi + 0.017)/(q-2) < 1.291/(q-2). The mass term has derivative q^(-9/5)(q/5 + 4/5) >= (1/5) q^(-4/5), so it rises by at least (1/5)(q+1)^(-4/5) per step, and g climbs wherever (1/5)(q-2)(q+1)^(-4/5) >= 1.291. The left side increases with q and first reaches 1.291 at q = 11221, the floor that lab/rs/mertens-numerology prints at the exponent 4/5. □

Fact 8.3 (the one-step walls). At one avoided digit the chord form of (W) holds at every q >= 39363 and every e_0, and at every q >= 28352 at the end digits, and fails at some digit at 39362 and at the end digits at 28351. On q <= 2 * 10^5 this is the float scan of Fact 8.1, re-read at 40 digits at the walls and certified in interval arithmetic nowhere, which is why these walls are Verified and never Proved; above the scan, P'_q(e_0) < P_q(1) at every digit by the comparison after Lemma 2.6, and P_q(1) < (q-1) q^(-4/5) because g(92317) > 0 by Fact 8.1 and g climbs from 11221 on by Lemma 8.2. Generator: lab/py/mobius-dissection, verb wall.

Fact 8.4 (the digit budget). At q = 10^7, P_q(m) < (q - m) q^(-4/5) holds exactly for m <= 176, the maximum asserted maximal. Generator: lab/rs/mertens-numerology, the m-budget table at the exponent 4/5.

Proposition 8.5 (the chain clears 1/5 from 584). At one avoided digit and every q >= 3, F carries the certificate C_F = z_q, q^(alpha_1) = z_q q/(q-1), where z_q > 1 is the root of (z - 1)^3 = (2/pi)(log q) z + gamma' (z - 1) + (2/pi)(z - 1)^2/(qz - 1), gamma' = 0.9625229 rounded up. It is below 1/5, that is z_q < q^(1/5)(1 - 1/q), at every q >= 584 and every avoided digit, and not at 583.

Proof. Section 4 of the first-base paper, Lemmas 4.1 to 4.5 and Corollary 4.6 there, bounds the grid sum at every shift: sum_(a < q^N) |hat F_N(x + a/q^N)| <= q^N A_N at every real x, with A_0 = 1 and A_N = A_(N-1) + sum_(l < N) lambda_l A_(N-1-l) + lambda_N, and Lemma 4.4(iii) there gives lambda_l <= lambda_l^+ = (2/pi) l log q + gamma' + (2/pi) q^(-l). Let z = z_q be the root of z = 1 + sum_(l >= 1) lambda_l^+ z^(-l), which is the cubic above after summing the three series. Induction gives A_N <= z^(N+1): A_N <= z^N + sum_(l=1)^N lambda_l^+ z^(N-l) = z^N (1 + sum_(l <= N) lambda_l^+ z^(-l)) <= z^(N+1). So the grid sum is at most z (zq)^N = C_F k^N q^(N alpha_1) with k = q - 1. The inequality z_q < q^(1/5)(1 - 1/q) is certified in interval arithmetic at 120 bits at every 584 <= q <= 1272, tightest margin 6.0170 * 10^(-3) in the cubic's units at 584, where 1/5 - alpha_1 >= 1.79 * 10^(-5), and fails at 583 with margin -8.3138 * 10^(-3). From 1272 on, the cubic with z_q <= 2(z_q - 1) gives the cap z_q <= 1 + sqrt(2 (2/pi) log q + 0.97), as in the proof of Theorem 4.7 there, which sits below q^(1/5)(1 - 1/q) at 1272 with gap 4.6127 * 10^(-4), and the gap grows from q = 100 on because q^(1/5)(1 - 1/q) sqrt(2 (2/pi) log q + 0.97) > 5 (2/pi) there (lab/py/prime-dissection, verb wall). □

Fact 8.6 (the window certificates). The same digit-uniform majorant run through the window machine of the first-base paper at two window digits, one outward-rounded Collatz-Wielandt certificate per base covering every avoided digit at once, bounds the grid sum at every shift with an effective constant and puts alpha_1 < 1/5 at every 301 <= q <= 583, the largest bound 0.199923 at 301; it reads 0.200021 at 300. Generator: lab/py/prime-dissection, verb window.

Fact 8.7 (the per-digit certificates). The window machine of the first-base paper run on |hat F| itself, each cell supremum bounded by its midpoint value plus the local first and second derivatives rounded outward, is a shifted-grid certificate with an effective constant, since the q^L points of any shifted grid fall in distinct cells of depth L. It certifies alpha_1 < 1/5 at every avoided digit of every base 115 <= q <= 300, at two window digits or, at 115 to 122, three, among them base 124 missing 61 at alpha_1 < 0.1999974. Base 114 missing 56 reads [0.2000730, 0.2001280] at three window digits, and every base 3 <= q <= 114 carries a digit certified above 1/5; the first base with some digit below 1/5 is 65, missing 0 at alpha_1 < 0.1996822, and all 1054 distinct sets of the bases 3 to 64 are certified above. Generator: lab/py/digit-transform-norms, verbs fifth and fifthbelow.

Proof of Corollary 1.2. From 584, Proposition 8.5, with two consecutive digits automatic at one avoided digit; on 301 <= q <= 583, Fact 8.6; on 115 <= q <= 300, Fact 8.7, whose base 114 gives the floor. The statements on (W) are Fact 8.3, Lemma 6.2 and Fact 8.4. □

The exponent 4/5 is the one Maynard (2022) calls a limit of its own method, "ultimately related to the 4/5 exponent of Lemma 4.2 for an exponential sum over primes", and that paper carries no Type I or Type II estimate that would move it. The certificate is the lever this section pulls, and its one-step form has a floor.

Proposition 8.8 (the floor at 33). At one avoided digit, if some constant P_W satisfies B_q(F) <= q P_W and P_W < (q-1) q^(-4/5), then q >= 33. So no bound on the one-step constant, however sharp, brings the wall of this dissection below 33.

Proof. Lemma 2.3(iii) gives q P_W >= B >= 2(q - 1), so (q-1) q^(-4/5) > 2(q-1)/q, that is q^(1/5) > 2, that is q > 32. □

The floor is a floor for one-step constants only. Region A itself pays the unshifted mass sum_(a mod q^n) |hat F_n(a/q^n)| at all n digits at once, whose only proved floor is Parseval's q^n, and a bound on that mass is not a power of a one-step constant. Its growth is not bounded here: the ratios of successive masses read 74.1654 at base 33 missing 16 and 70.5661 missing 0 from three to four digits, against 2(q-1) = 64, and 35.5234 at base 17 missing 8 from four to five, against 32, readings above the one-step floor that prove nothing (lab/py/mobius-dissection, verb wall). The chain of Proposition 8.5 is such a multi-digit bound, and it is why the wall sits below 39363; the exact supremum B is not certified at any base here.

What is in print

Every source below is listed in the references, and Maynard (2019), Maynard (2022), Nath (2024), Leng and Sawhney (2025), Basak, Robles and Zaharescu (2023), Koukoulopoulos (2019) and Chang and Martin (2019) are read at source.

Maynard (2022), Theorem 1.1, proves for q > 2000000, one avoided digit and every A > 0 that sum_(n < q^j) Lambda(n) 1_(S_F)(n) = kappa_q(e_0) (q-1)^j + O_A((q-1)^j (log q^j)^(-A)), and remarks that "a more involved calculation shows that q > 2500 is sufficient by the same method". It also remarks that its estimates are used only at highly composite moduli, where Siegel zeros play no role, so the error terms could be replaced by effective ones of size O((q-1)^j exp(-c j^(1/2))). Its prime input is its Lemma 4.2, sum_(n < x) Lambda(n) e(n alpha) << (x^(4/5) + x^(1/2) |d beta|^(-1/2) + x |d beta|^(1/2)) (log x)^4 at alpha = a/d + beta, the exponent 4/5 it calls a limit of its method; the set enters only through four Fourier norms, an l^1 bound (Lemma 5.1), a large sieve (Lemma 5.2), a hybrid bound (Lemma 5.3) and an l^infinity bound with an unsized constant (Lemma 5.4); its Theorem 1.3 takes s < q^(1/5 - eps) avoided digits. It carries no Type I or Type II estimate and no Mobius or Mertens sum.

The correspondence with this paper is close. Region A is his Lemma 4.2 paid against his Lemma 5.1, with the bound of Basak, Robles and Zaharescu (2023) for mu, Lemma 4.2 here, in place of the prime bound, whose 4/5 it shares. Lemma 5.1 here plays the part of his Lemma 5.3, derived from the one-step constant with no large sieve, and Lemma 6.1 plays the part of his Lemma 5.4, with an explicit constant. Region C2, mu in progressions to moduli dividing a power of the base with the exceptional characters confined to a finite family, is the Mobius form of his remark on Siegel zeros. So the shape of Theorem 1.1 is in print for Lambda, a Mobius version at q > 2 * 10^6 is a routine adaptation of it, and that adaptation is not written in the paper or in any other source read here. The theorem here holds from the proved wall 584, and from 115 by certificate, below his written 2 * 10^6, read in arXiv:1510.07711v1, and below his remarked 2500. His own gate is the same alpha_q < 1/5, for the constant of his Lemmas 5.1 and 5.3, which first drops below 1/5 at q = 1520573 (lab/py/prime-dissection, verb wall); the gain here lies wholly in the l^1 input. For the primes the asymptotic itself is older: Maynard (2019), p. 3 of the arXiv version, remarks that at one avoided digit its methods give #{p in A'} = (kappa + o(1)) #A'/log X for every q >= 12, with the o(1) unquantified. So for Lambda what the large-base argument adds is the error shape, and for mu it is the object itself with an effective constant.

Around it: Maynard (2019) proves the primes infinite in base 10 with one avoided digit, through a Type I estimate for the set and an l^infinity bound, Lemma 8.2 there; Nath (2024) proves Bombieri-Vinogradov theorems for Lambda 1_(S_F) at large base; Leng and Sawhney (2025) prove ternary Goldbach on the set; Erdos, Mauduit and Sarkozy (1998) distribute such sets in residue classes. Among the sources read here, the nearest multiplicative function computed over a missing-digit set is the divisor function of Kim (2024), whose own framing is that the lack of multiplicative structure blocks the standard approaches. A whole-text search of Maynard (2019), Maynard (2022) and Nath (2024) finds Mobius once, as an inversion step inside a proof, Liouville nowhere, and Mertens only as Mertens' theorem on a product over primes: none of the sources read here carries a Mobius or Mertens sum over a missing-digit set.

The prime count. On the coprimality page the same dissection is run with Lambda in place of mu, and under the same hypotheses, a certificate below 1/5 and two consecutive digits, it proves |sum_(n <= x, n in S_F) Lambda(n) - kappa_F A_F(x)| <= C A_F(x) exp(-c sqrt(log x)) at every x >= 2, with kappa_F = (q/phi(q)) #{f in F : (f, q) = 1}/k and C, c effective; so it holds at every one-avoided-digit set from 584 by a proof, and from 115 by the certificates of Facts 8.6 and 8.7. Its proof changes four things, the main term, which is the principal character at the C2 points, the minor-arc input, which is Lemma 4.2 of Maynard (2022), the constant of Lemma 5.1, and the arithmetic input, primes in progressions to moduli dividing a power of q with the exceptional zeros confined as in Lemma 6.5; it is not restated here. The asymptotic itself is remarked by Maynard (2019) at every q >= 12 with an unquantified o(1), and proved by Maynard (2022) at q > 2 * 10^6 with the effective shape only remarked, so what the dissection adds for the primes is an effective error of zero-free-region shape from 584.

In companion work the same expansion without the dissection proves, under the generalized Riemann hypothesis, a power saving against A_F(x) at one avoided digit from q = 1499 on, 1032 at the end digits, on the Mobius page, and from q = 34 on with the certified multi-digit exponent of the first-base paper, Theorem 5.1 and Corollary 6.5 there; its Proposition 5.3 shows that without the hypothesis and without a dissection that expansion returns nothing. The price of dropping the hypothesis, paid entirely in the base, is the step from 1499 to the wall of Corollary 1.2, and the saving falls from a power to exp(-c sqrt(log x)).

The check and the limits

The proof is a chain of bookkeeping, and every link that can be computed was computed on whole grids small enough to hold, far below the range of the theorem, where a wrong count or a wrong algebraic step would show.

Fact 10.1 (the falsification). On every residue of the grids of base 10 missing 5 and missing 0 at level 6, and base 5 missing 2 at level 9, with Z = 16, 16 and 12 and each fraction the last continued-fraction convergent at Q = y^(3/5): the four region sums of Lemma 3.1 add to the exact sum of mu over the strings to float precision; C1 and C2 hold at most 3 Z^2 and 3 Z N_Z residues; every C2 denominator divides y, with |j'| d = h < Z; the second approximation of Lemma 4.3 lands in [y/(2h), 2y/h] and differs from l/d at all 74778 and 128942 residues of B with h >= 1; and the bound of Lemma 5.1 holds at every class meeting the two conditions its proof uses, V_1 V_2 <= y and 16 D H <= y, a wider set than its hypothesis, 73 classes at each base-10 grid and 86 at base 5, the largest ratio past the class of a = 0 at most 0.001356 and the step 4D^2 ||f||_1 + ||f'||_1, both norms read on a grid, at most 0.110814 of its bound. Lemma 6.1 holds at 3000 seeded grid points of level 30 in base 10 and 3000 of level 45 in base 5, its logarithm exceeding that of |hat F_n(a/y)| by at least 27.9 at every one. The split of Lemma 3.2 meets M_F(x) exactly at 400 random x below 2 * 10^6 and at every q^e - 1 in five one-avoided-digit families, with at most k sets at one scale and A_F(x) >= k^(L-1) - 1 throughout. Generator: lab/py/mobius-dissection, verbs regions and blocks.

gridexactABC1C2
base 10 missing 5, level 68-156.4922+153.9294-0.1652+10.7280
base 10 missing 0, level 6172-105.7933-84.0187-0.0074+361.8194
base 5 missing 2, level 9-252-117.3766-34.6141+0.5583-100.5676

Table 1. The exact sum of mu over the strings and its four region shares on the three grids of Fact 10.1, residue counts 924406, 75314, 168 and 112 at base 10 and 1823896, 129062, 124 and 43 at base 5. The shares are float readings of an exact identity; the minor-arc bound carries an unstated constant, so the grid readings of |S_P| against it bound nothing. Generator: lab/py/mobius-dissection, verb regions.

What the theorem does not say. The saving is exp(-c sqrt(log x)) with a tiny c, never a power; the theorem is silent at every set without a certificate below 1/5, and so at every set of bounded fill; and it says nothing about the true size of M_F(x), for which square-root cancellation against A_F(x) is the natural guess. The dense sets it reaches and the sparse sets where that guess is interesting do not meet, and the next statement says the method cannot make them meet.

Proposition 10.2 (the route needs dense sets). Run on any bound for the one-step constant, region B asks alpha_1 < 1/4 whatever the uniform exponent of the minor-arc input, since that bar comes from the decay max(d, h)^(-1/2) alone, and region A asks alpha_1 < 1/2 even from a square-root uniform bound. Lemma 2.3(iii) gives alpha_1 >= log_q(B/k) >= 1 - alpha at every digit set, so the route needs k > q^(3/4). At F = {0, 1} in base 3, B >= 2(q-1) = 4 gives alpha_1 >= log 2/log 3 = 0.630929, above both bars.

Proof. In Proposition 5.2 the classes with max(D, H) near y^(2/5) weigh k^n y^(-1) x max(D, H)^(2 alpha_1 - 1/2) up to logarithms, a saving only when 2 alpha_1 < 1/2, whatever exponent the uniform term carries. In Proposition 4.4 the mass k^n y^(alpha_1) meets y^(b - 1) for a uniform exponent b, and no uniform exponent below 1/2 exists, since the mean square of S_mu(X, theta) over theta is sum_(v <= X) mu(v)^2 >> X; so the product saves only when alpha_1 < 1 - b <= 1/2. The inequalities are Lemma 2.3(iii) and log_q(B/k) <= alpha_1. □

Open problems

The wall is set by the certificate. Region A reads the transform only through the certificate at shift 0 and Lemma 5.1 through the certificate at every shift, so a sharper certificate moves the wall and nothing else changes. The chain of Proposition 8.5 runs on a majorant that forgets the avoided digit, and the per-digit certificates of Fact 8.7 bring the verified reach to 115, the floor of this route at one avoided digit, with a single digit certifying from base 65. A proof below 584, digit-uniform or digit by digit, is not written. The other lever is the exponent 4/5 of the minor-arc input on y^(2/5) <= d <= y^(3/5), which neither this paper nor Maynard (2022) moves. The floor 33 of Proposition 8.8 binds only one-step constants, and below 115 the theorem holds only set by set, where a digit certifies. The shape exp(-c sqrt(log x)) is not examined for improvement, the size of c is not optimised, and the sparse sets of bounded fill, where Proposition 10.2 shows this route dead at every strength of its inputs, have no unconditional Mertens bound in any source read here.

Reproducibility

Five studies print every number of Sections 2, 5, 8, 9 and 10, each run from the repository root with one verb and raising if any check fails. uv run python research/lab/py/prime-dissection/primes.py wall in 2 seconds prints Proposition 8.5's certificate and cap and Maynard's crossing 1520573, and window in 21 seconds prints Fact 8.6; uv run python research/lab/py/digit-transform-norms/norms.py fifth in five minutes and fifthbelow in 23 seconds print Fact 8.7. uv run python research/lab/py/mobius-dissection/dissection.py wall in under a second prints Fact 8.1 and the mass readings after Proposition 8.8; blocks in 18 seconds prints the split of Fact 10.1; regions in five seconds prints the rest of Fact 10.1 and Table 1. uv run python research/lab/py/mrly-pairing/pairing.py onestep in 29 seconds prints the threshold Psi'_q >= (1 + pi) q/2 from q = 36 used in Lemma 2.6 and the wall 1499 of Section 5. CARGO_BUILD_JOBS=4 cargo run --release -p mertens-numerology, in milliseconds, prints the floor 11221 of Lemma 8.2 and the budget 176 of Fact 8.4. The walls are float scans whose least gap is re-read at 40 digits in mpmath, and 1/5 - alpha_1 is printed from that gap through log(1 + gap/P_W)/log q, never by differencing two numbers of size 1. The figure is bash scripts/figures.sh paper-unconditional-mertens-at-large-base, under half a second a theme, its binary asserting the four region counts 732, 202, 26 and 40 of the 1000 frequencies.

References

  • Basak, Robles and Zaharescu 2023, Exponential sums over Mobius convolutions with applications to partitions. arxiv.org/abs/2312.17435
  • Koukoulopoulos 2019, The Distribution of Prime Numbers, Graduate Studies in Mathematics 203, American Mathematical Society, read in the author's preliminary version. dms.umontreal.ca
  • Chang and Martin 2019, The smallest invariant factor of the multiplicative group. arxiv.org/abs/1908.00035
  • Maynard 2022, Primes and polynomials with restricted digits, Int. Math. Res. Not. 2022, 10626-10648. doi.org/10.1093/imrn/rnab002
  • Maynard 2019, Primes with restricted digits, Invent. Math. 217, 127-218. link.springer.com
  • Nath 2024, Primes with a missing digit: distribution in arithmetic progressions and an application in sieve theory, J. London Math. Soc. 109, e12837. arxiv.org/abs/2108.09212
  • Leng and Sawhney 2025, Vinogradov's theorem for primes with restricted digits. arxiv.org/abs/2409.06894
  • Erdos, Mauduit and Sarkozy 1998, On arithmetic properties of integers with missing digits I, J. Number Theory 70, 99-120. doi.org/10.1006/jnth.1998.2229
  • Kim 2024, The divisor function over integers with a missing digit. arxiv.org/abs/2411.09076