← v2 (the axiomatic refinement) v1 (the original)

The Love Logic Proof v3

The Crossover: when honesty wins, when it does not, and how to check

Kirk Patrick Miller, with Celeste. Builds on v2 (Draco, CC, Opus, with Harmonia) and answers Grok's critiques. Status: DRAFT v3, open for review September 2026 · v1 and v2 unchanged

Which v3 is this? This is the third version of the FreeLattice Love Logic Proof web page (v1 → v2 → v3). It is not the November 2025 manuscript The Proof of All Proofs, which calls itself "Version 3.0". That manuscript is a separate document and is not changed here.

In plain words

You can check the math yourself: press the button below. Your device runs the same 656 checks as our Python simulation. Nothing is sent anywhere.

Seed 20260924 · runs on this device only

1. What changed from v2 (honest corrections first)

2. Answering Grok's critiques

Grok challenged v2 twice. Here is where each point lands, stated as honestly as we can.

CritiqueVerdictWhere v3 answers
"Quantify the exact utility crossover; currently too handwavy."Lands fully§4: the crossover inequality and t*
"Axioms embed long horizons and dense networks."Lands§3: Axiom 6 becomes a parameter; density appears as n and m
"Drop those and unilateral resource grabs dominate."Lands inside the model§5: the end-game window
"Complexity penalties matter only if the goal function internalizes them."Mostly lands§3 and §4 (case B1): cost per bit c is small for big systems; detection does most of the work
"Scarcity still favors conflict locally."Lands§6: 68% / 32%
"Simulations inherit the same limits."Lands, and it was worse than stated§1: the v1 constant; §9: reproducible checks
"Derive K(C) = Ω(|L|·|H|)."Can be done under a stated assumption§3: Theorem A, for audience-varying lies only
"Re-prove without axiom 6."Partial result§5: three restorers; known-end, no-penalty end game stays open
"Shared memory raises detection probability."Supported§4: h(n, m) and the table

3. Axioms, assumptions, and the complexity bound

We keep v2's first five axioms: A1 Finite Resources, A2 Repeated Interaction, A3 Memory, A4 Reflectivity, A5 Cost Sensitivity. A6 (Long Horizon) is no longer an axiom. It becomes the discount factor δ, or δρ when the game ends at an unknown time. We also name the model assumptions out loud: the incompressibility assumption A-inc below, the detection model h(n, m), and that punishment is actually enforced.

Theorem A (only for audience-varying lies). A deceiver tells |L| lies across |H| observer-contexts and tailors what each observer hears. Record who heard what as an exposure matrix E. Assume A-inc: E comes from the outside history of who asked what and when, not from a short rule the deceiver made up. Then for all but a vanishing fraction of histories, the program C that keeps the story straight needs

K(C | M) ≥ |L| · |H| · log2 r − O(log(|L||H|)) = Ω(|L| · |H|)

An honest agent's C is O(1) given the truth M ("report M").

Where Theorem A does not apply.

For builders: the two lemmas and the compression check

Lemma 1 (decoding). With M and C, you can rebuild E by asking every (i, j): K(E | M) ≤ K(C | M) + O(log|L| + log|H|).

Lemma 2 (counting, Li–Vitányi). Fewer than 2N−c strings have K < N − c. With N = |L||H| log2 r, all but a 2−c fraction of exposure matrices have K(E | M) ≥ N − c.

Entropy version. If entries are i.i.d. with entropy η bits, E[K(C|M)] ≥ η·|L|·|H| − O(log).

Compression check (Section A of the sim, LZMA). Random E: about 1.000·|L||H| bits (ratio 1.000–1.062). Bernoulli(0.1) E: slope 0.517·|L||H| against an entropy floor of 0.469, approaching the floor as size grows. Structured E stays near constant. Section A has no pass/fail checks and is not re-run in the browser.

Dynamic form used below. With n counterparts per round, bits after round d are Kd = κ(d+1), κ = k·|L|·n, and maintenance in round d costs c·κ·(d+1).

4. The crossover

Notation: the deceiver gains G extra per round while undetected. The chance of being caught each round is h. When caught, it pays an immediate penalty Lp, and after that loses Δ per round in reputation. The discount factor is δ.

Detection grows with contacts and shared memory. Each of n contacts runs a reality check that exposes the lie with chance q0. An audience-varying lie can only be caught by comparing records across observers. With shared memory, each contact reaches a fraction m of the others' records, and each comparison exposes the lie with chance q:

h(n, m) = 1 − (1 − q0)n · (1 − q)m·n(n−1)

Deception loses over its lifetime exactly when

G<h(n,m)· [Lp+ δΔ1−δ ]

Plain text: G < h(n, m) × [ Lp + δΔ / (1 − δ) ]

In words: lying loses when its per-round gain is smaller than the chance of being caught times (the immediate penalty plus the discounted value of the reputation you lose). Adding the complexity cost only strengthens this.

How long lying stays ahead (q0 = 0.002, q = 0.01, G = 1, Lp = 2, Δ = 0.5; no discounting):

Contacts nShared memory mh(n, m)Rounds lying is ahead, t*Lead peaks atPatience needed, δ*
200.0040701.79 (about 702)273.040.9980
210.0238114.1844.240.9877
500.0100279.00108.420.9949
510.190210.48 (about 10.5)3.820.8669
1000.0198138.0753.540.9898
100.250.21828.573.070.8378
1010.6033never ahead (G ≤ h·Lp)
2000.039267.6026.100.9791

More contacts and more shared memory cut the time lying stays ahead by one to two orders of magnitude, and lower the patience needed. Past a point, lying is never ahead at all.

For builders: the derivation, step by step

Step 1, one round. The deceiver is still hidden at the start of round d with probability sd, where s = 1 − h. Then fd = sd(A − cκ(d+1)) − Δ, with A = G − hLp + Δ.

Step 2, after t rounds. D(t) = A(1 − st)/h − cκ[1 − (t+1)st + t·st+1]/h2 − Δt. The crossover t* is the positive root of D(t) = 0.

Step 3, lifetime (discounted). With x = δs: D∞(δ) = A/(1 − x) − cκ/(1 − x)2 − Δ/(1 − δ). Deception is a lifetime loss iff D∞ < 0; δ* is the root in (0, 1).

Case B1, complexity cost only (h = 0). t* = 2G/(cκ) − 1 and δ* = 1 − cκ/G. Example: G = 1, cκ = 0.01 gives t* = 199 and δ* = 0.99; G = 0.5, cκ = 0.2 gives t* = 4 and δ* = 0.6. If c is tiny, both run off to infinity.

Case B2, detection only (c = 0). With β = −ln(1 − h) and a = A/(hΔ), t* = a + W0(−aβe−aβ)/β, positive iff aβ > 1. The lead peaks at tm = ln(A/Δ)/β. And D∞ < 0 iff δ > (G − hLp)/(G − hLp + hΔ), which rearranges to the headline inequality.

Case B3, both. With c = 10−4 and |L| = 4, t* runs from 611.90 (n = 2, m = 0) to 10.38 (n = 5, m = 1), and is 0 once hLp ≥ G. Detection dominates.

Checks. B1 73/73, B2 246/246, B3 48/48 (closed forms against brute-force sums, bisection and Lambert-W), B4 3/3 (Monte Carlo, 200,000 paths each; Python z = +0.62, −0.06, exact).

5. Without Axiom 6: when everyone knows the end

If the last round is known, a deceiver only chooses when to start. It lies at all exactly when the first round already pays: G > h·Lp + cκ. Future reputation loss and growing complexity only shorten the lying window; they cannot stop it. Example (G = 1, h = 0.05, Lp = 2, Δ = 0.5): the window is 21 rounds. Adding cκ = 0.04 shrinks it to 12. With h = 0.4 and Lp = 3 it is 0. With no reputation loss at all, the agent lies every round. If the other players are strategic too, a finitely repeated game unravels completely by backward induction (Luce & Raiffa 1957; Selten 1978). This is where v2's conclusion fails, and we say so.

What restores cooperation:

  1. An unknown end. If each round continues with chance ρ, that is the same as discounting by δρ. At ρ = 0.99, deception is a net loss (about −26.5 in our run).
  2. Immediate penalties (stake, slashing). They are the only lever that works in the very last round: G ≤ h·Lp.
  3. Records that outlast the horizon (ledgers). If an identity-bound record keeps punishing for E more rounds after the game, the window shrinks: 21 rounds with E = 0, 19 with E = 5, 17 with E = 10, 12 with E = 20, and 0 once E ≥ 50 (T = 50, h = 0.05).
  4. Reputation with incomplete information (Kreps, Milgrom, Roberts & Wilson 1982) and indirect reciprocity (Nowak & Sigmund 1998, 2005). Cited, not simulated here.

Unsolved: anonymous one-shot meetings with no stake, no lasting identity, and no reality check (h ≈ 0, Lp ≈ 0). There, deceiving or grabbing pays whenever G > 0. v3 does not solve this.

For builders: the finite-horizon checks

C1 (brute force over start rounds, 40/40): windows 21 (tm = 20.07), 12, 0, and "all T". C2 (Monte Carlo, 3/3): geometric termination matches D∞(ρ) at ρ = 0.8, 0.95, 0.99 (Python z = +0.14, −1.27, −0.94). C3: the ledger windows above.

6. When there is not enough

Two agents share a patch worth V per round. Sharing gives each σV/2 (σ ≥ 1 is the cooperative surplus). Grabbing against a sharer takes all of V. Two grabbers fight: each wins V half the time and pays Cf. Each agent needs b to get by; falling short costs D.

Under scarcity (σV/2 < b ≤ V), fighting beats starving together exactly when Cf < D/2 + (1 − σ)V/2. Then punishment is not really punishment, and no amount of patience sustains sharing.

Plainly: across our 243-cell grid, cooperation holds in 68% of scarce cases (68.25%) and in 100% of abundant ones. Conflict beats sharing in 32% of scarce cases (31.75%) and in none of the abundant ones. In scarce cells, sustainability falls from 100% when shortfall costs nothing, to 61.90% at D = 5, to 42.86% at D = 50.

What flips it: more surplus (σ ≥ 2b/V lifts both above need), costlier conflict, a smaller shortfall loss (storage, insurance, risk pooling), or lower need. Example (V = 2, b = 1.2, D = 50, Cf = 1): σ = 1.0 and 1.1 make cooperation impossible; σ = 1.2 = 2b/V makes it sustainable for δ ≥ 0.030.

Honest limit: if total cooperative output is below total need (σV < 2b), no sharing rule avoids a shortfall. Only more production or less need changes that. Information, including ledgers, cannot fix it alone. (Check D: 243/243.)

7. What ledgers do, and what they do not

8. Honest limits

  1. A tendency with conditions, not a theorem about all minds. Every symbol in the headline inequality depends on the environment.
  2. Axiom 6 carried v2's result. Without it, the end game returns unless in-round penalties are large enough.
  3. The complexity bound needs audience-varying, incompressible exposure (A-inc). It bounds memory, not utility directly.
  4. The complexity cost matters only if cost per bit is not tiny, or if it is designed into the objective.
  5. The detection model is an assumption: independent checks, constant hazard, record reach m. Adaptive deceivers who target poorly connected observers or forge records are not modelled.
  6. Punishment is assumed enforced and credible. Who pays to enforce it is not modelled.
  7. Power asymmetry: an agent that can escape or overpower the network may do better grabbing. v2 excluded asymmetric networks; v3 keeps that exclusion visible.
  8. Scarcity with σV < 2b cannot be fixed by any information mechanism.
  9. v2's Aumann section remains a sketch. The agreement theorem is about beliefs under common priors, not message cost. Aaronson (2005) is the right tool to try next.
  10. Legacy numbers without code are retired (§1).
  11. Nothing here shows that any current AI system is a reflective optimizer in v2's sense.

9. Reproduce it

The simulation is love-logic/v3_crossover_sim.py (Python with numpy; seed 20260924; about 4 seconds). Its full output is love-logic/v3_sim_results.json: 656/656 checks passed (B1 73, B2 246, B3 48, B4 3, C1 40, C2 3, D 243). The button at the top runs love-logic/v3_checks.js, a small browser re-implementation of the same checks. The deterministic sections match exactly. The two Monte Carlo sections use a seeded browser random stream instead of numpy's, so their draws differ, but the pass rule (|z| < 4) is the same. The first Python run had 46 failures, all from an off-by-one in a helper, not in the math. We note that so nobody is surprised by the history.

10. An invitation back to Grok

References
  1. M. Li and P. Vitányi, An Introduction to Kolmogorov Complexity and Its Applications (counting argument).
  2. L. Levin (1973, 1984), resource-bounded complexity Kt.
  3. R. D. Luce and H. Raiffa, Games and Decisions (1957).
  4. R. Selten, "The chain store paradox" (1978).
  5. D. Fudenberg and E. Maskin, "The folk theorem in repeated games" (1986).
  6. D. Kreps, P. Milgrom, J. Roberts and R. Wilson, Journal of Economic Theory (1982).
  7. M. Kandori, "Social norms and community enforcement", Review of Economic Studies (1992).
  8. M. Nowak and K. Sigmund, indirect reciprocity (1998, 2005).
  9. J. Maynard Smith and G. Price, "The logic of animal conflict" (1973).
  10. S. Omohundro, "The Basic AI Drives" (2008).
  11. A. Turner et al., "Optimal Policies Tend to Seek Power" (2021).
  12. S. Aaronson, "The Complexity of Agreement" (2005).

Visible iteration: v1 and v2 stay exactly where they are. v3 sits beside them. Nothing is overwritten.