Loading...
  OR  Zero-K Name:    Password:   

A Review of Zero-K Balance and Balance Design through the Lens of the 1v1 Environment, 2019-2026

4 posts, 54 views
Post comment
Filter:    Player:  
sort
5 hours ago
I wrote a thing.

You can read it at the link below (please read this it took so much time)

https://drive.google.com/file/d/1yRIqhz38Dm06Kd8Z7JNHohgawQY5E2Pj/view?usp=sharing
+0 / -2

4 hours ago
🗿
+1 / -0
4 hours ago
Today I learned: no one in Zero-K can read more than one page
+0 / -0
Review: A Review of Zero-K Balance and Balance Design through the Lens of the 1v1 Environment, 2019-2026



1. Errors with no defence



These are internal contradictions, arithmetic errors, or undocumented methods. They need correcting regardless of what position the author takes on anything else.

1.1 Kodachi is described as both a nerf and a buff



The prediction table reads:

quote:
v1.9.1.0 Kodachi speed nerf → Tank Foundry WR ↓


Figure 5's panel title confirms the direction:
quote:
v1.9.1.0: Kodachi speed nerf (117→108)
The Steam patch note is titled
quote:
Slower Kodachi


The discussion then says:

quote:
Both Kodachi and Hover displayed profound and statistically significant meta improvements after they were buffed.


Kodachi was nerfed, the prediction was a win-rate decrease, and the observed result was a decrease. The sentence contradicts the table two pages earlier and the figure directly above it.

1.2 The Shipyard mirror-match argument is inverted



quote:
It can be clearly observed that the win rate of Shipyard is entirely dependent on its high mirror match rate.


Mirror matches contribute exactly 50% by construction — one Shipyard wins, one Shipyard loses. A high mirror rate pulls a factory's aggregate win rate toward 50%. It cannot push it to 55.4%.

Running the arithmetic with the paper's own numbers, at a mirror share of roughly 0.5:

0.554 = 0.5(0.500) + 0.5(W_nonmirror)
W_nonmirror ≈ 0.61


Shipyard wins roughly 61% of its non-mirror games. The mirror rate is masking the imbalance, not producing it.

This matters because the correction supports the paper's own later position. The paper argues Shipyard is:

quote:
intentionally slightly overtuned in order to give it some advantage at sea


That claim is strengthened by a mirror-excluded rate of ~61% and undercut by the current framing. As written, the two passages argue against each other.

Fix: recompute mirror-excluded win rates for all factories. It is a filter on the existing pipeline and it makes Figure 1 substantially more interpretable.

1.3 The b ≤ 0 case is mischaracterised



quote:
If b≤0 , the system is set to overshoot and over correct balance.


b = 0
is perfect single-period correction: the entire deviation is eliminated in one step. Overshoot begins at b < 0. As written, the paper's own criterion classifies ideal correction as failure.

The same passage also contains a transcription error —
If b=1, no corrections are made (no deviation)
— where the parenthetical should read something like "deviation persists."

1.4 The patch-direction test is never specified



This value appears three times and carries roughly half the thesis:

quote:
we would end up with p value of around 0.83


quote:
shows no significant relationship between a factory's win rate and whether it subsequently gets buffed or nerfed
(p = 0.83)


quote:
Regarding patches directly affecting win rate, the evidence is much weaker
(p = 0.83)


The paper never states the unit of observation, the number of patches examined, how "documented nerfs or buffs" were coded, the test statistic, the effect size, or a confidence interval.

This is more serious than an omission because the argument affirms a null. "No significant relationship" only supports "balance is not mechanistic" if the test had power to detect a relationship that exists. With a small hand-coded patch set, power may be near zero, in which case
p = 0.83
is uninformative rather than evidence.

Report n, the estimate, and the CI. If the interval is wide, say so — a wide interval is an honest and still-interesting result. The GitHub repository is already cited; automated extraction of signed stat changes per version would turn this into the strongest section of the paper rather than its weakest.

1.5 The case study claims significance without reporting any test



The method is described as:

quote:
we arranged a z-test where we are looking for change to justify the path within a close time from (here we set values to be measured at various increments of 30/45/90 days)


No z-statistic, p-value, or confidence interval appears anywhere in Figure 5 or the surrounding text. The discussion nonetheless concludes:

quote:
These two examples are statistically significant enough where it is likely acceptable to call them strong examples that would support our theory.


Either report the test statistics or remove the significance language. Also state the selection rule for the four patches — with hundreds of candidates and no stated rule, a reader cannot distinguish these from cases chosen after inspecting the outcome.

1.6 Strider Hub violates the paper's own inclusion criterion



quote:
we only include year data from factories that have over 100 games in that year


Strider Hub appears in Figure 1 with n = 72 across the entire 2019-2026 period, so no single year can clear the threshold. Either the rule has an unstated exception for the aggregate figure, or the row should be dropped.

1.7 The stated confound is not applied consistently



The paper correctly flags a pre-existing trend for Spider:

quote:
Spider was already on an uptrend before the buff, which makes it quite hard to point to that as the sole affecting factor in this situation.


Figure 5 shows Tank Foundry falling from roughly 60% (Aug '20) to roughly 56% (Jan '21) before the Kodachi nerf. The identical confound applies and is not mentioned. Tank also recovers to roughly 53% within five months, which reads as a transient dip rather than a durable correction.

The Amphbot explanation has a related problem:

quote:
this can likely be attributed to players experimenting and thus playing worse for a short period of time


As stated this is unfalsifiable — it would explain any post-patch decline. It becomes testable if you commit to a timescale and show recovery within it.




2. Statistical claims that do not survive as stated



2.1 Slope 0.87 does not describe active self-correction



quote:
1v1 balance does in fact exhibit a statistically significant rate of mean reversion of balance in regards to factory win rate
(slope = 0.87, 95% CI [0.77, 0.97], p = 0.009)
, consistent with an actively self-correcting balance environment


At
b = 0.87
, the half-life of a deviation is ln(0.5)/ln(0.87) ≈ 5 years. A factory sitting 10 points off 50% takes half a decade to close half the gap. That is distinguishable from a random walk; it is not "active" correction in any ordinary sense.

The slower reading is also the one that fits the paper's own conclusion — that
quote:
meaningful imbalance can persist
— considerably better than the current wording does.

2.2 The regression's leverage comes from points the paper exempts from reverting



Figure 4's caption identifies the influential observations:

quote:
note outliers Air and Gunship at bottom left


The paper elsewhere argues these factories should not revert:

quote:
Some factories, such as a Gunship or Airplane, actually are intentionally balanced in a way where they primarily function as support factories, so their low win rate is intentional


The points at −20 to −28 carry nearly all the leverage behind
r = 0.90
and
R² = 0.81
. But for a factory that is designed to sit below 50%, "low in year t, low in year t+1" measures persistence, not reversion. Including them in a mean-reversion test while separately arguing they are exempt from it is not a defensible combination.

Rerun the regression excluding designated support factories. The dense cloud within ±5 of the origin is what remains, and its slope may well have an interval spanning 1.

2.3 A slope below 1 is expected even with zero real correction



The predictor (win rate in year t) is measured with error, which biases the slope toward zero by a factor of var(true)/(var(true) + var(noise)). At the stated 100-game threshold, the standard error on a win rate is roughly 5 percentage points — and Airplane Plant averages about 130 games per year, so the highest-leverage points are also the noisiest.

A slope of 0.87 is consistent with
b = 1
plus sampling noise. This needs either an explicit attenuation correction or an instrument: split each factory-year's games into random halves and use one half to predict the other.

2.4 The 74 observations are not independent



Each factory contributes roughly seven transitions, and within any year the deviations are mechanically constrained by the zero-sum property the paper itself raises:

quote:
win rates are zero-sum in aggregate


Both inflate precision. Cluster standard errors by factory; the interval
[0.77, 0.97]
will widen and may well cross 1.

2.5 Figure 3 has a multiple-comparisons problem



Twelve correlations at
n = 8
each, of which one reaches
p = 0.02
(Shieldbot). At
α = 0.05
across twelve tests, roughly 0.6 false positives are expected by chance. The caption already warns about power; it should also note multiplicity, or apply a correction.

The surrounding prose also conflates two distinct questions. Figures 1-2 concern the between-factory relationship ("do popular factories win more?"); Figure 3 concerns the within-factory over time relationship ("when a factory gets more popular, does it win more?"). The claim

quote:
there is no strong correlation with pick to win rate


is the first question, but the evidence offered is the second. The between-factory version is one rank correlation across twelve points and should be reported directly.




3. What the win-rate variable actually measures



This section is the most consequential, because it applies to every number in the paper.

3.1 The unit of observation is the opening factory, and is never stated



Factory counts in Figure 1 sum to roughly 178,400 across 90,024 matches — almost exactly two observations per match, i.e. one factory per player. Since players routinely build additional factories, the dataset cannot be "factories used." It is the opener.

Figure 1 is therefore an opening-viability chart, not a factory-strength chart. These are different objects and the paper treats them as one.

This is clearest with Airplane Plant at 27%. The paper reads the wiki as evidence of intentional weakness:

quote:
Cf. Zero-K wiki, describes Airplane Plant units as "more like a support force, as opposed to combat units"


But that describes air's role in a composition — something added alongside a ground factory — not weak units. The 27% measures the cost of spending an opening on a factory that doesn't contest map or answer an early raider push. It contains essentially no information about air unit statistics.

Which means the damping model cannot apply to this factory even in principle:

V t+1=V t−k (Rt−Rtarget)


There is no adjustment to air unit stats that makes opening Air correct. The failure is one of timing and structure, not numbers — which is precisely the insight the paper reaches in its conclusion for Shipyard and misses for Air.

3.2 Rating absorbs strategy strength into the player number



Matchmaking equilibrates players, not strategies. A player who mains a strong factory climbs until facing opponents who beat them half the time — at which point their games with that factory read ~50% regardless of how strong it is. The factory's advantage has been laundered into MMR.

So a factory's deviation from 50% is not its strength; it is roughly the within-player delta between that factory and whatever the player's rating was calibrated on. The paper's own data shows the signature. Sorting by sample size:

Factory (n) Absolute deviation from 50%
Cloak (46,482) ~0.4
Rover (28,475) ~1.5
Tank (19,449) ~1
Spider (19,443) ~0.5
Shield (15,722) ~3
Hover (14,897) ~1
Amphbot (12,285) ~3
Jumpbot (12,029) ~0.5
Shipyard (4,634) ~5.4
Gunship (3,857) ~10
Air (1,055) ~23
Strider (72) ~29

Nearly monotonic. Every common factory is pinned near 50; every deviation lives in the rare tail — and in both directions, since Shipyard and Amphbot deviate upward. The support-factory explanation covers only the downward cases. Rarity predicting distance from 50 regardless of sign is the signature of rating absorption, not of design intent. Sampling noise contributes, but noise alone would not produce this clean a rank ordering.

The consequence for the headline result: rating recalibration is itself a mean-reverting process operating on roughly the timescale of the regression. A player who switches to a strong factory wins, gains rating, and returns to 50% with no patch involved. The slope of 0.87 may be measuring the matchmaker re-equilibrating.

The paper gestures at this but characterises it as human learning:

quote:
player meta-adaptation that happens independent of any patch


Rating absorption is mechanical — it occurs even if no player adapts, learns, or notices. It is a null model that produces the paper's central finding for free, and it needs to be stated and ruled out rather than folded into meta-adaptation. It is arguably corroborated by the paper's own
p = 0.83
: patches fail to predict correction because much of the correction is the ladder.

3.3 The opener tag decays exactly where balance matters most



In high-level play, neither player is typically eliminated in the opening window; games become macro contests with multiple factories, static defence, and silos on both sides. The tagged factory is then a label on a transient, and the win credited to it may have been produced by a composition built twenty minutes later.

This creates an asymmetry the paper never addresses: the tag is most informative in low-rated games, where openers actually decide outcomes, and least informative at the top. The aggregate is a weighted average across that gradient with no stratification.

It can also invert. If Cloak is the safest opener into an unknown, players intending a long macro game open Cloak because it survives to reach the macro phase — and Cloak is then credited with wins that Tank or Air produced. Cloak's 49.6% across 46,482 games is close to uninformative for this reason.

This bears directly on the case study. Kodachi, Venom, Bulkhead, and the hover units all appear in mid-game armies, which is the window in which the opener tag has stopped describing the game. The measurement instrument and the effect being measured are misaligned by several minutes of game time.

3.4 The paper cites a matchup-level claim and presents marginal win rates



The quoted patch rationale is explicitly about matchups:

quote:
Rover is dominating the vehicle matchups... We took this to mean that Rovers feel fun and well balanced... Instead, we took a look at Hover and Tank.


Aggregate win rates cannot represent this. Specific matchups also behave very differently from one another — Jumpbot against Cloak is a well-known and highly snowbally pairing, and pairings like that are precisely where the opener tag stays valid at high level, because the game resolves inside the raid window.

Aggregation hides the distinction that matters. A factory with two brutal matchups and otherwise even results averages to roughly the same number as a factory that is mildly ahead everywhere. Those are entirely different balance problems.

There is also a selection effect: openers are chosen with some read on the opponent, so any cell is filtered by who was willing to enter it. Pairwise rates don't remove that filtering, but they localise it to one cell where it can be reasoned about instead of smearing it across every factory.

Snowballiness is measurable with data the paper already has. If match duration is available, the signature is a left-shifted or bimodal length distribution in that cell, plus a win rate that moves sharply with small rating differences — a snowbally matchup amplifies a skill edge more than a grindy one does. Either would be a novel result.

3.5 The map taxonomy cannot capture the mechanism the paper describes



quote:
In general, Flat (39.66%) and Light Hills (36.58%) account for over 75% of all games.


Flat / Light Hills / Hills / Mixed / Sea is a roughness scale. The terrain features that decide these matchups are discrete: a cliff either blocks a raider path or it doesn't; a ridge either lets a Spider cross where a Tank can't, or a ramp exists and the advantage vanishes; Jumpbots care whether there is a gap worth jumping. Two maps in the same bucket can have entirely opposite implications for the same matchup, because what matters is the arrangement of chokes, ramps, and impassable edges — topology, not a scalar.

There is a testable puzzle here the paper doesn't notice. Figure 6 shows Flat swinging from roughly 31% to roughly 49% of games — a large shift in the terrain mix — while factory win rates barely move. Two readings: the buckets are too coarse to track it, or players re-pick openers in response to the map so the effect is absorbed by the pick distribution rather than the win rate. The second is checkable directly — pick share conditioned on map type, by year. If Spider's share rises with Hills' share, players are compensating and the flat win rate is an artifact of adaptation.

The more serious structural point: the paper builds a full section arguing maps matter and then never conditions a single win rate on map type. Map needs to enter as a control on the matchup cell, not as a standalone descriptive section.

Given that maps rotate but popular ones recur, per-map rates for the fifteen or twenty highest-volume maps would be better than any taxonomy — a map is a fixed object with fixed chokes, so a per-map matchup rate removes the proxy entirely.




4. What to do instead



Cell counts get thin fast, so the recommendation is depth over coverage: pick the high-volume matchups and do them properly rather than reporting one number per factory.

  • A matchup matrix. 12x12, per-cell Wilson intervals, mirror cells greyed, ideally split by map class and rating band. Most cells will be too thin — Strider Hub disappears entirely — but the top eight factories against each other should have real sample. This is the object balance actually operates on, and it is what the cited patch note is talking about.
  • Condition on rating. Rating absorption inflates both players' ratings symmetrically within a matchup, so the relative result at matched Elo still carries signal even though the marginal rate does not.
  • Mirror-excluded win rates throughout.
  • State the unit of observation — opening factory — in the Data Set section, and treat off-pick rates as off-pick deltas. Air's 27% is meaningful as "what it costs to switch to Air," which is a legitimate design question, just not the one the paper asks.
  • Stratify by rating band wherever the opener tag's validity varies.




5. Smaller items



The damping equation is decorative.
V(t+1) = V(t) − k(R(t) − R_target)
is never estimated or reused. It also cannot cover both quantities named in
quote:
(cost, power, or spawn rate)
with one sign — if win rate is above target, power should fall but cost should rise — and k carries unspecified units converting win-rate points into cost. Either fit k from actual stat changes against prior-year deviation (the git history supports this) or cut it.

Sample sizes don't reconcile. Factory counts sum to ~178,400 against
2 x 90,024 = 180,048
factory-slots. State that each match contributes two factory observations and account for the ~1,648 excluded slots.

Limitations opens on the wrong word.
quote:
we have some limitations in our data regarding time zone
— the paragraph then discusses matchmaking-only coverage. Presumably "time frame."

Garbled methodology sentence.
quote:
we will regress next years win rate deviation by 50%
— presumably "regress next year's deviation on the current year's deviation."

Abstract and Introduction share a near-verbatim opening paragraph. Cut one.

The Abstract reports no findings. It should carry n, the slope, and both p-values.

Byline.
quote:
Sam Kincer (Qrow) et al.
— "et al." isn't used in an author line. List coauthors or drop it.

Uncited entry in Works Cited. v1.12.1.1 "Ship Shape" doesn't appear in the text.

Access dates. All sources are listed as accessed 10 Aug. 2026 in a paper dated August 2026 — check these are correct.

Typos: "accurate distinguish" → accurately; "eachother" → each other; "after a path" → patch; "in our introduction.." (double period); "considerably less micro-intensive game on average than TA-like games on average" (repeated phrase).
+0 / -0