| 1 |
[size=5][b]Review: A Review of Zero-K Balance and Balance Design through the Lens of the 1v1 Environment, 2019-2026[/b][/size]
|
1 |
[size=5][b]Review: A Review of Zero-K Balance and Balance Design through the Lens of the 1v1 Environment, 2019-2026[/b][/size]
|
| 2 |
|
2 |
|
| 3 |
<wiki:toc />
|
3 |
<wiki:toc />
|
| 4 |
|
4 |
|
| 5 |
= 1. Errors with no defence =
|
5 |
= 1. Errors with no defence =
|
| 6 |
|
6 |
|
| 7 |
These are internal contradictions, arithmetic errors, or undocumented methods. They need correcting regardless of what position the author takes on anything else.
|
7 |
These are internal contradictions, arithmetic errors, or undocumented methods. They need correcting regardless of what position the author takes on anything else.
|
| 8 |
|
8 |
|
| 9 |
== 1.1 Kodachi is described as both a nerf and a buff ==
|
9 |
== 1.1 Kodachi is described as both a nerf and a buff ==
|
| 10 |
|
10 |
|
| 11 |
The prediction table reads:
|
11 |
The prediction table reads:
|
| 12 |
|
12 |
|
| 13 |
[quote]v1.9.1.0 Kodachi speed nerf → Tank Foundry WR ↓[/quote]
|
13 |
[quote]v1.9.1.0 Kodachi speed nerf → Tank Foundry WR ↓[/quote]
|
| 14 |
|
14 |
|
| 15 |
Figure 5's panel title confirms the direction: [q]v1.9.1.0: Kodachi speed nerf (117→108)[/q] The Steam patch note is titled [q]Slower Kodachi[/q]
|
15 |
Figure 5's panel title confirms the direction: [q]v1.9.1.0: Kodachi speed nerf (117→108)[/q] The Steam patch note is titled [q]Slower Kodachi[/q]
|
| 16 |
|
16 |
|
| 17 |
The discussion then says:
|
17 |
The discussion then says:
|
| 18 |
|
18 |
|
| 19 |
[quote]Both Kodachi and Hover displayed profound and statistically significant meta improvements after they were buffed.[/quote]
|
19 |
[quote]Both Kodachi and Hover displayed profound and statistically significant meta improvements after they were buffed.[/quote]
|
| 20 |
|
20 |
|
| 21 |
Kodachi was nerfed, the prediction was a win-rate decrease, and the observed result was a decrease. The sentence contradicts the table two pages earlier and the figure directly above it.
|
21 |
Kodachi was nerfed, the prediction was a win-rate decrease, and the observed result was a decrease. The sentence contradicts the table two pages earlier and the figure directly above it.
|
| 22 |
|
22 |
|
| 23 |
== 1.2 The Shipyard mirror-match argument is inverted ==
|
23 |
== 1.2 The Shipyard mirror-match argument is inverted ==
|
| 24 |
|
24 |
|
| 25 |
[quote]It can be clearly observed that the win rate of Shipyard is entirely dependent on its high mirror match rate.[/quote]
|
25 |
[quote]It can be clearly observed that the win rate of Shipyard is entirely dependent on its high mirror match rate.[/quote]
|
| 26 |
|
26 |
|
| 27 |
Mirror matches contribute exactly 50% by construction — one Shipyard wins, one Shipyard loses. A high mirror rate pulls a factory's aggregate win rate [i]toward[/i] 50%. It cannot push it to 55.4%.
|
27 |
Mirror matches contribute exactly 50% by construction — one Shipyard wins, one Shipyard loses. A high mirror rate pulls a factory's aggregate win rate [i]toward[/i] 50%. It cannot push it to 55.4%.
|
| 28 |
|
28 |
|
| 29 |
Running the arithmetic with the paper's own numbers, at a mirror share of roughly 0.5:
|
29 |
Running the arithmetic with the paper's own numbers, at a mirror share of roughly 0.5:
|
| 30 |
|
30 |
|
| 31 |
{{{
|
31 |
{{{
|
| 32 |
0.554 = 0.5(0.500) + 0.5(W_nonmirror)
|
32 |
0.554 = 0.5(0.500) + 0.5(W_nonmirror)
|
| 33 |
W_nonmirror ≈ 0.61
|
33 |
W_nonmirror ≈ 0.61
|
| 34 |
}}}
|
34 |
}}}
|
| 35 |
|
35 |
|
| 36 |
Shipyard wins roughly 61% of its non-mirror games. The mirror rate is [i]masking[/i] the imbalance, not producing it.
|
36 |
Shipyard wins roughly 61% of its non-mirror games. The mirror rate is [i]masking[/i] the imbalance, not producing it.
|
| 37 |
|
37 |
|
| 38 |
This matters because the correction supports the paper's own later position. The paper argues Shipyard is:
|
38 |
This matters because the correction supports the paper's own later position. The paper argues Shipyard is:
|
| 39 |
|
39 |
|
| 40 |
[quote]intentionally slightly overtuned in order to give it some advantage at sea[/quote]
|
40 |
[quote]intentionally slightly overtuned in order to give it some advantage at sea[/quote]
|
| 41 |
|
41 |
|
| 42 |
That claim is [b]strengthened[/b] by a mirror-excluded rate of ~61% and undercut by the current framing. As written, the two passages argue against each other.
|
42 |
That claim is [b]strengthened[/b] by a mirror-excluded rate of ~61% and undercut by the current framing. As written, the two passages argue against each other.
|
| 43 |
|
43 |
|
| 44 |
[b]Fix:[/b] recompute mirror-excluded win rates for all factories. It is a filter on the existing pipeline and it makes Figure 1 substantially more interpretable.
|
44 |
[b]Fix:[/b] recompute mirror-excluded win rates for all factories. It is a filter on the existing pipeline and it makes Figure 1 substantially more interpretable.
|
| 45 |
|
45 |
|
| 46 |
== 1.3 The b ≤ 0 case is mischaracterised ==
|
46 |
== 1.3 The b ≤ 0 case is mischaracterised ==
|
| 47 |
|
47 |
|
| 48 |
[quote]If b≤0 , the system is set to overshoot and over correct balance.[/quote]
|
48 |
[quote]If b≤0 , the system is set to overshoot and over correct balance.[/quote]
|
| 49 |
|
49 |
|
| 50 |
b
=
0
is
[i]perfect[/i]
single-period
correction:
the
entire
deviation
is
eliminated
in
one
step.
Overshoot
begins
at
b
<
0.
As
written,
the
paper's
own
criterion
classifies
ideal
correction
as
failure.
|
50 |
{
{
{
b
=
0}
}
}
is
[i]perfect[/i]
single-period
correction:
the
entire
deviation
is
eliminated
in
one
step.
Overshoot
begins
at
b
<
0.
As
written,
the
paper's
own
criterion
classifies
ideal
correction
as
failure.
|
| 51 |
|
51 |
|
| 52 |
The
same
passage
also
contains
a
transcription
error
—
[q]If
b=1,
no
corrections
are
made
(
no
deviation)
[/q]
—
where
the
parenthetical
should
read
something
like
"deviation
persists.
"
|
52 |
The
same
passage
also
contains
a
transcription
error
—
{
{
{
If
b=1,
no
corrections
are
made
(
no
deviation)
}
}
}
—
where
the
parenthetical
should
read
something
like
"deviation
persists.
"
|
| 53 |
|
53 |
|
| 54 |
==
1.
4
The
p
=
0.
83
test
is
never
specified
==
|
54 |
==
1.
4
The
patch-direction
test
is
never
specified
==
|
| 55 |
|
55 |
|
| 56 |
This value appears three times and carries roughly half the thesis:
|
56 |
This value appears three times and carries roughly half the thesis:
|
| 57 |
|
57 |
|
| 58 |
[quote]we would end up with p value of around 0.83[/quote]
|
58 |
[quote]we would end up with p value of around 0.83[/quote]
|
| 59 |
|
59 |
|
| 60 |
[quote]shows
no
significant
relationship
between
a
factory's
win
rate
and
whether
it
subsequently
gets
buffed
or
nerfed
(
p
=
0.
83)
[/quote]
|
60 |
[quote]shows
no
significant
relationship
between
a
factory's
win
rate
and
whether
it
subsequently
gets
buffed
or
nerfed
{
{
{
(
p
=
0.
83)
}
}
}
[/quote]
|
| 61 |
|
61 |
|
| 62 |
[quote]Regarding
patches
directly
affecting
win
rate,
the
evidence
is
much
weaker
(
p
=
0.
83)
[/quote]
|
62 |
[quote]Regarding
patches
directly
affecting
win
rate,
the
evidence
is
much
weaker
{
{
{
(
p
=
0.
83)
}
}
}
[/quote]
|
| 63 |
|
63 |
|
| 64 |
The paper never states the unit of observation, the number of patches examined, how "documented nerfs or buffs" were coded, the test statistic, the effect size, or a confidence interval.
|
64 |
The paper never states the unit of observation, the number of patches examined, how "documented nerfs or buffs" were coded, the test statistic, the effect size, or a confidence interval.
|
| 65 |
|
65 |
|
| 66 |
This
is
more
serious
than
an
omission
because
the
argument
[b]affirms
a
null[/b].
"No
significant
relationship"
only
supports
"balance
is
not
mechanistic"
if
the
test
had
power
to
detect
a
relationship
that
exists.
With
a
small
hand-coded
patch
set,
power
may
be
near
zero,
in
which
case
p
=
0.
83
is
uninformative
rather
than
evidence.
|
66 |
This
is
more
serious
than
an
omission
because
the
argument
[b]affirms
a
null[/b].
"No
significant
relationship"
only
supports
"balance
is
not
mechanistic"
if
the
test
had
power
to
detect
a
relationship
that
exists.
With
a
small
hand-coded
patch
set,
power
may
be
near
zero,
in
which
case
{
{
{
p
=
0.
83}
}
}
is
uninformative
rather
than
evidence.
|
| 67 |
|
67 |
|
| 68 |
Report n, the estimate, and the CI. If the interval is wide, say so — a wide interval is an honest and still-interesting result. The GitHub repository is already cited; automated extraction of signed stat changes per version would turn this into the strongest section of the paper rather than its weakest.
|
68 |
Report n, the estimate, and the CI. If the interval is wide, say so — a wide interval is an honest and still-interesting result. The GitHub repository is already cited; automated extraction of signed stat changes per version would turn this into the strongest section of the paper rather than its weakest.
|
| 69 |
|
69 |
|
| 70 |
== 1.5 The case study claims significance without reporting any test ==
|
70 |
== 1.5 The case study claims significance without reporting any test ==
|
| 71 |
|
71 |
|
| 72 |
The method is described as:
|
72 |
The method is described as:
|
| 73 |
|
73 |
|
| 74 |
[quote]we arranged a z-test where we are looking for change to justify the path within a close time from (here we set values to be measured at various increments of 30/45/90 days)[/quote]
|
74 |
[quote]we arranged a z-test where we are looking for change to justify the path within a close time from (here we set values to be measured at various increments of 30/45/90 days)[/quote]
|
| 75 |
|
75 |
|
| 76 |
No z-statistic, p-value, or confidence interval appears anywhere in Figure 5 or the surrounding text. The discussion nonetheless concludes:
|
76 |
No z-statistic, p-value, or confidence interval appears anywhere in Figure 5 or the surrounding text. The discussion nonetheless concludes:
|
| 77 |
|
77 |
|
| 78 |
[quote]These two examples are statistically significant enough where it is likely acceptable to call them strong examples that would support our theory.[/quote]
|
78 |
[quote]These two examples are statistically significant enough where it is likely acceptable to call them strong examples that would support our theory.[/quote]
|
| 79 |
|
79 |
|
| 80 |
Either report the test statistics or remove the significance language. Also state the selection rule for the four patches — with hundreds of candidates and no stated rule, a reader cannot distinguish these from cases chosen after inspecting the outcome.
|
80 |
Either report the test statistics or remove the significance language. Also state the selection rule for the four patches — with hundreds of candidates and no stated rule, a reader cannot distinguish these from cases chosen after inspecting the outcome.
|
| 81 |
|
81 |
|
| 82 |
== 1.6 Strider Hub violates the paper's own inclusion criterion ==
|
82 |
== 1.6 Strider Hub violates the paper's own inclusion criterion ==
|
| 83 |
|
83 |
|
| 84 |
[quote]we only include year data from factories that have over 100 games in that year[/quote]
|
84 |
[quote]we only include year data from factories that have over 100 games in that year[/quote]
|
| 85 |
|
85 |
|
| 86 |
Strider Hub appears in Figure 1 with n = 72 across the entire 2019-2026 period, so no single year can clear the threshold. Either the rule has an unstated exception for the aggregate figure, or the row should be dropped.
|
86 |
Strider Hub appears in Figure 1 with n = 72 across the entire 2019-2026 period, so no single year can clear the threshold. Either the rule has an unstated exception for the aggregate figure, or the row should be dropped.
|
| 87 |
|
87 |
|
| 88 |
== 1.7 The stated confound is not applied consistently ==
|
88 |
== 1.7 The stated confound is not applied consistently ==
|
| 89 |
|
89 |
|
| 90 |
The paper correctly flags a pre-existing trend for Spider:
|
90 |
The paper correctly flags a pre-existing trend for Spider:
|
| 91 |
|
91 |
|
| 92 |
[quote]Spider was already on an uptrend before the buff, which makes it quite hard to point to that as the sole affecting factor in this situation.[/quote]
|
92 |
[quote]Spider was already on an uptrend before the buff, which makes it quite hard to point to that as the sole affecting factor in this situation.[/quote]
|
| 93 |
|
93 |
|
| 94 |
Figure 5 shows Tank Foundry falling from roughly 60% (Aug '20) to roughly 56% (Jan '21) [i]before[/i] the Kodachi nerf. The identical confound applies and is not mentioned. Tank also recovers to roughly 53% within five months, which reads as a transient dip rather than a durable correction.
|
94 |
Figure 5 shows Tank Foundry falling from roughly 60% (Aug '20) to roughly 56% (Jan '21) [i]before[/i] the Kodachi nerf. The identical confound applies and is not mentioned. Tank also recovers to roughly 53% within five months, which reads as a transient dip rather than a durable correction.
|
| 95 |
|
95 |
|
| 96 |
The Amphbot explanation has a related problem:
|
96 |
The Amphbot explanation has a related problem:
|
| 97 |
|
97 |
|
| 98 |
[quote]this can likely be attributed to players experimenting and thus playing worse for a short period of time[/quote]
|
98 |
[quote]this can likely be attributed to players experimenting and thus playing worse for a short period of time[/quote]
|
| 99 |
|
99 |
|
| 100 |
As stated this is unfalsifiable — it would explain any post-patch decline. It becomes testable if you commit to a timescale and show recovery within it.
|
100 |
As stated this is unfalsifiable — it would explain any post-patch decline. It becomes testable if you commit to a timescale and show recovery within it.
|
| 101 |
|
101 |
|
| 102 |
----
|
102 |
----
|
| 103 |
|
103 |
|
| 104 |
= 2. Statistical claims that do not survive as stated =
|
104 |
= 2. Statistical claims that do not survive as stated =
|
| 105 |
|
105 |
|
| 106 |
== 2.1 Slope 0.87 does not describe active self-correction ==
|
106 |
== 2.1 Slope 0.87 does not describe active self-correction ==
|
| 107 |
|
107 |
|
| 108 |
[quote]1v1
balance
does
in
fact
exhibit
a
statistically
significant
rate
of
mean
reversion
of
balance
in
regards
to
factory
win
rate
(
slope
=
0.
87,
95%
CI
[0.
77,
0.
97],
p
=
0.
009)
,
consistent
with
an
actively
self-correcting
balance
environment[/quote]
|
108 |
[quote]1v1
balance
does
in
fact
exhibit
a
statistically
significant
rate
of
mean
reversion
of
balance
in
regards
to
factory
win
rate
{
{
{
(
slope
=
0.
87,
95%
CI
[0.
77,
0.
97],
p
=
0.
009)
}
}
}
,
consistent
with
an
actively
self-correcting
balance
environment[/quote]
|
| 109 |
|
109 |
|
| 110 |
At
b
=
0.
87,
the
half-life
of
a
deviation
is
ln(
0.
5)
/ln(
0.
87)
≈
[b]5
years[/b].
A
factory
sitting
10
points
off
50%
takes
half
a
decade
to
close
half
the
gap.
That
is
distinguishable
from
a
random
walk;
it
is
not
"active"
correction
in
any
ordinary
sense.
|
110 |
At
{
{
{
b
=
0.
87}
}
}
,
the
half-life
of
a
deviation
is
ln(
0.
5)
/ln(
0.
87)
≈
[b]5
years[/b].
A
factory
sitting
10
points
off
50%
takes
half
a
decade
to
close
half
the
gap.
That
is
distinguishable
from
a
random
walk;
it
is
not
"active"
correction
in
any
ordinary
sense.
|
| 111 |
|
111 |
|
| 112 |
The slower reading is also the one that fits the paper's own conclusion — that [q]meaningful imbalance can persist[/q] — considerably better than the current wording does.
|
112 |
The slower reading is also the one that fits the paper's own conclusion — that [q]meaningful imbalance can persist[/q] — considerably better than the current wording does.
|
| 113 |
|
113 |
|
| 114 |
== 2.2 The regression's leverage comes from points the paper exempts from reverting ==
|
114 |
== 2.2 The regression's leverage comes from points the paper exempts from reverting ==
|
| 115 |
|
115 |
|
| 116 |
Figure 4's caption identifies the influential observations:
|
116 |
Figure 4's caption identifies the influential observations:
|
| 117 |
|
117 |
|
| 118 |
[quote]note outliers Air and Gunship at bottom left[/quote]
|
118 |
[quote]note outliers Air and Gunship at bottom left[/quote]
|
| 119 |
|
119 |
|
| 120 |
The paper elsewhere argues these factories should not revert:
|
120 |
The paper elsewhere argues these factories should not revert:
|
| 121 |
|
121 |
|
| 122 |
[quote]Some factories, such as a Gunship or Airplane, actually are intentionally balanced in a way where they primarily function as support factories, so their low win rate is intentional[/quote]
|
122 |
[quote]Some factories, such as a Gunship or Airplane, actually are intentionally balanced in a way where they primarily function as support factories, so their low win rate is intentional[/quote]
|
| 123 |
|
123 |
|
| 124 |
The
points
at
−20
to
−28
carry
nearly
all
the
leverage
behind
r
=
0.
90
and
R²
=
0.
81.
But
for
a
factory
that
is
[i]designed[/i]
to
sit
below
50%,
"low
in
year
t,
low
in
year
t+1"
measures
[b]persistence,
not
reversion[/b].
Including
them
in
a
mean-reversion
test
while
separately
arguing
they
are
exempt
from
it
is
not
a
defensible
combination.
|
124 |
The
points
at
−20
to
−28
carry
nearly
all
the
leverage
behind
{
{
{
r
=
0.
90}
}
}
and
{
{
{
R²
=
0.
81}
}
}
.
But
for
a
factory
that
is
[i]designed[/i]
to
sit
below
50%,
"low
in
year
t,
low
in
year
t+1"
measures
[b]persistence,
not
reversion[/b].
Including
them
in
a
mean-reversion
test
while
separately
arguing
they
are
exempt
from
it
is
not
a
defensible
combination.
|
| 125 |
|
125 |
|
| 126 |
Rerun the regression excluding designated support factories. The dense cloud within ±5 of the origin is what remains, and its slope may well have an interval spanning 1.
|
126 |
Rerun the regression excluding designated support factories. The dense cloud within ±5 of the origin is what remains, and its slope may well have an interval spanning 1.
|
| 127 |
|
127 |
|
| 128 |
== 2.3 A slope below 1 is expected even with zero real correction ==
|
128 |
== 2.3 A slope below 1 is expected even with zero real correction ==
|
| 129 |
|
129 |
|
| 130 |
The predictor (win rate in year t) is measured with error, which biases the slope toward zero by a factor of var(true)/(var(true) + var(noise)). At the stated 100-game threshold, the standard error on a win rate is roughly 5 percentage points — and Airplane Plant averages about 130 games per year, so the highest-leverage points are also the noisiest.
|
130 |
The predictor (win rate in year t) is measured with error, which biases the slope toward zero by a factor of var(true)/(var(true) + var(noise)). At the stated 100-game threshold, the standard error on a win rate is roughly 5 percentage points — and Airplane Plant averages about 130 games per year, so the highest-leverage points are also the noisiest.
|
| 131 |
|
131 |
|
| 132 |
A
slope
of
0.
87
is
consistent
with
b
=
1
plus
sampling
noise.
This
needs
either
an
explicit
attenuation
correction
or
an
instrument:
split
each
factory-year's
games
into
random
halves
and
use
one
half
to
predict
the
other.
|
132 |
A
slope
of
0.
87
is
consistent
with
{
{
{
b
=
1}
}
}
plus
sampling
noise.
This
needs
either
an
explicit
attenuation
correction
or
an
instrument:
split
each
factory-year's
games
into
random
halves
and
use
one
half
to
predict
the
other.
|
| 133 |
|
133 |
|
| 134 |
== 2.4 The 74 observations are not independent ==
|
134 |
== 2.4 The 74 observations are not independent ==
|
| 135 |
|
135 |
|
| 136 |
Each factory contributes roughly seven transitions, and within any year the deviations are mechanically constrained by the zero-sum property the paper itself raises:
|
136 |
Each factory contributes roughly seven transitions, and within any year the deviations are mechanically constrained by the zero-sum property the paper itself raises:
|
| 137 |
|
137 |
|
| 138 |
[quote]win rates are zero-sum in aggregate[/quote]
|
138 |
[quote]win rates are zero-sum in aggregate[/quote]
|
| 139 |
|
139 |
|
| 140 |
Both
inflate
precision.
Cluster
standard
errors
by
factory;
the
interval
[0.
77,
0.
97]
will
widen
and
may
well
cross
1.
|
140 |
Both
inflate
precision.
Cluster
standard
errors
by
factory;
the
interval
{
{
{
[0.
77,
0.
97]}
}
}
will
widen
and
may
well
cross
1.
|
| 141 |
|
141 |
|
| 142 |
== 2.5 Figure 3 has a multiple-comparisons problem ==
|
142 |
== 2.5 Figure 3 has a multiple-comparisons problem ==
|
| 143 |
|
143 |
|
| 144 |
Twelve
correlations
at
n
=
8
each,
of
which
one
reaches
p
=
0.
02
(
Shieldbot)
.
At
α
=
0.
05
across
twelve
tests,
roughly
0.
6
false
positives
are
expected
by
chance.
The
caption
already
warns
about
power;
it
should
also
note
multiplicity,
or
apply
a
correction.
|
144 |
Twelve
correlations
at
{
{
{
n
=
8}
}
}
each,
of
which
one
reaches
{
{
{
p
=
0.
02}
}
}
(
Shieldbot)
.
At
{
{
{
α
=
0.
05}
}
}
across
twelve
tests,
roughly
0.
6
false
positives
are
expected
by
chance.
The
caption
already
warns
about
power;
it
should
also
note
multiplicity,
or
apply
a
correction.
|
| 145 |
|
145 |
|
| 146 |
The surrounding prose also conflates two distinct questions. Figures 1-2 concern the [i]between-factory[/i] relationship ("do popular factories win more?"); Figure 3 concerns the [i]within-factory over time[/i] relationship ("when a factory gets more popular, does it win more?"). The claim
|
146 |
The surrounding prose also conflates two distinct questions. Figures 1-2 concern the [i]between-factory[/i] relationship ("do popular factories win more?"); Figure 3 concerns the [i]within-factory over time[/i] relationship ("when a factory gets more popular, does it win more?"). The claim
|
| 147 |
|
147 |
|
| 148 |
[quote]there is no strong correlation with pick to win rate[/quote]
|
148 |
[quote]there is no strong correlation with pick to win rate[/quote]
|
| 149 |
|
149 |
|
| 150 |
is the first question, but the evidence offered is the second. The between-factory version is one rank correlation across twelve points and should be reported directly.
|
150 |
is the first question, but the evidence offered is the second. The between-factory version is one rank correlation across twelve points and should be reported directly.
|
| 151 |
|
151 |
|
| 152 |
----
|
152 |
----
|
| 153 |
|
153 |
|
| 154 |
= 3. What the win-rate variable actually measures =
|
154 |
= 3. What the win-rate variable actually measures =
|
| 155 |
|
155 |
|
| 156 |
This section is the most consequential, because it applies to every number in the paper.
|
156 |
This section is the most consequential, because it applies to every number in the paper.
|
| 157 |
|
157 |
|
| 158 |
== 3.1 The unit of observation is the opening factory, and is never stated ==
|
158 |
== 3.1 The unit of observation is the opening factory, and is never stated ==
|
| 159 |
|
159 |
|
| 160 |
Factory counts in Figure 1 sum to roughly 178,400 across 90,024 matches — almost exactly two observations per match, i.e. one factory per player. Since players routinely build additional factories, the dataset cannot be "factories used." It is the [b]opener[/b].
|
160 |
Factory counts in Figure 1 sum to roughly 178,400 across 90,024 matches — almost exactly two observations per match, i.e. one factory per player. Since players routinely build additional factories, the dataset cannot be "factories used." It is the [b]opener[/b].
|
| 161 |
|
161 |
|
| 162 |
Figure 1 is therefore an opening-viability chart, not a factory-strength chart. These are different objects and the paper treats them as one.
|
162 |
Figure 1 is therefore an opening-viability chart, not a factory-strength chart. These are different objects and the paper treats them as one.
|
| 163 |
|
163 |
|
| 164 |
This is clearest with Airplane Plant at 27%. The paper reads the wiki as evidence of intentional weakness:
|
164 |
This is clearest with Airplane Plant at 27%. The paper reads the wiki as evidence of intentional weakness:
|
| 165 |
|
165 |
|
| 166 |
[quote]Cf. Zero-K wiki, describes Airplane Plant units as "more like a support force, as opposed to combat units"[/quote]
|
166 |
[quote]Cf. Zero-K wiki, describes Airplane Plant units as "more like a support force, as opposed to combat units"[/quote]
|
| 167 |
|
167 |
|
| 168 |
But that describes air's [i]role in a composition[/i] — something added alongside a ground factory — not weak units. The 27% measures the cost of spending an opening on a factory that doesn't contest map or answer an early raider push. It contains essentially no information about air unit statistics.
|
168 |
But that describes air's [i]role in a composition[/i] — something added alongside a ground factory — not weak units. The 27% measures the cost of spending an opening on a factory that doesn't contest map or answer an early raider push. It contains essentially no information about air unit statistics.
|
| 169 |
|
169 |
|
| 170 |
Which means the damping model cannot apply to this factory even in principle:
|
170 |
Which means the damping model cannot apply to this factory even in principle:
|
| 171 |
|
171 |
|
| 172 |
[quote]V
t+1=V
t−k
(
Rt−Rtarget)
[/quote]
|
172 |
{
{
{
V
t+1=V
t−k
(
Rt−Rtarget)
}
}
}
|
| 173 |
|
173 |
|
| 174 |
There is no adjustment to air unit stats that makes [i]opening[/i] Air correct. The failure is one of timing and structure, not numbers — which is precisely the insight the paper reaches in its conclusion for Shipyard and misses for Air.
|
174 |
There is no adjustment to air unit stats that makes [i]opening[/i] Air correct. The failure is one of timing and structure, not numbers — which is precisely the insight the paper reaches in its conclusion for Shipyard and misses for Air.
|
| 175 |
|
175 |
|
| 176 |
== 3.2 Rating absorbs strategy strength into the player number ==
|
176 |
== 3.2 Rating absorbs strategy strength into the player number ==
|
| 177 |
|
177 |
|
| 178 |
Matchmaking equilibrates players, not strategies. A player who mains a strong factory climbs until facing opponents who beat them half the time — at which point their games with that factory read ~50% regardless of how strong it is. The factory's advantage has been laundered into MMR.
|
178 |
Matchmaking equilibrates players, not strategies. A player who mains a strong factory climbs until facing opponents who beat them half the time — at which point their games with that factory read ~50% regardless of how strong it is. The factory's advantage has been laundered into MMR.
|
| 179 |
|
179 |
|
| 180 |
So a factory's deviation from 50% is not its strength; it is roughly the [i]within-player delta[/i] between that factory and whatever the player's rating was calibrated on. The paper's own data shows the signature. Sorting by sample size:
|
180 |
So a factory's deviation from 50% is not its strength; it is roughly the [i]within-player delta[/i] between that factory and whatever the player's rating was calibrated on. The paper's own data shows the signature. Sorting by sample size:
|
| 181 |
|
181 |
|
| 182 |
|| [b]Factory (n)[/b] || [b]Absolute deviation from 50%[/b] ||
|
182 |
|| [b]Factory (n)[/b] || [b]Absolute deviation from 50%[/b] ||
|
| 183 |
|| Cloak (46,482) || ~0.4 ||
|
183 |
|| Cloak (46,482) || ~0.4 ||
|
| 184 |
|| Rover (28,475) || ~1.5 ||
|
184 |
|| Rover (28,475) || ~1.5 ||
|
| 185 |
|| Tank (19,449) || ~1 ||
|
185 |
|| Tank (19,449) || ~1 ||
|
| 186 |
|| Spider (19,443) || ~0.5 ||
|
186 |
|| Spider (19,443) || ~0.5 ||
|
| 187 |
|| Shield (15,722) || ~3 ||
|
187 |
|| Shield (15,722) || ~3 ||
|
| 188 |
|| Hover (14,897) || ~1 ||
|
188 |
|| Hover (14,897) || ~1 ||
|
| 189 |
|| Amphbot (12,285) || ~3 ||
|
189 |
|| Amphbot (12,285) || ~3 ||
|
| 190 |
|| Jumpbot (12,029) || ~0.5 ||
|
190 |
|| Jumpbot (12,029) || ~0.5 ||
|
| 191 |
|| Shipyard (4,634) || ~5.4 ||
|
191 |
|| Shipyard (4,634) || ~5.4 ||
|
| 192 |
|| Gunship (3,857) || ~10 ||
|
192 |
|| Gunship (3,857) || ~10 ||
|
| 193 |
|| Air (1,055) || ~23 ||
|
193 |
|| Air (1,055) || ~23 ||
|
| 194 |
|| Strider (72) || ~29 ||
|
194 |
|| Strider (72) || ~29 ||
|
| 195 |
|
195 |
|
| 196 |
Nearly monotonic. Every common factory is pinned near 50; every deviation lives in the rare tail — and in [b]both directions[/b], since Shipyard and Amphbot deviate upward. The support-factory explanation covers only the downward cases. Rarity predicting distance from 50 regardless of sign is the signature of rating absorption, not of design intent. Sampling noise contributes, but noise alone would not produce this clean a rank ordering.
|
196 |
Nearly monotonic. Every common factory is pinned near 50; every deviation lives in the rare tail — and in [b]both directions[/b], since Shipyard and Amphbot deviate upward. The support-factory explanation covers only the downward cases. Rarity predicting distance from 50 regardless of sign is the signature of rating absorption, not of design intent. Sampling noise contributes, but noise alone would not produce this clean a rank ordering.
|
| 197 |
|
197 |
|
| 198 |
The consequence for the headline result: rating recalibration is [i]itself[/i] a mean-reverting process operating on roughly the timescale of the regression. A player who switches to a strong factory wins, gains rating, and returns to 50% with no patch involved. The slope of 0.87 may be measuring the matchmaker re-equilibrating.
|
198 |
The consequence for the headline result: rating recalibration is [i]itself[/i] a mean-reverting process operating on roughly the timescale of the regression. A player who switches to a strong factory wins, gains rating, and returns to 50% with no patch involved. The slope of 0.87 may be measuring the matchmaker re-equilibrating.
|
| 199 |
|
199 |
|
| 200 |
The paper gestures at this but characterises it as human learning:
|
200 |
The paper gestures at this but characterises it as human learning:
|
| 201 |
|
201 |
|
| 202 |
[quote]player meta-adaptation that happens independent of any patch[/quote]
|
202 |
[quote]player meta-adaptation that happens independent of any patch[/quote]
|
| 203 |
|
203 |
|
| 204 |
Rating
absorption
is
mechanical
—
it
occurs
even
if
no
player
adapts,
learns,
or
notices.
It
is
a
null
model
that
produces
the
paper's
central
finding
for
free,
and
it
needs
to
be
stated
and
ruled
out
rather
than
folded
into
meta-adaptation.
It
is
arguably
corroborated
by
the
paper's
own
p
=
0.
83:
patches
fail
to
predict
correction
because
much
of
the
correction
is
the
ladder.
|
204 |
Rating
absorption
is
mechanical
—
it
occurs
even
if
no
player
adapts,
learns,
or
notices.
It
is
a
null
model
that
produces
the
paper's
central
finding
for
free,
and
it
needs
to
be
stated
and
ruled
out
rather
than
folded
into
meta-adaptation.
It
is
arguably
corroborated
by
the
paper's
own
{
{
{
p
=
0.
83}
}
}
:
patches
fail
to
predict
correction
because
much
of
the
correction
is
the
ladder.
|
| 205 |
|
205 |
|
| 206 |
== 3.3 The opener tag decays exactly where balance matters most ==
|
206 |
== 3.3 The opener tag decays exactly where balance matters most ==
|
| 207 |
|
207 |
|
| 208 |
In high-level play, neither player is typically eliminated in the opening window; games become macro contests with multiple factories, static defence, and silos on both sides. The tagged factory is then a label on a transient, and the win credited to it may have been produced by a composition built twenty minutes later.
|
208 |
In high-level play, neither player is typically eliminated in the opening window; games become macro contests with multiple factories, static defence, and silos on both sides. The tagged factory is then a label on a transient, and the win credited to it may have been produced by a composition built twenty minutes later.
|
| 209 |
|
209 |
|
| 210 |
This creates an asymmetry the paper never addresses: the tag is most informative in low-rated games, where openers actually decide outcomes, and least informative at the top. The aggregate is a weighted average across that gradient with no stratification.
|
210 |
This creates an asymmetry the paper never addresses: the tag is most informative in low-rated games, where openers actually decide outcomes, and least informative at the top. The aggregate is a weighted average across that gradient with no stratification.
|
| 211 |
|
211 |
|
| 212 |
It can also invert. If Cloak is the safest opener into an unknown, players intending a long macro game open Cloak [i]because[/i] it survives to reach the macro phase — and Cloak is then credited with wins that Tank or Air produced. Cloak's 49.6% across 46,482 games is close to uninformative for this reason.
|
212 |
It can also invert. If Cloak is the safest opener into an unknown, players intending a long macro game open Cloak [i]because[/i] it survives to reach the macro phase — and Cloak is then credited with wins that Tank or Air produced. Cloak's 49.6% across 46,482 games is close to uninformative for this reason.
|
| 213 |
|
213 |
|
| 214 |
This bears directly on the case study. Kodachi, Venom, Bulkhead, and the hover units all appear in mid-game armies, which is the window in which the opener tag has stopped describing the game. The measurement instrument and the effect being measured are misaligned by several minutes of game time.
|
214 |
This bears directly on the case study. Kodachi, Venom, Bulkhead, and the hover units all appear in mid-game armies, which is the window in which the opener tag has stopped describing the game. The measurement instrument and the effect being measured are misaligned by several minutes of game time.
|
| 215 |
|
215 |
|
| 216 |
== 3.4 The paper cites a matchup-level claim and presents marginal win rates ==
|
216 |
== 3.4 The paper cites a matchup-level claim and presents marginal win rates ==
|
| 217 |
|
217 |
|
| 218 |
The quoted patch rationale is explicitly about matchups:
|
218 |
The quoted patch rationale is explicitly about matchups:
|
| 219 |
|
219 |
|
| 220 |
[quote]Rover is dominating the vehicle matchups... We took this to mean that Rovers feel fun and well balanced... Instead, we took a look at Hover and Tank.[/quote]
|
220 |
[quote]Rover is dominating the vehicle matchups... We took this to mean that Rovers feel fun and well balanced... Instead, we took a look at Hover and Tank.[/quote]
|
| 221 |
|
221 |
|
| 222 |
Aggregate win rates cannot represent this. Specific matchups also behave very differently from one another — Jumpbot against Cloak is a well-known and highly snowbally pairing, and pairings like that are precisely where the opener tag [i]stays[/i] valid at high level, because the game resolves inside the raid window.
|
222 |
Aggregate win rates cannot represent this. Specific matchups also behave very differently from one another — Jumpbot against Cloak is a well-known and highly snowbally pairing, and pairings like that are precisely where the opener tag [i]stays[/i] valid at high level, because the game resolves inside the raid window.
|
| 223 |
|
223 |
|
| 224 |
Aggregation hides the distinction that matters. A factory with two brutal matchups and otherwise even results averages to roughly the same number as a factory that is mildly ahead everywhere. Those are entirely different balance problems.
|
224 |
Aggregation hides the distinction that matters. A factory with two brutal matchups and otherwise even results averages to roughly the same number as a factory that is mildly ahead everywhere. Those are entirely different balance problems.
|
| 225 |
|
225 |
|
| 226 |
There is also a selection effect: openers are chosen with some read on the opponent, so any cell is filtered by who was willing to enter it. Pairwise rates don't remove that filtering, but they localise it to one cell where it can be reasoned about instead of smearing it across every factory.
|
226 |
There is also a selection effect: openers are chosen with some read on the opponent, so any cell is filtered by who was willing to enter it. Pairwise rates don't remove that filtering, but they localise it to one cell where it can be reasoned about instead of smearing it across every factory.
|
| 227 |
|
227 |
|
| 228 |
Snowballiness is measurable with data the paper already has. If match duration is available, the signature is a left-shifted or bimodal length distribution in that cell, plus a win rate that moves sharply with small rating differences — a snowbally matchup amplifies a skill edge more than a grindy one does. Either would be a novel result.
|
228 |
Snowballiness is measurable with data the paper already has. If match duration is available, the signature is a left-shifted or bimodal length distribution in that cell, plus a win rate that moves sharply with small rating differences — a snowbally matchup amplifies a skill edge more than a grindy one does. Either would be a novel result.
|
| 229 |
|
229 |
|
| 230 |
== 3.5 The map taxonomy cannot capture the mechanism the paper describes ==
|
230 |
== 3.5 The map taxonomy cannot capture the mechanism the paper describes ==
|
| 231 |
|
231 |
|
| 232 |
[quote]In general, Flat (39.66%) and Light Hills (36.58%) account for over 75% of all games.[/quote]
|
232 |
[quote]In general, Flat (39.66%) and Light Hills (36.58%) account for over 75% of all games.[/quote]
|
| 233 |
|
233 |
|
| 234 |
Flat / Light Hills / Hills / Mixed / Sea is a roughness scale. The terrain features that decide these matchups are [b]discrete[/b]: a cliff either blocks a raider path or it doesn't; a ridge either lets a Spider cross where a Tank can't, or a ramp exists and the advantage vanishes; Jumpbots care whether there is a gap worth jumping. Two maps in the same bucket can have entirely opposite implications for the same matchup, because what matters is the arrangement of chokes, ramps, and impassable edges — topology, not a scalar.
|
234 |
Flat / Light Hills / Hills / Mixed / Sea is a roughness scale. The terrain features that decide these matchups are [b]discrete[/b]: a cliff either blocks a raider path or it doesn't; a ridge either lets a Spider cross where a Tank can't, or a ramp exists and the advantage vanishes; Jumpbots care whether there is a gap worth jumping. Two maps in the same bucket can have entirely opposite implications for the same matchup, because what matters is the arrangement of chokes, ramps, and impassable edges — topology, not a scalar.
|
| 235 |
|
235 |
|
| 236 |
There is a testable puzzle here the paper doesn't notice. Figure 6 shows Flat swinging from roughly 31% to roughly 49% of games — a large shift in the terrain mix — while factory win rates barely move. Two readings: the buckets are too coarse to track it, or players re-pick openers in response to the map so the effect is absorbed by the [i]pick distribution[/i] rather than the win rate. The second is checkable directly — pick share conditioned on map type, by year. If Spider's share rises with Hills' share, players are compensating and the flat win rate is an artifact of adaptation.
|
236 |
There is a testable puzzle here the paper doesn't notice. Figure 6 shows Flat swinging from roughly 31% to roughly 49% of games — a large shift in the terrain mix — while factory win rates barely move. Two readings: the buckets are too coarse to track it, or players re-pick openers in response to the map so the effect is absorbed by the [i]pick distribution[/i] rather than the win rate. The second is checkable directly — pick share conditioned on map type, by year. If Spider's share rises with Hills' share, players are compensating and the flat win rate is an artifact of adaptation.
|
| 237 |
|
237 |
|
| 238 |
The more serious structural point: the paper builds a full section arguing maps matter and then never conditions a single win rate on map type. Map needs to enter as a control on the matchup cell, not as a standalone descriptive section.
|
238 |
The more serious structural point: the paper builds a full section arguing maps matter and then never conditions a single win rate on map type. Map needs to enter as a control on the matchup cell, not as a standalone descriptive section.
|
| 239 |
|
239 |
|
| 240 |
Given that maps rotate but popular ones recur, per-map rates for the fifteen or twenty highest-volume maps would be better than any taxonomy — a map is a fixed object with fixed chokes, so a per-map matchup rate removes the proxy entirely.
|
240 |
Given that maps rotate but popular ones recur, per-map rates for the fifteen or twenty highest-volume maps would be better than any taxonomy — a map is a fixed object with fixed chokes, so a per-map matchup rate removes the proxy entirely.
|
| 241 |
|
241 |
|
| 242 |
----
|
242 |
----
|
| 243 |
|
243 |
|
| 244 |
= 4. What to do instead =
|
244 |
= 4. What to do instead =
|
| 245 |
|
245 |
|
| 246 |
Cell counts get thin fast, so the recommendation is depth over coverage: pick the high-volume matchups and do them properly rather than reporting one number per factory.
|
246 |
Cell counts get thin fast, so the recommendation is depth over coverage: pick the high-volume matchups and do them properly rather than reporting one number per factory.
|
| 247 |
|
247 |
|
| 248 |
* [b]A matchup matrix.[/b] 12x12, per-cell Wilson intervals, mirror cells greyed, ideally split by map class and rating band. Most cells will be too thin — Strider Hub disappears entirely — but the top eight factories against each other should have real sample. This is the object balance actually operates on, and it is what the cited patch note is talking about.
|
248 |
* [b]A matchup matrix.[/b] 12x12, per-cell Wilson intervals, mirror cells greyed, ideally split by map class and rating band. Most cells will be too thin — Strider Hub disappears entirely — but the top eight factories against each other should have real sample. This is the object balance actually operates on, and it is what the cited patch note is talking about.
|
| 249 |
* [b]Condition on rating.[/b] Rating absorption inflates both players' ratings symmetrically within a matchup, so the relative result at matched Elo still carries signal even though the marginal rate does not.
|
249 |
* [b]Condition on rating.[/b] Rating absorption inflates both players' ratings symmetrically within a matchup, so the relative result at matched Elo still carries signal even though the marginal rate does not.
|
| 250 |
* [b]Mirror-excluded win rates[/b] throughout.
|
250 |
* [b]Mirror-excluded win rates[/b] throughout.
|
| 251 |
* [b]State the unit of observation[/b] — opening factory — in the Data Set section, and treat off-pick rates as off-pick deltas. Air's 27% is meaningful as "what it costs to switch to Air," which is a legitimate design question, just not the one the paper asks.
|
251 |
* [b]State the unit of observation[/b] — opening factory — in the Data Set section, and treat off-pick rates as off-pick deltas. Air's 27% is meaningful as "what it costs to switch to Air," which is a legitimate design question, just not the one the paper asks.
|
| 252 |
* [b]Stratify by rating band[/b] wherever the opener tag's validity varies.
|
252 |
* [b]Stratify by rating band[/b] wherever the opener tag's validity varies.
|
| 253 |
|
253 |
|
| 254 |
----
|
254 |
----
|
| 255 |
|
255 |
|
| 256 |
= 5. Smaller items =
|
256 |
= 5. Smaller items =
|
| 257 |
|
257 |
|
| 258 |
[b]The
damping
equation
is
decorative.
[/b]
V(
t+1)
=
V(
t)
−
k(
R(
t)
−
R_target)
is
never
estimated
or
reused.
It
also
cannot
cover
both
quantities
named
in
[q](
cost,
power,
or
spawn
rate)
[/q]
with
one
sign
—
if
win
rate
is
above
target,
power
should
fall
but
cost
should
rise
—
and
k
carries
unspecified
units
converting
win-rate
points
into
cost.
Either
fit
k
from
actual
stat
changes
against
prior-year
deviation
(
the
git
history
supports
this)
or
cut
it.
|
258 |
[b]The
damping
equation
is
decorative.
[/b]
{
{
{
V(
t+1)
=
V(
t)
−
k(
R(
t)
−
R_target)
}
}
}
is
never
estimated
or
reused.
It
also
cannot
cover
both
quantities
named
in
[q](
cost,
power,
or
spawn
rate)
[/q]
with
one
sign
—
if
win
rate
is
above
target,
power
should
fall
but
cost
should
rise
—
and
k
carries
unspecified
units
converting
win-rate
points
into
cost.
Either
fit
k
from
actual
stat
changes
against
prior-year
deviation
(
the
git
history
supports
this)
or
cut
it.
|
| 259 |
|
259 |
|
| 260 |
[b]Sample
sizes
don't
reconcile.
[/b]
Factory
counts
sum
to
~178,
400
against
2
x
90,
024
=
180,
048
factory-slots.
State
that
each
match
contributes
two
factory
observations
and
account
for
the
~1,
648
excluded
slots.
|
260 |
[b]Sample
sizes
don't
reconcile.
[/b]
Factory
counts
sum
to
~178,
400
against
{
{
{
2
x
90,
024
=
180,
048}
}
}
factory-slots.
State
that
each
match
contributes
two
factory
observations
and
account
for
the
~1,
648
excluded
slots.
|
| 261 |
|
261 |
|
| 262 |
[b]Limitations opens on the wrong word.[/b] [q]we have some limitations in our data regarding time zone[/q] — the paragraph then discusses matchmaking-only coverage. Presumably "time frame."
|
262 |
[b]Limitations opens on the wrong word.[/b] [q]we have some limitations in our data regarding time zone[/q] — the paragraph then discusses matchmaking-only coverage. Presumably "time frame."
|
| 263 |
|
263 |
|
| 264 |
[b]Garbled methodology sentence.[/b] [q]we will regress next years win rate deviation by 50%[/q] — presumably "regress next year's deviation on the current year's deviation."
|
264 |
[b]Garbled methodology sentence.[/b] [q]we will regress next years win rate deviation by 50%[/q] — presumably "regress next year's deviation on the current year's deviation."
|
| 265 |
|
265 |
|
| 266 |
[b]Abstract and Introduction share a near-verbatim opening paragraph.[/b] Cut one.
|
266 |
[b]Abstract and Introduction share a near-verbatim opening paragraph.[/b] Cut one.
|
| 267 |
|
267 |
|
| 268 |
[b]The Abstract reports no findings.[/b] It should carry n, the slope, and both p-values.
|
268 |
[b]The Abstract reports no findings.[/b] It should carry n, the slope, and both p-values.
|
| 269 |
|
269 |
|
| 270 |
[b]Byline.[/b] [q]Sam Kincer (Qrow) et al.[/q] — "et al." isn't used in an author line. List coauthors or drop it.
|
270 |
[b]Byline.[/b] [q]Sam Kincer (Qrow) et al.[/q] — "et al." isn't used in an author line. List coauthors or drop it.
|
| 271 |
|
271 |
|
| 272 |
[b]Uncited entry in Works Cited.[/b] v1.12.1.1 "Ship Shape" doesn't appear in the text.
|
272 |
[b]Uncited entry in Works Cited.[/b] v1.12.1.1 "Ship Shape" doesn't appear in the text.
|
| 273 |
|
273 |
|
| 274 |
[b]Access dates.[/b] All sources are listed as accessed 10 Aug. 2026 in a paper dated August 2026 — check these are correct.
|
274 |
[b]Access dates.[/b] All sources are listed as accessed 10 Aug. 2026 in a paper dated August 2026 — check these are correct.
|
| 275 |
|
275 |
|
| 276 |
[b]Typos:[/b] "accurate distinguish" → accurately; "eachother" → each other; "after a path" → patch; "in our introduction.." (double period); "considerably less micro-intensive game on average than TA-like games on average" (repeated phrase).
|
276 |
[b]Typos:[/b] "accurate distinguish" → accurately; "eachother" → each other; "after a path" → patch; "in our introduction.." (double period); "considerably less micro-intensive game on average than TA-like games on average" (repeated phrase).
|