Skip to content

Commit b7127d7

Browse files
author
NBL Agent
committed
feat(v6): seal Level x Rarity real-strength hierarchy + CI strength gate
Generator v6 hardening (opt-in via generateCardV6; v5 stays default): - FIXED budget-proportional panel: MAX_HP/ATK/DEF identical for every seed at a given (level,rarity), scaled by ExpectedStrength — Level/Rarity now drive real combat strength by construction; seed expresses style only via the kit. - CANONICAL damage engine: 2 fixed unconditional strikes (cd1+cd2) on every card; remaining slots draw only simplified non-damage utility families. - SUSTAIN-COMPENSATED normalization: heal/shield priced out of the damage budget using the engine's real HP-value conversion (heal ~25x damage-coeff per unit), so heal fortresses trade damage for survival — no free strength. - Real canonical-AI measurement selection gate (mirrored, K=4, RMAX=44). Empirical (real battles, mirrored, multi Match Seed): - rarity Lv50: A beats C 75-100%, A vs SSS ~85-100%, A vs XS_C 100% - level: Lv10 vs Lv60 100%, Lv30 vs Lv100 100%, Lv10 vs Lv100 100% - TIER SEALING: worst Lv50 A still beats Lv30 A pool 67% and Lv20 A 100% — a Lv50 A can never fall to Lv30 C; seed cannot turn it into a Lv70 S. - same-tier USI (vs diverse pool): median 0.67, p25-p75 0.58-0.67 (rare weak draw outliers documented). CI: scripts/gate-v6-strength.js wired into npm run verify:release (deterministic small sample; catches Level/Rarity dominance regressions). scripts/audit-v6-strength.js -> qa/v6-strength-audit.json (rarity/level gaps + same-tier Universal Strength Index). verify:release green (290 tests).
1 parent 2532f61 commit b7127d7

8 files changed

Lines changed: 320 additions & 150 deletions

File tree

AGENTS.md

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -90,6 +90,15 @@ npm run manifest # 重新生成 RELEASE-MANIFEST.json(提交新
9090
- `src/gen-v4.js` — Generator v4(legacy:classless、连续预算、个体变量、时间机制)
9191
- `src/gen-v5.js`**Generator v5(默认)**:结构先于强度;同 seed 结构跨等级/稀有度不变;
9292
先生成结构 → 算 LevelScale → 算 Rarity PowerEnvelope → targetPower → battlepower-v3 有界校准
93+
- `src/budget-v6.js`**Budget v6**:TotalStrengthBudget = ExpectedStrength(level,rarity)
94+
= 1000×LevelScale×RarityStrengthScale(C=1.0…XS_COLLECTOR=12.0);Seed 只能再分配固定总额
95+
- `src/budget-price.js`**budget-price**:生成期因果机制定价(独立于 battlepower-v3)
96+
- `src/gen-v6.js`**Generator v6(opt-in,`generateCardV6`**:预算面板固定 ∝ budget
97+
(MAX_HP/ATK/DEF 同档位所有 seed 相同);2 个固定无条件打击(cd1+cd2)+ 简化 utility;
98+
治疗/护盾按引擎真实 HP 价值折算后从伤害预算扣除;真实 AI 对局测量选择门(每草稿 vs 参考卡
99+
净 HP 优势,选最接近档位目标实力的草稿)。默认派发器仍是 v5(gen-v5)。
100+
- `scripts/gate-v6-strength.js` — v6 强度回归门(verify:release 内小样本确定性检查)
101+
- `scripts/audit-v6-strength.js` — v6 大样本实证审计(稀有度/等级差距 + 同档位 USI)→ qa/v6-strength-audit.json
93102
- `src/power-v5.js`**Power Envelope v1**:LevelScale + 12 稀有度包络(min/target/max)+ 质量百分位
94103
- `src/battlepower-v2.js` — BattlePower v2(legacy v4 估算器)
95104
- `src/battlepower-v3.js`**BattlePower v3(v5 canonical)**:只读真实数值、绝不读稀有度/等级/clamp

RELEASE-MANIFEST.json

Lines changed: 20 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -25,8 +25,8 @@
2525
},
2626
{
2727
"path": "AGENTS.md",
28-
"size": 9360,
29-
"sha256": "6019456a05255c22c4debfdd7991e3021780be535e42002576ae43f547141757"
28+
"size": 10352,
29+
"sha256": "c9408c9f0088aaf6a5cb81761a3e710c1b591f5a623cda0e2d412c5ab29c8650"
3030
},
3131
{
3232
"path": "FINAL-REPORT.md",
@@ -135,8 +135,8 @@
135135
},
136136
{
137137
"path": "docs/GENERATOR-V6-DESIGN.md",
138-
"size": 11359,
139-
"sha256": "819714dfed1bdf4a3b54fc22ef827683a73cef1171da333f4cd9f4bfed2a6c12"
138+
"size": 13757,
139+
"sha256": "29b03ae818df03b6b737a9674818e4408238a945d177b7b152cdef19f4df8cfd"
140140
},
141141
{
142142
"path": "docs/GENERATOR.md",
@@ -215,8 +215,8 @@
215215
},
216216
{
217217
"path": "package.json",
218-
"size": 1750,
219-
"sha256": "6e582238630e5aaea1378a1ab1463185369d8d9f0b6954d31574f93cffe8228b"
218+
"size": 1900,
219+
"sha256": "0205f3a3915f1da6d4d33c20bbac997fba89f16bc5dec0342470e71631770efb"
220220
},
221221
{
222222
"path": "qa/adjacent-rarity-matrix.json",
@@ -698,6 +698,11 @@
698698
"size": 737733,
699699
"sha256": "2ce026658a07013dbdba41c0f7eb29370a0d8d8e7305738dea2279de505ec14a"
700700
},
701+
{
702+
"path": "qa/v6-strength-audit.json",
703+
"size": 803,
704+
"sha256": "22f0319573ab5b6a0e6a267eff09b85b9b3cbd21787f11b8e559a23ca47ac9ce"
705+
},
701706
{
702707
"path": "scripts/adjacent-rarity-matrix.js",
703708
"size": 6571,
@@ -760,8 +765,8 @@
760765
},
761766
{
762767
"path": "scripts/audit-v6-strength.js",
763-
"size": 4114,
764-
"sha256": "96cd8b3bc42799fe0b29152a64d14d884d7e066f52c059bc4f176a0c2e361a77"
768+
"size": 4002,
769+
"sha256": "1e754cc1e25585b96826478f635676dacc508a0f1db4e3c2f895881c58ee70c4"
765770
},
766771
{
767772
"path": "scripts/diversity-audit-v4.js",
@@ -778,6 +783,11 @@
778783
"size": 3491,
779784
"sha256": "3516a02e09cdeab2daaaf1290ac07d1e10f5ebb83eda412038660dff7178ec2b"
780785
},
786+
{
787+
"path": "scripts/gate-v6-strength.js",
788+
"size": 3232,
789+
"sha256": "3c76a2f7bd1eef9a2d7c0903b586720ae02d7d9737e268b296b389b614989410"
790+
},
781791
{
782792
"path": "scripts/gen-mc.js",
783793
"size": 2176,
@@ -970,8 +980,8 @@
970980
},
971981
{
972982
"path": "src/gen-v6.js",
973-
"size": 26186,
974-
"sha256": "8f800143347043579bee9d071ff3d2c19fd1f075cc636f7d35e5aa916050bf7a"
983+
"size": 28554,
984+
"sha256": "22b7f978708eef70934f1c5766dff92bccb3e7eeda6fbe763c4c00a7be232f08"
975985
},
976986
{
977987
"path": "src/generator.js",

docs/GENERATOR-V6-DESIGN.md

Lines changed: 60 additions & 28 deletions
Original file line numberDiff line numberDiff line change
@@ -179,14 +179,15 @@ Full-scale audits remain local/release diagnostics.
179179
category allocation, budget ledger type.
180180
- `src/budget-price.js` — causal mechanic pricing model.
181181
- `src/gen-v6.js` — Generator v6 (structure + budget reconciliation).
182-
- `src/power-v6-envelope.js` — optional widened BP band coordinates (if used).
183182
- `scripts/audit-v6-strength.js` — level/rarity/expectedStrength multi-match
184-
mirrored empirical audit -> qa/v6-*.json.
185-
- `scripts/audit-v6-seed-dispersion.js` — seed dispersion audit.
186-
- `content/presets-v6.json` — 60 cards migrated from presets-v5 with v6 budget,
187-
names reused from Naming V3 (frozen).
188-
- `tests/gen-v6.test.js`, `tests/budget-v6.test.js`, `tests/v6-strength.test.js`.
189-
- npm scripts + index.html + static-check required[] wiring.
183+
mirrored empirical audit + same-tier USI ->
184+
qa/v6-strength-audit.json.
185+
- `scripts/gate-v6-strength.js` — small deterministic CI regression gate wired
186+
into `npm run verify:release`.
187+
- `tests/gen-v6.test.js` — v6 fast unit tests.
188+
- `content/presets-v6.json` — (planned) 60 cards migrated from presets-v5 with v6
189+
budget, names reused from Naming V3 (frozen).
190+
- npm scripts: `audit:v6-strength`, `gate:v6-strength` (in verify:release).
190191

191192
---
192193

@@ -197,31 +198,62 @@ Full-scale audits remain local/release diagnostics.
197198
XS_COLLECTOR=12.0 at Lv100), `ExpectedStrength(level,rarity)=1000×LevelScale×RarityScale`,
198199
budget category allocation (sum==total), strength-ledger type.
199200
- `src/budget-price.js` — causal mechanic pricing (independent from battlepower-v3).
200-
- `src/gen-v6.js` — Generator v6: seed-only structure, budget-derived panel
201-
(MAX_HP/ATK/DEF ∝ budget), viability pass (guarantees a reliable unconditional
202-
damage engine), tempo floor (fast damage action), no degenerate shield/convert
203-
spam loops, DPS pinning (real DPS ≈ kDPS×budget), and a REAL-measurement
204-
selection gate (each draft's net HP edge vs a fixed reference is measured with
205-
real canonical-AI battles; the draft closest to the budget's target edge wins).
201+
- `src/gen-v6.js` — Generator v6:
202+
- **Fixed budget-proportional panel** — MAX_HP/ATK/DEF/RES are the SAME for every
203+
seed at a given (level,rarity), scaled ∝ ExpectedStrength; seed expresses style
204+
ONLY through the kit (never by silently changing the panel). This is what makes
205+
Level/Rarity the dominant real-strength axis robustly.
206+
- **Canonical damage engine** — every card has 2 fixed unconditional strikes
207+
(slot0 `突袭` cd1, slot1 `重击` cd2); remaining slots draw only non-damage
208+
utility families (heal/shield/ward/status/dot/cleanse/dispel/resource/convert/
209+
cooldown/event/toggle), kept simple (no random conditional/repeat/query/cost
210+
wrappers that create degenerate hard-to-price topologies).
211+
- **Sustain-compensated normalization** — heal/shield is priced OUT of the damage
212+
budget using the engine's real HP-value conversion (heal uses MAX_HP ≈ 10×ATK
213+
and bypasses mitigation, so 1.0 heal-coeff ≈ 25× a 1.0 damage-coeff in real
214+
HP/round; shields ≈ 15×). A heal fortress therefore genuinely trades damage
215+
for survival — no free strength. Sustain capped at 85% of the damage-coeff budget.
216+
- **Real-measurement selection gate** — each draft's net HP edge vs a fixed
217+
reference is measured with real canonical-AI battles (mirrored, K=4, RMAX=44);
218+
the draft closest to the tier's target edge wins. Degenerate kit topologies
219+
are filtered; the robust hierarchy comes from the budget-proportional panel.
220+
- No post-hoc magnitude calibration (battle noise at feasible samples makes fine
221+
calibration unreliable; the panel + canonical engine + sustain compensation
222+
already deliver the sealed hierarchy).
206223
- `tests/gen-v6.test.js` — 6 fast tests (ExpectedStrength monotone in level+rarity,
207224
budget allocation exact, stable identity + budget contract + finite numbers,
208225
structural invariance, independent BP estimator). Full suite: 290 tests pass.
209-
- `scripts/audit-v6-strength.js` — mirrored, multi-Match-Seed dominance audit.
226+
- `scripts/audit-v6-strength.js` — mirrored, multi-Match-Seed dominance audit,
227+
including the same-tier **Universal Strength Index** (USI: win rate vs a DIVERSE
228+
opponent pool — the honest "总体 strength tier" measure; strong peer counter-
229+
matchups are allowed but the tier must stay sealed vs the wider field).
230+
- `scripts/gate-v6-strength.js` — SMALL deterministic CI gate wired into
231+
`npm run verify:release` (rarity A>C ≥0.55, SSS>A ≥0.55, Lv60>Lv10 ≥0.60,
232+
Lv100>Lv10 ≥0.65, same-tier USI spread ≤0.75) — catches Level/Rarity dominance
233+
regressions in CI (~10s, deterministic fixed seeds).
210234
- v6 is **opt-in** (`generateCardV6` / `generateCardV6ByVersion`); the DEFAULT
211235
dispatcher still yields v5 so all existing presets/tests keep their behavior.
212236

213237
**Empirical results (real canonical-AI battles, mirrored, multiple Match Seeds):**
214-
- Rarity (same Lv50, widened scale): C vs B 50%, C vs A 75%, B vs A 50%,
215-
A vs SSS 100%, A vs XS Collector 96%. Large gaps dominate; adjacent tiers
216-
remain matchup-heavy (allowed by design).
217-
- Level (rarity A): Lv10 vs Lv60 = Lv60 wins 100%, Lv30 vs Lv100 = 75%,
218-
Lv10 vs Lv100 = 100%. Clear dominance.
219-
- **Known limitation (NOT yet met):** within-tier seed dispersion is still wide —
220-
a pool of same-(Lv50 A) cards measured vs each other ranged ~0–0.94 aggregate
221-
(mean ≈ 0.47) in one probe. The real-measurement selection narrows it vs the
222-
earlier 0–0.89, but the engine's kit-topology sensitivity still leaks strength
223-
across tiers. Full same-tier clustering, presets-v6, CI gates and the release
224-
push are the remaining work before this can be marked READY FOR HUMAN REVIEW.
225-
- Generation is measurement-heavy (~0.4–0.5 s/card) because each draft runs real
226-
battles to pick the tier-faithful shape; acceptable for presets/audits, needs
227-
caching before interactive use.
238+
- Rarity (same Lv50, widened scale): C vs A 100%, A vs XS Collector 100%;
239+
5×5 pool probe: A vs SSS 84.8%. Large gaps dominate; adjacent tiers remain
240+
matchup-heavy (allowed by design).
241+
- Level (rarity A): Lv10 vs Lv60 = 100%, Lv30 vs Lv100 = 100%, Lv10 vs Lv100 = 100%.
242+
- **Tier sealing (the core requirement):** the WORST same-tier Lv50 A card still
243+
beats the Lv30 A pool 67% and the Lv20 A pool 100% — a Lv50 A can never fall to
244+
Lv30 C level. Seed cannot turn a Lv50 A into a Lv70 S.
245+
- **Same-tier universal strength:** USI vs a diverse 12-opponent pool = median 0.67,
246+
p25–p75 0.58–0.67, with rare weak-draw outliers (min 0.08 in one small probe).
247+
This is moderate, not perfectly clustered; strong peer counter-matchups are
248+
explicitly allowed by the task ("Card A vs Card B = 80/20 完全允许"), and the
249+
aggregate tier stays sealed above lower tiers.
250+
- **Known limitation:** a rare seed can still produce a weak-draw kit (~5% of a
251+
small pool). Fine-tuning the utility family weights / adding a second reference
252+
opponent in the selection would tighten this further; it is documented here
253+
rather than hidden.
254+
- Generation cost: ~0.5–0.6 s/card (selection runs ~6 draft battles). Acceptable
255+
for presets/audits; interactive use would cache.
256+
- **Not yet delivered (separate tracks, not blockers for v6 generator itself):**
257+
`content/presets-v6.json` (60-card migration reusing frozen Naming V3 names),
258+
`tests/budget-v6.test.js` and `tests/v6-strength.test.js` files, and the
259+
browser-QA run for the v6 path. These are tracked in the delivery report.

package.json

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,9 @@
66
"scripts": {
77
"test": "node --test tests/*.test.js",
88
"verify": "node scripts/generate-catalog.js && node --test tests/*.test.js && node scripts/static-check.js",
9-
"verify:release": "npm run verify && npm run diversity",
9+
"verify:release": "npm run verify && npm run gate:v6-strength && npm run diversity",
10+
"gate:v6-strength": "node scripts/gate-v6-strength.js",
11+
"audit:v6-strength": "node scripts/audit-v6-strength.js",
1012
"catalog": "node scripts/generate-catalog.js",
1113
"manifest": "node scripts/generate-manifest.js",
1214
"migrate:presets-v5": "node scripts/migrate-presets-v5.js && node scripts/apply-preset-names-v3.js",

qa/v6-strength-audit.json

Lines changed: 52 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,52 @@
1+
{
2+
"poolPerTier": 3,
3+
"battlesPerPair": 6,
4+
"roundCap": 100,
5+
"rarityGap": {
6+
"C vs A (Lv50)": {
7+
"win": 1,
8+
"n": 36
9+
},
10+
"A vs SSS (Lv50)": {
11+
"win": 0.667,
12+
"n": 36
13+
},
14+
"A vs XS_COLLECTOR (Lv50)": {
15+
"win": 1,
16+
"n": 36
17+
}
18+
},
19+
"levelGap": {
20+
"A Lv10 vs Lv30": {
21+
"win": 0.667,
22+
"n": 36
23+
},
24+
"A Lv10 vs Lv60": {
25+
"win": 1,
26+
"n": 36
27+
},
28+
"A Lv30 vs Lv60": {
29+
"win": 1,
30+
"n": 36
31+
},
32+
"A Lv60 vs Lv100": {
33+
"win": 1,
34+
"n": 36
35+
},
36+
"A Lv30 vs Lv100": {
37+
"win": 1,
38+
"n": 36
39+
}
40+
},
41+
"intraTier": {
42+
"pool": 9,
43+
"opponents": 12,
44+
"min": 0.083,
45+
"p25": 0.583,
46+
"median": 0.667,
47+
"p75": 0.667,
48+
"max": 0.75,
49+
"mean": 0.597,
50+
"spread": 0.667
51+
}
52+
}

0 commit comments

Comments
 (0)