Claim ID: THD-AI-001
Version: 1.0
Status: ACTIVE
Domain: Artificial Intelligence / Optimization
Quick Link
Claim Summary
For this registered test, THD uses $P_{\mathrm{THD}}\approx 6n^2H$ with $n=2$ and $H=1$, giving $P_{\mathrm{THD}}=24$ epochs.
The registered THD target window is $P_{\mathrm{IN}}={23,24,25}$.
The six registered outside challengers are $P_{\mathrm{OUT}}={0,8,16,32,48,96}$.
The complete registered set is $P_{\mathrm{REG}}={0,8,16,23,24,25,32,48,96}$.
This is a schedule-allocation claim, not an isolated causal claim about plateau duration. With a fixed 102-epoch budget, increasing $P$ lengthens the high-rate plateau and shortens the decay phase through $D=96-P$.
THD-AI-001 predicts that no registered outside schedule will reproducibly outperform all three schedules in the THD target window under the frozen system below.
Exact Scientific Prediction
Let $A(P,s)$ be final top-1 accuracy for schedule allocation $P$ and registered seed $s$ on the frozen 10,000-image confirmatory holdout after 102 epochs.
A registered outside challenger $p$ contradicts THD-AI-001 only if it establishes:
$\mu(p)>\mu(23)$, $\mu(p)>\mu(24)$, and $\mu(p)>\mu(25)$
in PRIMARY and the same frozen challenger repeats that result under fresh REPLICATION seeds.
The claim does not assert global optimality over every possible learning-rate schedule or every untested value of $P$.
Frozen Experimental System
Data
Use the official CIFAR-10 training set as the source population.
Within each class, rank original training indices deterministically using SHA-256 with split salt THD-AI-001-ISSP-PROBE-v1, then assign:
- 2,000 images per class to the Optimization Set — 20,000 images;
- the next 500 per class to the Diagnostic Set — 5,000 images;
- the next 1,000 per class to the untouched Confirmatory Holdout — 10,000 images.
The remaining training images are unused.
The official CIFAR-10 test split is not used for adjudication.
Model
Use CIFAR-ResNet-20 with:
- 3×3 input convolution, 16 channels;
- three residual stages with widths 16, 32, and 64;
- three BasicBlocks per stage;
- stride 2 at the first block of stages 2 and 3;
- adaptive global average pooling;
- 10-class linear output.
Training
Use:
- SGD;
- momentum $0.9$;
- weight decay $5\times10^{-4}$;
- batch size $256$;
- FP32;
- no early stopping;
- 102 total epochs.
Learning-rate schedule:
- epochs 1–6: linear warm-up from $0.001$ to $0.1$;
- then $P$ plateau epochs at $0.1$;
- the remaining $D=96-P$ epochs decay exponentially to $0.001$.
All registered schedules use the same total training budget.
Paired randomness
Within a stage, the same registered seed determines initialization, training-example order, crop choices, and horizontal flips for every $P$. The value of $P$ is not part of those random-choice derivations.
PRIMARY Test
The registered sample size is $N=24$ paired seeds.
Run all nine registered schedules under all 24 PRIMARY seeds.
For each outside challenger $p$, run three one-sided paired $t$-tests against 23, 24, and 25.
The challenger-level intersection-union $p$-value is the maximum of those three component $p$-values.
Apply Holm step-down correction across the six outside challenger-level $p$-values at familywise $\alpha=0.05$.
A challenger qualifies only if its challenger-level hypothesis is rejected.
If more than one challenger qualifies, freeze $P^\dagger$ using:
- largest minimum observed PRIMARY mean advantage against 23, 24, and 25;
- higher mean PRIMARY top-1 accuracy;
- lower mean PRIMARY cross-entropy;
- smaller numerical $P$.
If no challenger qualifies, the determination is:
NO FALSIFICATION ESTABLISHED.
REPLICATION Test
If PRIMARY freezes $P^\dagger$, run only ${P^\dagger,23,24,25}$ under 24 fresh registered REPLICATION seeds.
Compute the three one-sided paired $t$-tests against 23, 24, and 25.
Replication succeeds only if the maximum of those three component $p$-values is $\le0.05$.
The challenger cannot be replaced after PRIMARY.
Exact Falsifier
THD-AI-001 v1.0 is scientifically falsified for this registered test if and only if:
- $P^\dagger\in{0,8,16,32,48,96}$;
- $P^\dagger$ qualifies under the PRIMARY intersection-union/Holm rule;
- $P^\dagger$ is frozen before replication;
- the same $P^\dagger$ again beats 23, 24, and 25 under the REPLICATION intersection-union rule.
If all four conditions hold:
THD-AI-001 v1.0 IS SCIENTIFICALLY FALSIFIED FOR THIS REGISTERED TEST.
Design Validation
Before activation, a non-confirmatory 9-run variance probe tested the stability of this ResNet-20 / 20,000-image instrument using $P={8,24,48}$ across three paired seeds.
Observed paired SDs were:
- $P=8$ vs. 24: 0.2053 pp
- $P=8$ vs. 48: 0.4850 pp
- $P=24$ vs. 48: 0.3029 pp
Maximum paired SD: 0.4850 pp — STRONG PASS.
Using 25% variance inflation, the conservative SD is 0.60625 pp. With $N=24$, the frozen full-path simulation estimated minimum power of approximately 78.3% to carry a true +0.5 pp outside advantage through PRIMARY selection and REPLICATION.
This sensitivity estimate is a design limitation, not the falsification threshold. Smaller effects may still falsify the claim if they satisfy the registered evidence rule.
The variance probe did not inspect the confirmatory holdout and schedule rankings from that probe are not THD evidence.
Interpretation Limits
If falsified, the warranted conclusion is:
At least one prospectively registered outside schedule reproducibly outperformed all three members of the THD 23–25 target window under the frozen CIFAR-10 / ResNet-20 system.
If not falsified, the warranted conclusion is only:
THD-AI-001 v1.0 was not falsified by its six registered outside challengers under this test.
Do not conclude that THD is generally proven, that 23–25 is globally optimal, or that untested schedules cannot outperform the target.
Independent challenges and bounty submissions are handled through the main THD Bounty page.
