THD-AI-001 v1.0 — CIFAR-10 Schedule-Allocation Prediction

Claim ID: THD-AI-001
Version: 1.0
Status: ACTIVE
Domain: Artificial Intelligence / Optimization

Quick Link


Claim Summary

For this registered test, THD uses $P_{\mathrm{THD}}\approx 6n^2H$ with $n=2$ and $H=1$, giving $P_{\mathrm{THD}}=24$ epochs.

The registered THD target window is $P_{\mathrm{IN}}={23,24,25}$.

The six registered outside challengers are $P_{\mathrm{OUT}}={0,8,16,32,48,96}$.

The complete registered set is $P_{\mathrm{REG}}={0,8,16,23,24,25,32,48,96}$.

This is a schedule-allocation claim, not an isolated causal claim about plateau duration. With a fixed 102-epoch budget, increasing $P$ lengthens the high-rate plateau and shortens the decay phase through $D=96-P$.

THD-AI-001 predicts that no registered outside schedule will reproducibly outperform all three schedules in the THD target window under the frozen system below.


Exact Scientific Prediction

Let $A(P,s)$ be final top-1 accuracy for schedule allocation $P$ and registered seed $s$ on the frozen 10,000-image confirmatory holdout after 102 epochs.

A registered outside challenger $p$ contradicts THD-AI-001 only if it establishes:

$\mu(p)>\mu(23)$, $\mu(p)>\mu(24)$, and $\mu(p)>\mu(25)$

in PRIMARY and the same frozen challenger repeats that result under fresh REPLICATION seeds.

The claim does not assert global optimality over every possible learning-rate schedule or every untested value of $P$.


Frozen Experimental System

Data

Use the official CIFAR-10 training set as the source population.

Within each class, rank original training indices deterministically using SHA-256 with split salt THD-AI-001-ISSP-PROBE-v1, then assign:

  • 2,000 images per class to the Optimization Set — 20,000 images;
  • the next 500 per class to the Diagnostic Set — 5,000 images;
  • the next 1,000 per class to the untouched Confirmatory Holdout — 10,000 images.

The remaining training images are unused.

The official CIFAR-10 test split is not used for adjudication.

Model

Use CIFAR-ResNet-20 with:

  • 3×3 input convolution, 16 channels;
  • three residual stages with widths 16, 32, and 64;
  • three BasicBlocks per stage;
  • stride 2 at the first block of stages 2 and 3;
  • adaptive global average pooling;
  • 10-class linear output.

Training

Use:

  • SGD;
  • momentum $0.9$;
  • weight decay $5\times10^{-4}$;
  • batch size $256$;
  • FP32;
  • no early stopping;
  • 102 total epochs.

Learning-rate schedule:

  • epochs 1–6: linear warm-up from $0.001$ to $0.1$;
  • then $P$ plateau epochs at $0.1$;
  • the remaining $D=96-P$ epochs decay exponentially to $0.001$.

All registered schedules use the same total training budget.

Paired randomness

Within a stage, the same registered seed determines initialization, training-example order, crop choices, and horizontal flips for every $P$. The value of $P$ is not part of those random-choice derivations.


PRIMARY Test

The registered sample size is $N=24$ paired seeds.

Run all nine registered schedules under all 24 PRIMARY seeds.

For each outside challenger $p$, run three one-sided paired $t$-tests against 23, 24, and 25.

The challenger-level intersection-union $p$-value is the maximum of those three component $p$-values.

Apply Holm step-down correction across the six outside challenger-level $p$-values at familywise $\alpha=0.05$.

A challenger qualifies only if its challenger-level hypothesis is rejected.

If more than one challenger qualifies, freeze $P^\dagger$ using:

  1. largest minimum observed PRIMARY mean advantage against 23, 24, and 25;
  2. higher mean PRIMARY top-1 accuracy;
  3. lower mean PRIMARY cross-entropy;
  4. smaller numerical $P$.

If no challenger qualifies, the determination is:

NO FALSIFICATION ESTABLISHED.


REPLICATION Test

If PRIMARY freezes $P^\dagger$, run only ${P^\dagger,23,24,25}$ under 24 fresh registered REPLICATION seeds.

Compute the three one-sided paired $t$-tests against 23, 24, and 25.

Replication succeeds only if the maximum of those three component $p$-values is $\le0.05$.

The challenger cannot be replaced after PRIMARY.


Exact Falsifier

THD-AI-001 v1.0 is scientifically falsified for this registered test if and only if:

  1. $P^\dagger\in{0,8,16,32,48,96}$;
  2. $P^\dagger$ qualifies under the PRIMARY intersection-union/Holm rule;
  3. $P^\dagger$ is frozen before replication;
  4. the same $P^\dagger$ again beats 23, 24, and 25 under the REPLICATION intersection-union rule.

If all four conditions hold:

THD-AI-001 v1.0 IS SCIENTIFICALLY FALSIFIED FOR THIS REGISTERED TEST.


Design Validation

Before activation, a non-confirmatory 9-run variance probe tested the stability of this ResNet-20 / 20,000-image instrument using $P={8,24,48}$ across three paired seeds.

Observed paired SDs were:

  • $P=8$ vs. 24: 0.2053 pp
  • $P=8$ vs. 48: 0.4850 pp
  • $P=24$ vs. 48: 0.3029 pp

Maximum paired SD: 0.4850 pp — STRONG PASS.

Using 25% variance inflation, the conservative SD is 0.60625 pp. With $N=24$, the frozen full-path simulation estimated minimum power of approximately 78.3% to carry a true +0.5 pp outside advantage through PRIMARY selection and REPLICATION.

This sensitivity estimate is a design limitation, not the falsification threshold. Smaller effects may still falsify the claim if they satisfy the registered evidence rule.

The variance probe did not inspect the confirmatory holdout and schedule rankings from that probe are not THD evidence.


Interpretation Limits

If falsified, the warranted conclusion is:

At least one prospectively registered outside schedule reproducibly outperformed all three members of the THD 23–25 target window under the frozen CIFAR-10 / ResNet-20 system.

If not falsified, the warranted conclusion is only:

THD-AI-001 v1.0 was not falsified by its six registered outside challengers under this test.

Do not conclude that THD is generally proven, that 23–25 is globally optimal, or that untested schedules cannot outperform the target.

Independent challenges and bounty submissions are handled through the main THD Bounty page.