REPORTED
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
arXiv:2607.25091 · evidence revision b5dbeb008c98
- SPEC-FROZEN
- RUNNING
- RUNS-COMPLETE
- ANALYSIS-OPEN
- REPORTED
Claim under test
A complete matrix, not a favorable checkpoint.
The released experiments support stable PPO convergence and a robust capacity-headroom explanation of reward improvement across the reported model and corpus matrix.
The registered family contains 30 selected arms across 2 separately interpreted tracks. Every selected arm reached a terminal, claim-ready state before the verdict gate opened.
Frozen protocol
“Exact” is a versioned statement.
- Execution protocol
- v1.0.0 with prospectively logged amendments
- Freeze revision
8a54ff56ab10d0795ae4e6845d61022b2916de97- Matrix SHA-256
0a55e04ecd7e2c3fca8a58326b050c91a4a41410630cc625569918d22aabfdb2- Config SHA-256
520b60137a885d1f80c66a81c217bd133507cba0944dc4315f051c28b7e7c40b
15/15 claim-ready
15/15 claim-ready
Hardware and stack provenance
Observed hardware, neutral lab identities.
| Hardware | Host | Profile | Interpretation |
|---|---|---|---|
| GeForce RTX 4090 | LAB WORKSTATION | EXACT | Paper-pinned execution profile |
| GeForce RTX 3090 | LAB DEVBOX | EXACT | Paper-pinned execution profile |
| RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090 | LAB DEVBOX | COMPAT | Compatibility training; exact-stack claim evaluation |
The publication export contains observed GPU labels and neutral host aliases. Device UUIDs, private paths, and unrelated operational records are excluded.
Final run ledger
30 selected arms · 2 tracks
30 done
reported Jul 31, 2026, 3:46 PM UTC
| Arm | Track | State | Configuration | GPU | Profile | Release Δ [95% CI] | Direction |
|---|---|---|---|---|---|---|---|
| 001 | R | Done. Selected run is terminal and claim-ready. | Pythia 70MTinyStories | GeForce RTX 4090lab-workstation | EXACT | +0.0571[-0.1872, +0.3045] | INCONCLUSIVE |
| 002 | R | Done. Selected run is terminal and claim-ready. | Pythia 70MCNN / DailyMail | GeForce RTX 3090lab-devbox | EXACT | +0.0788[-0.1511, +0.3116] | INCONCLUSIVE |
| 003 | R | Done. Selected run is terminal and claim-ready. | Pythia 70MWikiText | GeForce RTX 4090lab-workstation | EXACT | -0.1026[-0.3301, +0.1348] | INCONCLUSIVE |
| 004 | R | Done. Selected run is terminal and claim-ready. | Pythia 160MTinyStories | GeForce RTX 4090lab-workstation | EXACT | -0.0477[-0.6025, +0.4992] | INCONCLUSIVE |
| 005 | R | Done. Selected run is terminal and claim-ready. | Pythia 160MCNN / DailyMail | GeForce RTX 4090lab-workstation | EXACT | -0.0822[-0.2649, +0.0992] | INCONCLUSIVE |
| 006 | R | Done. Selected run is terminal and claim-ready. | Pythia 160MWikiText | GeForce RTX 3090lab-devbox | EXACT | -0.1920[-0.6555, +0.2695] | INCONCLUSIVE |
| 007 | R | Done. Selected run is terminal and claim-ready. | Pythia 410MTinyStories | GeForce RTX 3090lab-devbox | EXACT | +1.4715[+0.7398, +2.2281] | MATCH |
| 008 | R | Done. Selected run is terminal and claim-ready. | Pythia 410MCNN / DailyMail | GeForce RTX 4090lab-workstation | EXACT | -0.5445[-0.8633, -0.2379] | MATCH |
| 009 | R | Done. Selected run is terminal and claim-ready. | Pythia 410MWikiText | GeForce RTX 3090lab-devbox | EXACT | -0.9716[-1.4751, -0.4611] | MATCH |
| 010 | R | Done. Selected run is terminal and claim-ready. | SmolLM2 135MTinyStories | GeForce RTX 4090lab-workstation | EXACT | +0.3138[-0.1171, +0.7437] | INCONCLUSIVE |
| 011 | R | Done. Selected run is terminal and claim-ready. | SmolLM2 135MCNN / DailyMail | GeForce RTX 3090lab-devbox | EXACT | -0.0328[-0.3554, +0.2823] | INCONCLUSIVE |
| 012 | R | Done. Selected run is terminal and claim-ready. | SmolLM2 135MWikiText | GeForce RTX 4090lab-workstation | EXACT | -0.0423[-0.2696, +0.1773] | INCONCLUSIVE |
| 013 | R | Done. Selected run is terminal and claim-ready. | SmolLM2 360MTinyStories | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | +0.2579[-0.1099, +0.6378] | INCONCLUSIVE |
| 014 | R | Done. Selected run is terminal and claim-ready. | SmolLM2 360MCNN / DailyMail | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | +0.3203[+0.1287, +0.5104] | DIVERGES |
| 015 | R | Done. Selected run is terminal and claim-ready. | SmolLM2 360MWikiText | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | +0.3979[+0.1950, +0.6038] | MATCH |
| 016 | M | Done. Selected run is terminal and claim-ready. | Pythia 70MTinyStories | GeForce RTX 3090lab-devbox | EXACT | +0.1968[-0.4734, +0.8717] | INCONCLUSIVE |
| 017 | M | Done. Selected run is terminal and claim-ready. | Pythia 70MCNN / DailyMail | GeForce RTX 4090lab-workstation | EXACT | +0.1405[-0.1036, +0.3994] | INCONCLUSIVE |
| 018 | M | Done. Selected run is terminal and claim-ready. | Pythia 70MWikiText | GeForce RTX 4090lab-workstation | EXACT | +0.2714[-0.0259, +0.5656] | INCONCLUSIVE |
| 019 | M | Done. Selected run is terminal and claim-ready. | Pythia 160MTinyStories | GeForce RTX 4090lab-workstation | EXACT | -0.1149[-0.6334, +0.4042] | INCONCLUSIVE |
| 020 | M | Done. Selected run is terminal and claim-ready. | Pythia 160MCNN / DailyMail | GeForce RTX 3090lab-devbox | EXACT | +0.1572[-0.0706, +0.3854] | INCONCLUSIVE |
| 021 | M | Done. Selected run is terminal and claim-ready. | Pythia 160MWikiText | GeForce RTX 4090lab-workstation | EXACT | +0.0032[-0.2653, +0.2738] | INCONCLUSIVE |
| 022 | M | Done. Selected run is terminal and claim-ready. | Pythia 410MTinyStories | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | -0.2119[-0.8537, +0.4294] | INCONCLUSIVE |
| 023 | M | Done. Selected run is terminal and claim-ready. | Pythia 410MCNN / DailyMail | GeForce RTX 3090lab-devbox | EXACT | +0.0538[-0.4590, +0.5522] | INCONCLUSIVE |
| 024 | M | Done. Selected run is terminal and claim-ready. | Pythia 410MWikiText | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | +1.6959[+1.0325, +2.3337] | DIVERGES |
| 025 | M | Done. Selected run is terminal and claim-ready. | SmolLM2 135MTinyStories | GeForce RTX 4090lab-workstation | EXACT | +2.2563[+1.9795, +2.5336] | MATCH |
| 026 | M | Done. Selected run is terminal and claim-ready. | SmolLM2 135MCNN / DailyMail | GeForce RTX 3090lab-devbox | EXACT | +0.2020[-0.0804, +0.4749] | INCONCLUSIVE |
| 027 | M | Done. Selected run is terminal and claim-ready. | SmolLM2 135MWikiText | GeForce RTX 4090lab-workstation | EXACT | +0.4184[+0.2013, +0.6319] | MATCH |
| 028 | M | Done. Selected run is terminal and claim-ready. | SmolLM2 360MTinyStories | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | +0.2992[-0.0621, +0.6600] | INCONCLUSIVE |
| 029 | M | Done. Selected run is terminal and claim-ready. | SmolLM2 360MCNN / DailyMail | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | +0.1204[-0.0074, +0.2521] | INCONCLUSIVE |
| 030 | M | Done. Selected run is terminal and claim-ready. | SmolLM2 360MWikiText | RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devbox | COMPAT1 | +0.5556[+0.3170, +0.7953] | MATCH |
1 COMPAT marks training on a documented compatibility stack. Claim-level metrics for all such arms were reevaluated on the exact paper stack.
Intervals are conditional on fixed checkpoints and retained generations. They do not estimate training-to-training or decoding-to-decoding variance.
Deviation register
Material differences stay attached to the result.
- D-RECIPE-CONFLICTmethod interpretation
The manuscript, README, and released launchers specify conflicting budgets and operations, so no single executable recipe represents every authoritative source.
Control: The frozen protocol reports released-code Track R and manuscript-stated Track M separately and never substitutes one for the other.
- D-ENVIRONMENT-LOCKsoftware environment
The authors did not release a complete dependency lock capable of reconstructing the paper-era software environment verbatim.
Control: A reconstructed exact-stack lock, import checks, CUDA smoke tests, and fully resolved analysis and compatibility locks are retained and containerized.
- D-COMPAT-TRAININGeight training arms
Eight selected checkpoints were trained under the disclosed Blackwell-compatible stack because the exact paper stack could not execute correctly on that accelerator.
Control: All eight received exact-stack claim-level reevaluation on compatible hardware; native results remain descriptive and exact-stack retraining is proposed as replication strengthening.
- D-MISSING-UPSTREAM-EVIDENCEreleased artifacts
The released model collection omits reward checkpoints, equivalent PPO argument records, and the raw per-prompt records needed for a direct audit of the published table.
Control: We retrained the unavailable stages, retained all local per-prompt records, and distinguish compatible numerical reproduction from identity with the unreleased original runs.
Results · reported
partially reproducedReleased numbers broadly reproduced; stronger claims remain unconfirmed
We broadly reproduced the released numerical table, but did not confirm the paper's stronger claims of universal stable convergence or a robust capacity-headroom rule; actual output-quality improvement remains unresolved.
What the evidence says
- Twelve of fifteen published Track R deltas fell inside conditional prompt intervals, with three misses and substantial directional uncertainty.
- The prescribed PPO budget was reached in eleven of fifteen Track R arms and twelve of fifteen Track M arms.
- The explicit PPL necessary condition held in Track R but failed in Track M, while reward-signal informativeness was not operationalized for the full joint claim.
- Internal reward scores are not independently calibrated across tracks, and the Qwen reviewer did not pass its outer-teacher reliability audit.
Limits on the claim
- Each configuration has one reward-model and PPO training realization at seed 42, so training-to-training variance and seed-by-method interactions are not estimated.
- Release-style prompt intervals bootstrap 200 retained prompt pairs conditional on one sampled generation per prompt; they do not include fresh decoding or fresh training variability.
- Track M changes three operations together and trains a different reward model, so its component effects and cross-track reward-scale comparability are unresolved.
- The release does not operationalize or retain a claim-level measure of reward-signal informativeness, preventing a complete test of the joint capacity-headroom hypothesis.
- The blinded Qwen reviewer showed substantial position sensitivity under outer-teacher audit and is retained only as a fallible diagnostic of output quality.
21fab5e1de4eFull matrix report0a3bda3b36a5Machine-readable analysis91cfca155d47Extension roadmapd3d93f8372b5Website handoff contract17ab55d5b10aPublic bundle bound to evidence revision b5dbeb008c98833091b040ffa61463bbc17504c7. Artifact links are verified against the SHA-256 digests shown above before deployment.
Extensions remain fenced
New evidence may extend this record, never rewrite it.
Reviewer diagnostics, additional seeds, ablations, and future hardware sensitivity checks remain separate from this frozen primary verdict. Any extension receives its own provenance and cannot silently change the completed run manifest.
FROZEN PRIMARY RESULT · 30/30 CLAIM-READY · REWRITABLE: NO
Follow or challenge the work