RUNNING
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
arXiv:2607.25091 · released source 64acb621037c
- SPEC-FROZEN
- RUNNING
- RUNS-COMPLETE
- ANALYSIS-OPEN
- REPORTED
Claim under test
Does capacity headroom predict where PPO helps?
The paper reports a capacity-headroom pattern: reinforcement learning gains should depend jointly on model scale and how well supervised fine-tuning already fits the task.
This is a claim-level matrix test, not a search for one favorable checkpoint. Numerical, directional, and family-wise assessments are specified in advance. The overall interpretation remains locked until all 15 Track R arms complete.
Frozen protocol
“Exact” is a versioned statement.
- Protocol
- 2607.25091-protocol-v1.0.0
- Freeze commit
8a54ff56ab10d0795ae4e6845d61022b2916de97- Matrix SHA-256
0a55e04ecd7e2c3fca8a58326b050c91a4a41410630cc625569918d22aabfdb2- Config SHA-256
520b60137a885d1f80c66a81c217bd133507cba0944dc4315f051c28b7e7c40b
Track R follows the released 250-step recipe. Track M separately tests manuscript-stated operations contradicted or absent in the executable release. Extensions cannot alter either primary track.
Hardware and stack provenance
Three cards, two execution profiles.
| Hardware | Host | Profile | Role |
|---|---|---|---|
| RTX 3090 · 24 GB | wtatum84 | EXACT | Paper-pinned Ampere execution |
| RTX 4090 · 24 GB | MonkeyPC | EXACT | Paper-pinned Ada execution |
| RTX PRO 6000 · 96 GB | wtatum84 | COMPAT | Blackwell throughput; exact-stack re-evaluation required |
Faster hardware changes wall time, not the frozen data, seed, hyperparameters, or evaluation. The execution profile remains attached to every arm.
Live run ledger
15 frozen configurations
9 done · 1 running · 1 queued · 4 failed
as of Jul 31, 2026, 1:29 AM UTC
| Arm | State | Configuration | GPU | Provenance | Verdict |
|---|---|---|---|---|---|
| 001 | Done. Run finished; no study verdict implied. | Pythia 70MTinyStories | RTX 4090MonkeyPC | EXACT | — |
| 002 | Done. Run finished; no study verdict implied. | Pythia 70MCNN / DailyMail | RTX 3090wtatum84 | EXACT | — |
| 003 | Done. Run finished; no study verdict implied. | Pythia 70MWikiText | RTX 4090MonkeyPC | EXACT | — |
| 004 | Done. Run finished; no study verdict implied. | Pythia 160MTinyStories | RTX 4090MonkeyPC | EXACT | — |
| 005 | Done. Run finished; no study verdict implied. | Pythia 160MCNN / DailyMail | RTX 4090MonkeyPC | EXACT | — |
| 006 | Done. Run finished; no study verdict implied. | Pythia 160MWikiText | RTX 3090wtatum84 | EXACT | — |
| 007 | Done. Run finished; no study verdict implied. | Pythia 410MTinyStories | RTX 3090wtatum84 | EXACT | — |
| 008 | Done. Run finished; no study verdict implied. | Pythia 410MCNN / DailyMail | RTX 4090MonkeyPC | EXACT | — |
| 009 | Done. Run finished; no study verdict implied. | Pythia 410MWikiText | RTX 3090wtatum84 | EXACT | — |
| 010 | Failed. Run ended without a valid completion. | SmolLM2 135MTinyStories | RTX 4090MonkeyPC | EXACT | — |
| 011 | Running. Actively executing; no result implied. | SmolLM2 135MCNN / DailyMail | RTX 3090wtatum84 | EXACT | — |
| 012 | Queued. Assigned, not yet started. | SmolLM2 135MWikiText | RTX 4090MonkeyPC | EXACT | — |
| 013 | Failed. Run ended without a valid completion. | SmolLM2 360MTinyStories | RTX PRO 6000wtatum84 | COMPAT1 | — |
| 014 | Failed. Run ended without a valid completion. | SmolLM2 360MCNN / DailyMail | RTX PRO 6000wtatum84 | COMPAT1 | — |
| 015 | Failed. Run ended without a valid completion. | SmolLM2 360MWikiText | RTX PRO 6000wtatum84 | COMPAT1 | — |
1 COMPAT marks RTX PRO 6000 Blackwell arms. The paper-pinned PyTorch build cannot target sm_120; the substitution and required exact-stack re-evaluation are recorded as D-001.
State is operational. Verdict remains blank until the frozen 15-configuration family is complete and the analysis gate opens.
Deviation register
Append-only by design.
- D-001RTX PRO 6000 arms only
The authors' pinned PyTorch 2.5.1 stack cannot execute CUDA kernels for Blackwell sm_120. These arms use PyTorch 2.9.1 with CUDA 12.8 while retaining the other paper pins.
Control: They are labeled COMPAT, kept out of exact-stack claims, and require exact-stack checkpoint re-evaluation before comparison.
- D-002Wall-clock execution only
The 15 configurations are scheduled concurrently across three locally controlled GPUs.
Control: Data, seed, model, hyperparameters, and evaluation remain unchanged; host, GPU, stack, pause events, and commit are recorded per arm.
- D-003Four SmolLM2 attempts
Four immutable attempts now have terminal failure manifests. The 135M TinyStories attempt exited after SFT; the three RTX PRO 6000 attempts ended before completing the full SFT, reward, PPO, and evaluation pipeline.
Control: Partial artifacts remain attached to their original attempt IDs and do not count as completed runs. Any recovery must create a new attempt with explicit provenance linkage; the claim-level result gate remains closed.
Results · gated
No claim-level result yet.
Publishing a moving conclusion would turn queue order into a scientific choice. The run ledger above is the only live status we report.
When all 15 arms finish, this gate opens first to the frozen analysis, then to an explicit conclusion: reproduced, partially reproduced, not reproduced, or inconclusive. Null results remain first-class results.
Local-compute extensions
Interesting, separate, unable to rewrite the replication.
After the paper-faithful matrix, we will test manuscript/code mismatches, local judges, readiness signals, and an outer teacher that reviews the Qwen 27B reviewer. These are fenced from primary outcomes in protocol and storage.
- manuscript-faithful reward initialization and dtype;
- Qwen 27B pairwise review of much smaller agents;
- Codex outer-teacher audit of the reviewer;
- additional seeds and stability diagnostics where justified.
Follow or challenge the work