Study 001 — running

9 done · 1 running · 1 queued · 4 failed · conclusions not final

Open ledger →
STUDY 001AS OF JUL 31, 2026, 1:29 AM UTC

RUNNING

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

arXiv:2607.25091 · released source 64acb621037c

  1. SPEC-FROZEN
  2. RUNNING
  3. RUNS-COMPLETE
  4. ANALYSIS-OPEN
  5. REPORTED
01

Claim under test

Does capacity headroom predict where PPO helps?

The paper reports a capacity-headroom pattern: reinforcement learning gains should depend jointly on model scale and how well supervised fine-tuning already fits the task.

This is a claim-level matrix test, not a search for one favorable checkpoint. Numerical, directional, and family-wise assessments are specified in advance. The overall interpretation remains locked until all 15 Track R arms complete.

02

Frozen protocol

“Exact” is a versioned statement.

Protocol
2607.25091-protocol-v1.0.0
Freeze commit
8a54ff56ab10d0795ae4e6845d61022b2916de97
Matrix SHA-256
0a55e04ecd7e2c3fca8a58326b050c91a4a41410630cc625569918d22aabfdb2
Config SHA-256
520b60137a885d1f80c66a81c217bd133507cba0944dc4315f051c28b7e7c40b

Track R follows the released 250-step recipe. Track M separately tests manuscript-stated operations contradicted or absent in the executable release. Extensions cannot alter either primary track.

03

Hardware and stack provenance

Three cards, two execution profiles.

HardwareHostProfileRole
RTX 3090 · 24 GBwtatum84EXACTPaper-pinned Ampere execution
RTX 4090 · 24 GBMonkeyPCEXACTPaper-pinned Ada execution
RTX PRO 6000 · 96 GBwtatum84COMPATBlackwell throughput; exact-stack re-evaluation required

Faster hardware changes wall time, not the frozen data, seed, hyperparameters, or evaluation. The execution profile remains attached to every arm.

04

Live run ledger

15 frozen configurations

9 done · 1 running · 1 queued · 4 failed
as of Jul 31, 2026, 1:29 AM UTC

Arm 1: Pythia 70M, TinyStories, DoneArm 2: Pythia 70M, CNN / DailyMail, DoneArm 3: Pythia 70M, WikiText, DoneArm 4: Pythia 160M, TinyStories, DoneArm 5: Pythia 160M, CNN / DailyMail, DoneArm 6: Pythia 160M, WikiText, DoneArm 7: Pythia 410M, TinyStories, DoneArm 8: Pythia 410M, CNN / DailyMail, DoneArm 9: Pythia 410M, WikiText, DoneArm 10: SmolLM2 135M, TinyStories, FailedArm 11: SmolLM2 135M, CNN / DailyMail, RunningArm 12: SmolLM2 135M, WikiText, QueuedArm 13: SmolLM2 360M, TinyStories, FailedArm 14: SmolLM2 360M, CNN / DailyMail, FailedArm 15: SmolLM2 360M, WikiText, Failed
Track R execution ledger. DONE means only that a run finished; it is not a replication verdict.
ArmStateConfigurationGPUProvenanceVerdict
001Done. Run finished; no study verdict implied.Pythia 70MTinyStoriesRTX 4090MonkeyPCEXACT
002Done. Run finished; no study verdict implied.Pythia 70MCNN / DailyMailRTX 3090wtatum84EXACT
003Done. Run finished; no study verdict implied.Pythia 70MWikiTextRTX 4090MonkeyPCEXACT
004Done. Run finished; no study verdict implied.Pythia 160MTinyStoriesRTX 4090MonkeyPCEXACT
005Done. Run finished; no study verdict implied.Pythia 160MCNN / DailyMailRTX 4090MonkeyPCEXACT
006Done. Run finished; no study verdict implied.Pythia 160MWikiTextRTX 3090wtatum84EXACT
007Done. Run finished; no study verdict implied.Pythia 410MTinyStoriesRTX 3090wtatum84EXACT
008Done. Run finished; no study verdict implied.Pythia 410MCNN / DailyMailRTX 4090MonkeyPCEXACT
009Done. Run finished; no study verdict implied.Pythia 410MWikiTextRTX 3090wtatum84EXACT
010Failed. Run ended without a valid completion.SmolLM2 135MTinyStoriesRTX 4090MonkeyPCEXACT
011Running. Actively executing; no result implied.SmolLM2 135MCNN / DailyMailRTX 3090wtatum84EXACT
012Queued. Assigned, not yet started.SmolLM2 135MWikiTextRTX 4090MonkeyPCEXACT
013Failed. Run ended without a valid completion.SmolLM2 360MTinyStoriesRTX PRO 6000wtatum84COMPAT1
014Failed. Run ended without a valid completion.SmolLM2 360MCNN / DailyMailRTX PRO 6000wtatum84COMPAT1
015Failed. Run ended without a valid completion.SmolLM2 360MWikiTextRTX PRO 6000wtatum84COMPAT1

1 COMPAT marks RTX PRO 6000 Blackwell arms. The paper-pinned PyTorch build cannot target sm_120; the substitution and required exact-stack re-evaluation are recorded as D-001.

State is operational. Verdict remains blank until the frozen 15-configuration family is complete and the analysis gate opens.

05

Deviation register

Append-only by design.

  1. D-001RTX PRO 6000 arms only

    The authors' pinned PyTorch 2.5.1 stack cannot execute CUDA kernels for Blackwell sm_120. These arms use PyTorch 2.9.1 with CUDA 12.8 while retaining the other paper pins.

    Control: They are labeled COMPAT, kept out of exact-stack claims, and require exact-stack checkpoint re-evaluation before comparison.

  2. D-002Wall-clock execution only

    The 15 configurations are scheduled concurrently across three locally controlled GPUs.

    Control: Data, seed, model, hyperparameters, and evaluation remain unchanged; host, GPU, stack, pause events, and commit are recorded per arm.

  3. D-003Four SmolLM2 attempts

    Four immutable attempts now have terminal failure manifests. The 135M TinyStories attempt exited after SFT; the three RTX PRO 6000 attempts ended before completing the full SFT, reward, PPO, and evaluation pipeline.

    Control: Partial artifacts remain attached to their original attempt IDs and do not count as completed runs. Any recovery must create a new attempt with explicit provenance linkage; the claim-level result gate remains closed.

06

Results · gated

No claim-level result yet.

Publishing a moving conclusion would turn queue order into a scientific choice. The run ledger above is the only live status we report.

When all 15 arms finish, this gate opens first to the frozen analysis, then to an explicit conclusion: reproduced, partially reproduced, not reproduced, or inconclusive. Null results remain first-class results.

EXT

Local-compute extensions

Interesting, separate, unable to rewrite the replication.

After the paper-faithful matrix, we will test manuscript/code mismatches, local judges, readiness signals, and an outer teacher that reviews the Qwen 27B reviewer. These are fenced from primary outcomes in protocol and storage.

  • manuscript-faithful reward initialization and dtype;
  • Qwen 27B pairwise review of much smaller agents;
  • Codex outer-teacher audit of the reviewer;
  • additional seeds and stability diagnostics where justified.

Follow or challenge the work

Every useful objection should become a reproducible test.