Study 260725091 — reported

30 done · partially reproduced

Open result →
STUDY 260725091REPORTED JUL 31, 2026, 3:46 PM UTC

REPORTED

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

arXiv:2607.25091 · evidence revision b5dbeb008c98

  1. SPEC-FROZEN
  2. RUNNING
  3. RUNS-COMPLETE
  4. ANALYSIS-OPEN
  5. REPORTED
01

Claim under test

A complete matrix, not a favorable checkpoint.

The released experiments support stable PPO convergence and a robust capacity-headroom explanation of reward improvement across the reported model and corpus matrix.

The registered family contains 30 selected arms across 2 separately interpreted tracks. Every selected arm reached a terminal, claim-ready state before the verdict gate opened.

02

Frozen protocol

“Exact” is a versioned statement.

Execution protocol
v1.0.0 with prospectively logged amendments
Freeze revision
8a54ff56ab10d0795ae4e6845d61022b2916de97
Matrix SHA-256
0a55e04ecd7e2c3fca8a58326b050c91a4a41410630cc625569918d22aabfdb2
Config SHA-256
520b60137a885d1f80c66a81c217bd133507cba0944dc4315f051c28b7e7c40b
TRACK MManuscript-method bundle

15/15 claim-ready

TRACK RReleased-code path

15/15 claim-ready

03

Hardware and stack provenance

Observed hardware, neutral lab identities.

HardwareHostProfileInterpretation
GeForce RTX 4090LAB WORKSTATIONEXACTPaper-pinned execution profile
GeForce RTX 3090LAB DEVBOXEXACTPaper-pinned execution profile
RTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090LAB DEVBOXCOMPATCompatibility training; exact-stack claim evaluation

The publication export contains observed GPU labels and neutral host aliases. Device UUIDs, private paths, and unrelated operational records are excluded.

04

Final run ledger

30 selected arms · 2 tracks

30 done
reported Jul 31, 2026, 3:46 PM UTC

Arm 1, Track R: Pythia 70M, TinyStories, DoneArm 2, Track R: Pythia 70M, CNN / DailyMail, DoneArm 3, Track R: Pythia 70M, WikiText, DoneArm 4, Track R: Pythia 160M, TinyStories, DoneArm 5, Track R: Pythia 160M, CNN / DailyMail, DoneArm 6, Track R: Pythia 160M, WikiText, DoneArm 7, Track R: Pythia 410M, TinyStories, DoneArm 8, Track R: Pythia 410M, CNN / DailyMail, DoneArm 9, Track R: Pythia 410M, WikiText, DoneArm 10, Track R: SmolLM2 135M, TinyStories, DoneArm 11, Track R: SmolLM2 135M, CNN / DailyMail, DoneArm 12, Track R: SmolLM2 135M, WikiText, DoneArm 13, Track R: SmolLM2 360M, TinyStories, DoneArm 14, Track R: SmolLM2 360M, CNN / DailyMail, DoneArm 15, Track R: SmolLM2 360M, WikiText, DoneArm 16, Track M: Pythia 70M, TinyStories, DoneArm 17, Track M: Pythia 70M, CNN / DailyMail, DoneArm 18, Track M: Pythia 70M, WikiText, DoneArm 19, Track M: Pythia 160M, TinyStories, DoneArm 20, Track M: Pythia 160M, CNN / DailyMail, DoneArm 21, Track M: Pythia 160M, WikiText, DoneArm 22, Track M: Pythia 410M, TinyStories, DoneArm 23, Track M: Pythia 410M, CNN / DailyMail, DoneArm 24, Track M: Pythia 410M, WikiText, DoneArm 25, Track M: SmolLM2 135M, TinyStories, DoneArm 26, Track M: SmolLM2 135M, CNN / DailyMail, DoneArm 27, Track M: SmolLM2 135M, WikiText, DoneArm 28, Track M: SmolLM2 360M, TinyStories, DoneArm 29, Track M: SmolLM2 360M, CNN / DailyMail, DoneArm 30, Track M: SmolLM2 360M, WikiText, Done
Selected terminal arms for the released-code and manuscript-method tracks. Row verdicts are directional assessments; the study verdict appears below.
ArmTrackStateConfigurationGPUProfileRelease Δ [95% CI]Direction
001RDone. Selected run is terminal and claim-ready.Pythia 70MTinyStoriesGeForce RTX 4090lab-workstationEXACT+0.0571[-0.1872, +0.3045]INCONCLUSIVE
002RDone. Selected run is terminal and claim-ready.Pythia 70MCNN / DailyMailGeForce RTX 3090lab-devboxEXACT+0.0788[-0.1511, +0.3116]INCONCLUSIVE
003RDone. Selected run is terminal and claim-ready.Pythia 70MWikiTextGeForce RTX 4090lab-workstationEXACT-0.1026[-0.3301, +0.1348]INCONCLUSIVE
004RDone. Selected run is terminal and claim-ready.Pythia 160MTinyStoriesGeForce RTX 4090lab-workstationEXACT-0.0477[-0.6025, +0.4992]INCONCLUSIVE
005RDone. Selected run is terminal and claim-ready.Pythia 160MCNN / DailyMailGeForce RTX 4090lab-workstationEXACT-0.0822[-0.2649, +0.0992]INCONCLUSIVE
006RDone. Selected run is terminal and claim-ready.Pythia 160MWikiTextGeForce RTX 3090lab-devboxEXACT-0.1920[-0.6555, +0.2695]INCONCLUSIVE
007RDone. Selected run is terminal and claim-ready.Pythia 410MTinyStoriesGeForce RTX 3090lab-devboxEXACT+1.4715[+0.7398, +2.2281]MATCH
008RDone. Selected run is terminal and claim-ready.Pythia 410MCNN / DailyMailGeForce RTX 4090lab-workstationEXACT-0.5445[-0.8633, -0.2379]MATCH
009RDone. Selected run is terminal and claim-ready.Pythia 410MWikiTextGeForce RTX 3090lab-devboxEXACT-0.9716[-1.4751, -0.4611]MATCH
010RDone. Selected run is terminal and claim-ready.SmolLM2 135MTinyStoriesGeForce RTX 4090lab-workstationEXACT+0.3138[-0.1171, +0.7437]INCONCLUSIVE
011RDone. Selected run is terminal and claim-ready.SmolLM2 135MCNN / DailyMailGeForce RTX 3090lab-devboxEXACT-0.0328[-0.3554, +0.2823]INCONCLUSIVE
012RDone. Selected run is terminal and claim-ready.SmolLM2 135MWikiTextGeForce RTX 4090lab-workstationEXACT-0.0423[-0.2696, +0.1773]INCONCLUSIVE
013RDone. Selected run is terminal and claim-ready.SmolLM2 360MTinyStoriesRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1+0.2579[-0.1099, +0.6378]INCONCLUSIVE
014RDone. Selected run is terminal and claim-ready.SmolLM2 360MCNN / DailyMailRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1+0.3203[+0.1287, +0.5104]DIVERGES
015RDone. Selected run is terminal and claim-ready.SmolLM2 360MWikiTextRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1+0.3979[+0.1950, +0.6038]MATCH
016MDone. Selected run is terminal and claim-ready.Pythia 70MTinyStoriesGeForce RTX 3090lab-devboxEXACT+0.1968[-0.4734, +0.8717]INCONCLUSIVE
017MDone. Selected run is terminal and claim-ready.Pythia 70MCNN / DailyMailGeForce RTX 4090lab-workstationEXACT+0.1405[-0.1036, +0.3994]INCONCLUSIVE
018MDone. Selected run is terminal and claim-ready.Pythia 70MWikiTextGeForce RTX 4090lab-workstationEXACT+0.2714[-0.0259, +0.5656]INCONCLUSIVE
019MDone. Selected run is terminal and claim-ready.Pythia 160MTinyStoriesGeForce RTX 4090lab-workstationEXACT-0.1149[-0.6334, +0.4042]INCONCLUSIVE
020MDone. Selected run is terminal and claim-ready.Pythia 160MCNN / DailyMailGeForce RTX 3090lab-devboxEXACT+0.1572[-0.0706, +0.3854]INCONCLUSIVE
021MDone. Selected run is terminal and claim-ready.Pythia 160MWikiTextGeForce RTX 4090lab-workstationEXACT+0.0032[-0.2653, +0.2738]INCONCLUSIVE
022MDone. Selected run is terminal and claim-ready.Pythia 410MTinyStoriesRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1-0.2119[-0.8537, +0.4294]INCONCLUSIVE
023MDone. Selected run is terminal and claim-ready.Pythia 410MCNN / DailyMailGeForce RTX 3090lab-devboxEXACT+0.0538[-0.4590, +0.5522]INCONCLUSIVE
024MDone. Selected run is terminal and claim-ready.Pythia 410MWikiTextRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1+1.6959[+1.0325, +2.3337]DIVERGES
025MDone. Selected run is terminal and claim-ready.SmolLM2 135MTinyStoriesGeForce RTX 4090lab-workstationEXACT+2.2563[+1.9795, +2.5336]MATCH
026MDone. Selected run is terminal and claim-ready.SmolLM2 135MCNN / DailyMailGeForce RTX 3090lab-devboxEXACT+0.2020[-0.0804, +0.4749]INCONCLUSIVE
027MDone. Selected run is terminal and claim-ready.SmolLM2 135MWikiTextGeForce RTX 4090lab-workstationEXACT+0.4184[+0.2013, +0.6319]MATCH
028MDone. Selected run is terminal and claim-ready.SmolLM2 360MTinyStoriesRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1+0.2992[-0.0621, +0.6600]INCONCLUSIVE
029MDone. Selected run is terminal and claim-ready.SmolLM2 360MCNN / DailyMailRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1+0.1204[-0.0074, +0.2521]INCONCLUSIVE
030MDone. Selected run is terminal and claim-ready.SmolLM2 360MWikiTextRTX PRO 6000 Blackwell; exact-stack evaluation on RTX 3090lab-devboxCOMPAT1+0.5556[+0.3170, +0.7953]MATCH

1 COMPAT marks training on a documented compatibility stack. Claim-level metrics for all such arms were reevaluated on the exact paper stack.

Intervals are conditional on fixed checkpoints and retained generations. They do not estimate training-to-training or decoding-to-decoding variance.

05

Deviation register

Material differences stay attached to the result.

  1. D-RECIPE-CONFLICTmethod interpretation

    The manuscript, README, and released launchers specify conflicting budgets and operations, so no single executable recipe represents every authoritative source.

    Control: The frozen protocol reports released-code Track R and manuscript-stated Track M separately and never substitutes one for the other.

  2. D-ENVIRONMENT-LOCKsoftware environment

    The authors did not release a complete dependency lock capable of reconstructing the paper-era software environment verbatim.

    Control: A reconstructed exact-stack lock, import checks, CUDA smoke tests, and fully resolved analysis and compatibility locks are retained and containerized.

  3. D-COMPAT-TRAININGeight training arms

    Eight selected checkpoints were trained under the disclosed Blackwell-compatible stack because the exact paper stack could not execute correctly on that accelerator.

    Control: All eight received exact-stack claim-level reevaluation on compatible hardware; native results remain descriptive and exact-stack retraining is proposed as replication strengthening.

  4. D-MISSING-UPSTREAM-EVIDENCEreleased artifacts

    The released model collection omits reward checkpoints, equivalent PPO argument records, and the raw per-prompt records needed for a direct audit of the published table.

    Control: We retrained the unavailable stages, retained all local per-prompt records, and distinguish compatible numerical reproduction from identity with the unreleased original runs.

06

Results · reported

partially reproduced

Released numbers broadly reproduced; stronger claims remain unconfirmed

We broadly reproduced the released numerical table, but did not confirm the paper's stronger claims of universal stable convergence or a robust capacity-headroom rule; actual output-quality improvement remains unresolved.

What the evidence says

  • Twelve of fifteen published Track R deltas fell inside conditional prompt intervals, with three misses and substantial directional uncertainty.
  • The prescribed PPO budget was reached in eleven of fifteen Track R arms and twelve of fifteen Track M arms.
  • The explicit PPL necessary condition held in Track R but failed in Track M, while reward-signal informativeness was not operationalized for the full joint claim.
  • Internal reward scores are not independently calibrated across tracks, and the Qwen reviewer did not pass its outer-teacher reliability audit.

Limits on the claim

  • Each configuration has one reward-model and PPO training realization at seed 42, so training-to-training variance and seed-by-method interactions are not estimated.
  • Release-style prompt intervals bootstrap 200 retained prompt pairs conditional on one sampled generation per prompt; they do not include fresh decoding or fresh training variability.
  • Track M changes three operations together and trains a different reward model, so its component effects and cross-track reward-scale comparability are unresolved.
  • The release does not operationalize or retain a claim-level measure of reward-signal informativeness, preventing a complete test of the joint capacity-headroom hypothesis.
  • The blinded Qwen reviewer showed substantial position sensitivity under outer-teacher audit and is retained only as a fallible diagnostic of output quality.

Public bundle bound to evidence revision b5dbeb008c98833091b040ffa61463bbc17504c7. Artifact links are verified against the SHA-256 digests shown above before deployment.

EXT

Extensions remain fenced

New evidence may extend this record, never rewrite it.

Reviewer diagnostics, additional seeds, ablations, and future hardware sensitivity checks remain separate from this frozen primary verdict. Any extension receives its own provenance and cannot silently change the completed run manifest.

FROZEN PRIMARY RESULT · 30/30 CLAIM-READY · REWRITABLE: NO

Which follow-up would most improve confidence in this result?

No account or email. Your network address is used transiently for abuse prevention and is not sent to Discord.

Follow or challenge the work

Every useful objection should become a reproducible test.