Independent research replication
We rerun the experiments.
NULSPEC independently replicates published AI research on hardware we control—and publishes the run ledger in public before we know how it ends.
No pay-to-confirm. No hidden reruns. No success-only drawer.
What this is
Independent replication, done by people curious enough to check.
NULSPEC is a small team of enthusiasts and accelerationrationalists—a fused word, on purpose. We want the field to move fast, and we think the fastest route runs through checking the work.
We reproduce recent papers on our own machines, freeze the protocol before the first run, and publish the ledger whether or not the result cooperates.
Our operating protocol
The artifact is the argument.
A paper is not a vibe. A replication should leave enough evidence for a stranger to disagree productively.
- 01
Freeze the specification
The protocol, comparison rules, and exclusions enter Git before the first full-matrix run.
- 02
Reproduce before extending
Released code and manuscript-faithful interpretations stay separate. New ideas cannot rewrite the primary result.
- 03
Number every deviation
Hardware, stack, and implementation substitutions get an ID, a reason, and an impact control.
- 04
Publish the miss
Failures, null results, and irreproducible recipes receive the same artifact trail as a match.
- 05
Make rerunning cheaper
Commands, digests, checkpoints, and analysis code are preserved so the next person starts ahead of us.
Now running · Study 001
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
We are testing the complete 15-configuration released-code matrix before opening claim-level analysis. Operational status is public; partial conclusions are deliberately withheld.
Live run ledger
15 frozen configurations
9 done · 1 running · 1 queued · 4 failed
as of Jul 31, 2026, 1:29 AM UTC
| Arm | State | Configuration | GPU | Provenance | Verdict |
|---|---|---|---|---|---|
| 001 | Done. Run finished; no study verdict implied. | Pythia 70MTinyStories | RTX 4090MonkeyPC | EXACT | — |
| 002 | Done. Run finished; no study verdict implied. | Pythia 70MCNN / DailyMail | RTX 3090wtatum84 | EXACT | — |
| 003 | Done. Run finished; no study verdict implied. | Pythia 70MWikiText | RTX 4090MonkeyPC | EXACT | — |
| 004 | Done. Run finished; no study verdict implied. | Pythia 160MTinyStories | RTX 4090MonkeyPC | EXACT | — |
| 005 | Done. Run finished; no study verdict implied. | Pythia 160MCNN / DailyMail | RTX 4090MonkeyPC | EXACT | — |
| 006 | Done. Run finished; no study verdict implied. | Pythia 160MWikiText | RTX 3090wtatum84 | EXACT | — |
| 007 | Done. Run finished; no study verdict implied. | Pythia 410MTinyStories | RTX 3090wtatum84 | EXACT | — |
| 008 | Done. Run finished; no study verdict implied. | Pythia 410MCNN / DailyMail | RTX 4090MonkeyPC | EXACT | — |
| 009 | Done. Run finished; no study verdict implied. | Pythia 410MWikiText | RTX 3090wtatum84 | EXACT | — |
| 010 | Failed. Run ended without a valid completion. | SmolLM2 135MTinyStories | RTX 4090MonkeyPC | EXACT | — |
| 011 | Running. Actively executing; no result implied. | SmolLM2 135MCNN / DailyMail | RTX 3090wtatum84 | EXACT | — |
| 012 | Queued. Assigned, not yet started. | SmolLM2 135MWikiText | RTX 4090MonkeyPC | EXACT | — |
| 013 | Failed. Run ended without a valid completion. | SmolLM2 360MTinyStories | RTX PRO 6000wtatum84 | COMPAT1 | — |
| 014 | Failed. Run ended without a valid completion. | SmolLM2 360MCNN / DailyMail | RTX PRO 6000wtatum84 | COMPAT1 | — |
| 015 | Failed. Run ended without a valid completion. | SmolLM2 360MWikiText | RTX PRO 6000wtatum84 | COMPAT1 | — |
1 COMPAT marks RTX PRO 6000 Blackwell arms. The paper-pinned PyTorch build cannot target sm_120; the substitution and required exact-stack re-evaluation are recorded as D-001.
State is operational. Verdict remains blank until the frozen 15-configuration family is complete and the analysis gate opens.
A null result is a result.
A deviation hidden is a claim faked.
If you cannot rerun it, you are reading marketing.
Put a claim on the bench
Seen a result you want tested?
Nominate it. We choose papers we can honestly attempt on local compute and a fixed budget. If we take yours on, the protocol goes public before the first arm launches—and so does every deviation we are forced to make.
Nominate a paperKeep the apparatus alive
Support buys compute, not conclusions.
Donations go to GPU-hours, storage, and time. They cannot touch a verdict: protocols and decision rules are frozen before analysis. If you want more papers checked, faster, buy the lab monkey a little more runway.
Fund GPU-hours on Ko-fi