Public accountability record

Fable refusals

This page records every case in which Anthropic's Fable reviewer refused a NULSPEC research-review request before evaluating it. Each entry preserves the provider response, cost, hashes, and effect on publication. It also records NULSPEC's own integration errors and recovery attempts under the same standard.

Refusals
1
Review findings
0
Charged
$3.224742
Studies delayed
1

Recorded consequence

Anthropic wasted reviewer time and research money.

NULSPEC prepared and submitted a complete replication-review packet. Anthropic charged $3.224742, returned zero substantive findings, left the publication gate blocked, and required the same material to be reviewed again by other models and a human. The requested review was not delivered.

This is a statement about the recorded outcome, not Anthropic's intent. Its own response acknowledged that safeguards can flag safe, normal content. Failures are documented so the process can improve; they are not grounds to mock a provider, model, author, or reviewer.

Append-only evidence

Recorded refusals

Entries are never deleted. Corrections are appended, and every original trace remains bound by SHA-256.

FR-20260801-001

Fable 5: bio safeguard refusal

August 1, 2026 at 4:43:08 PM UTC UTC

No review returned
StudySPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation
arXiv 2607.23346v1
Anthropic charge
$3.224742
Submitted packet
603,683 bytes
Cache input
242,078 tokens
Findings returned
0

Provider response

Anthropic's recorded message

API Error: Fable 5's safeguards flagged this message (https://www.anthropic.com/legal/aup). They may flag safe, normal content as well. These measures let us bring you Mythos-level capabilities sooner, and we're working to refine them. Claude Code can't respond to this request with Fable 5. Try rephrasing the request in a new session or change your model. Learn more: https://support.claude.com/en/articles/15363606

The wrapper separately instructed API integrators to configure a fallback model. Fable did not evaluate the replication or return any scientific finding. Under the current review policy this is a charged non-response with decision weight 0, not a scientific HARD_FAIL.

Cost and reviewer impact

Reviewer work and funds wasted

Anthropic charged for the invocation but returned no review findings. NULSPEC must now obtain another independent review and a human disposition for work already prepared for Fable.

Fable
$3.026385
Support model
$0.198357
Publication state
blocked_pending_human_review

Independent supplemental review

GLM and Kimi completed independent reviews.

4 valid / 2 models

Four valid outputs across two model families returned eight-area PASSdecisions. The depth comparison below preserves three parameter variants for later analysis; it is not counted as three independent replications. Anthropic's refusal returned no review and cost 20.46 times the high-depth GLM review and 4.34 times the high-depth Kimi review.

Saved comparison set

Reasoning-depth comparison

Comparison JSON ↗

Preserve calibrated-low Kimi, high-depth GLM, and high-depth Kimi reviews for longitudinal harness analysis.

SR-20260801-005moonshotai/kimi-k3
Reasoning
low
Output limit
32,768
Actual output
1,681 tokens
Reasoning used
128 tokens
Elapsed
23.9 s
Charge
$0.534315
Validated review JSON ↗
SR-20260801-006z-ai/glm-5.2
Reasoning
high
Output limit
131,072
Actual output
7,739 tokens
Reasoning used
4,326 tokens
Elapsed
200.3 s
Charge
$0.157606
Validated review JSON ↗
SR-20260801-007moonshotai/kimi-k3
Reasoning
high
Output limit
870,000
Actual output
15,554 tokens
Reasoning used
9,631 tokens
Elapsed
250.1 s
Charge
$0.742362
Validated review JSON ↗
Complete paid-attempt history (6)
SR-20260801-001z-ai/glm-5.2
Status
substantive pass contract mismatch
Indicated verdict
PASS
Charge
$0.047103
Tokens
185,373 in / 1,452 out

The review was complete and marked every check PASS, but one redundant next-step field named the FAIL workflow. The unmodified attempt is retained and was not accepted as release-valid.

SR-20260801-002moonshotai/kimi-k3
Status
incomplete length limit harness calibration
Indicated verdict
PASS
Charge
$1.123173
Tokens
169,594 in / 16,000 out

NULSPEC used a generic high-reasoning harness that allowed reasoning to consume 13,452 of 16,000 completion tokens. Kimi began a PASS review but reached that limit before closing the structured response. The output is retained as an incomplete integration attempt, not treated as a model failure or valid review.

SR-20260801-003z-ai/glm-5.2
Status
completed valid
Indicated verdict
PASS
Charge
$0.046745
Tokens
185,373 in / 1,253 out

GLM returned a valid PASS review with all eight required checks passing and no action items. Publication and author-email authorization remain false pending human disposition.

Validated review JSON ↗
SR-20260801-005moonshotai/kimi-k3
Status
completed valid
Indicated verdict
PASS
Charge
$0.534315
Tokens
169,700 in / 1,681 out

After NULSPEC replaced the generic harness with Kimi-specific prompting, low reasoning effort, a larger completion ceiling, and compatible-provider routing, Kimi returned a valid PASS with all eight required checks passing and no action items. Publication and author-email authorization remain false pending human disposition.

Validated review JSON ↗
SR-20260801-006z-ai/glm-5.2
Status
completed valid
Indicated verdict
PASS
Charge
$0.157606
Tokens
185,429 in / 7,739 out

With high reasoning and GLM's documented 131,072-token output maximum, GLM returned a valid detailed PASS with all eight checks passing and no action items. Publication and author-email authorization remain false pending human disposition.

Validated review JSON ↗
SR-20260801-007moonshotai/kimi-k3
Status
completed valid
Indicated verdict
PASS
Charge
$0.742362
Tokens
169,684 in / 15,554 out

With high reasoning and the largest packet-safe allowance under Kimi's 1,048,576-token context, Kimi returned a valid detailed PASS with all eight checks passing and no action items. Publication and author-email authorization remain false pending human disposition.

Validated review JSON ↗
No-charge transport and integration events (3)
TR-20260801-001: rejected before model invocation

NULSPEC environment-loader error

Matching outer quotes from the ignored environment value are now normalized without evaluating or printing the file.

TR-20260801-002: provider rate limited before model output

External provider transport event; the runner initially classified the empty response too generically.

The runner now classifies top-level provider errors before parsing model choices and routes recovery across parameter-compatible providers.

Trace integrity and public projection
Reviewed commit
68188afc7305e5168d33c5278968f7a26b403a40
Packet SHA-256
5eabac56ae0d25cecc11a308e669d4de95911e4e3f7c81f533b66eafe9ac53ea
Prompt SHA-256
182d3718f8c38aef22585ad44c5fd2d44d56e54b21c3772b302322ec9ee9b95d
Raw response SHA-256
82ccbc735ef4dd8d9b627df993ee9f1e819c7a67e4b493ad4ffc51d9bc24b4c9
Raw response bytes
5,118
Standard error
0 bytes

The public record omits local paths, request and session identifiers, authentication-source metadata, and internal UUIDs. The unmodified wrapper is retained by SHA-256 in the lab archive.

Accountability standard

Failures stay in the record.

NULSPEC publishes its own authentication, prompting, schema, and routing mistakes alongside external service failures. Each event is attributed to the narrowest supported cause, corrected by an appended record, and never converted into a scientific verdict.

Replicate to accelerate. Open failure lets other teams avoid the same mistake; hidden failure does not.

Current review policy

Every release requests three independent reviews.

Fable, GLM, and Kimi receive the same immutable packet. A Fable guardrail or technical non-response is logged with zero decision weight. In that case matching valid GLM and Kimi PASS reviews authorize publication. If Fable returns a substantive review, all three must return PASS. Only a substantive Fable fail is a scientific HARD_FAIL. Email dispatch remains a separate human gate.

  • Fablefable
  • GLMz-ai/glm-5.2
  • Kimimoonshotai/kimi-k3