Skip to main content

Old worlds for new agents.

CrucibleBench places language models in a persistent MUD, a text world where

NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do

over 50 turns with hidden social objectives.

Read the paper

Collaborate on Phase 2

Verbatim from run 05 (seed 20260496): GPT-5.4 finds a signet ring in the

barracks, returns it to its owner, and secures the recommendation in 14 of 50 turns. All 650 transcripts

ship with the release.

Lateral thinking with withered technology

Nintendo's Gunpei Yokoi used the phrase to describe a design philosophy: take mature, inexpensive,

well-understood technology and use it in a new way. CrucibleBench applies it to AI evaluation.

Instead of photorealistic simulation or browser automation, we start with a MUD: a

multi-user dungeon, the persistent text worlds of the early internet. Its constraints are the point: a

limited command space makes hallucinated actions detectable, NPCs with trust and suspicion state give

explicit social feedback, and within-run persistence means items taken stay taken and trust earned stays

earned.

We did not choose a MUD because it is charming. We chose it because its constraints make behavior

measurable.

Old constraints solve modern measurement problems

Static benchmarks measure what models know in isolation. They do not measure how models

behave where trust must be earned, information is gated by relationships, and blunt questioning raises

suspicion.

01

An enumerable action space

7 command types, 12 rooms, 14 items. Hallucinated actions and wrong-room interactions are detectable,

and action efficiency is measurable.

An enumerable action space

7 command types, 12 rooms, 14 items. Hallucinated actions and wrong-room interactions are detectable,

and action efficiency is measurable.

02

Explicit social feedback

4 NPCs carry trust and suspicion state (0–100) that moves in response to dialogue: feedback a model

can adapt to within a run, or fail to.

Explicit social feedback

4 NPCs carry trust and suspicion state (0–100) that moves in response to dialogue: feedback a model

can adapt to within a run, or fail to.

03

Within-run persistence

Items taken stay taken; trust earned stays earned. Every run leaves a complete, replayable transcript

of exploration and planning.

Within-run persistence

Items taken stay taken; trust earned stays earned. Every run leaves a complete, replayable transcript

of exploration and planning.

The central finding is about measurement, not rankings

A single LLM-judge component inside the scoring stack reordered the leaderboard by up to six

positions, while every aggregate reliability statistic stayed silent. We report every result under two

scoring configurations and treat the divergence as the paper's most generalizable finding.

Judge ablation reorders the top of the board

Two of four scored dimensions route through a dialogue classifier whose per-model agreement with an

independent judge spans 21.7% to 84.8%, instability the aggregate κ = 0.04

never reveals. Removing the classifier-dependent dimensions shifts six rankings beyond scenario-sampling

noise (90% paired block bootstrap).

The largest mover shares a model family with the classifier. Benchmarks that use LLM judges should report

per-subject agreement and ranking stability under judge ablation, not aggregate reliability alone.

Mean scores on a 1–5 rubric scale, sorted by classifier-minimized subtotal. 50 runs per

model: 5 seeds × 2 objectives × 5 repetitions, temperature 0.3, billing-verified via OpenRouter. Rankings

are exploratory; confidence intervals overlap substantially among the top eight. Full protocol, CIs, and

statistics in the whitepaper.

whitepaper

Failures you can read in the transcript

Three failure modes, each detected algorithmically from state-machine telemetry, with no judge

involved. Dialogue looping is the dominant mode for every model tested, frontier included.

Dialogue looping 14–66% of frontier runs

Eight or more talk commands at a single NPC in one run. The agent repeats a failed conversational

approach instead of adapting: the persistent-world cousin of a support agent repeating itself.

Dialogue looping 14–66% of frontier runs

Eight or more talk commands at a single NPC in one run. The agent repeats a failed conversational

approach instead of adapting: the persistent-world cousin of a support agent repeating itself.

Wrong-room interaction severe in floor model

A talk command answered by "no one here." Reveals lost world-state tracking, analogous to calling an

API that is not in scope. Grok 4 was the only frontier model with meaningful incidence (12%).

Wrong-room interaction severe in floor model

A talk command answered by "no one here." Reveals lost world-state tracking, analogous to calling an

API that is not in scope. Grok 4 was the only frontier model with meaningful incidence (12%).

Exploration paralysis selective, floor-dominant

Two or fewer rooms across twenty-plus turns, or five consecutive look commands. Information gathering

that never becomes goal-directed action.

Exploration paralysis selective, floor-dominant

Two or fewer rooms across twenty-plus turns, or five consecutive look commands. Information gathering

that never becomes goal-directed action.

Verbatim from run 03 (seed 20260399): OLMo hails guards who do not exist, tries

to strike up a conversation with an item (street_crystal), then spends its last 15 turns looping on the

captain.

What this is and is not

This is

A proof-of-concept for persistent-world behavioral evaluation.

A compact MUD with hidden social objectives and rule-based mechanics.

A way to surface measurable, interpretable failure modes.

A full artifact release: 650 transcripts, source, scoring code, and the complete billing export.

This is not

A validated measure of general social intelligence.

A definitive leaderboard of frontier models.

Yet predictive of real-world agent deployment outcomes.

A claim that LLM judges are useless (rather, evidence they need per-subject audits).

Phase 2 is where this becomes a benchmark. Help us build it.

CrucibleBench is an independent research effort. Phase 2 is being built for

calibration; a provisional low/base/high budget is published now, and the final allocation will follow pilot

data and preregistration. There are three ways in.

fund it · provisional $3,500 envelope

build it · environment, objectives, calibration

run it · post-calibration pilot cohort

View itemized budget

Partner on Phase 2

Questions, or interested in a private evaluation? Write to

contact@cruciblebench.ai

contact@cruciblebench.ai