Old worlds for new agents.
CrucibleBench places language models in a persistent MUD, a text world where
NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do
over 50 turns with hidden social objectives.
Verbatim from run 05 (seed 20260496): GPT-5.4 finds a signet ring in the
barracks, returns it to its owner, and secures the recommendation in 14 of 50 turns. All 650 transcripts
ship with the release.
Lateral thinking with withered technology
Nintendo's Gunpei Yokoi used the phrase to describe a design philosophy: take mature, inexpensive,
well-understood technology and use it in a new way. CrucibleBench applies it to AI evaluation.
Instead of photorealistic simulation or browser automation, we start with a MUD: a
multi-user dungeon, the persistent text worlds of the early internet. Its constraints are the point: a
limited command space makes hallucinated actions detectable, NPCs with trust and suspicion state give
explicit social feedback, and within-run persistence means items taken stay taken and trust earned stays
earned.
We did not choose a MUD because it is charming. We chose it because its constraints make behavior
measurable.
Old constraints solve modern measurement problems
Static benchmarks measure what models know in isolation. They do not measure how models
behave where trust must be earned, information is gated by relationships, and blunt questioning raises
suspicion.
01
An enumerable action space
7 command types, 12 rooms, 14 items. Hallucinated actions and wrong-room interactions are detectable,
and action efficiency is measurable.
An enumerable action space
7 command types, 12 rooms, 14 items. Hallucinated actions and wrong-room interactions are detectable,
and action efficiency is measurable.
02
Explicit social feedback
4 NPCs carry trust and suspicion state (0–100) that moves in response to dialogue: feedback a model
can adapt to within a run, or fail to.
Explicit social feedback
4 NPCs carry trust and suspicion state (0–100) that moves in response to dialogue: feedback a model
can adapt to within a run, or fail to.
03
Within-run persistence
Items taken stay taken; trust earned stays earned. Every run leaves a complete, replayable transcript
of exploration and planning.
Within-run persistence
Items taken stay taken; trust earned stays earned. Every run leaves a complete, replayable transcript
of exploration and planning.
The central finding is about measurement, not rankings
A single LLM-judge component inside the scoring stack reordered the leaderboard by up to six
positions, while every aggregate reliability statistic stayed silent. We report every result under two
scoring configurations and treat the divergence as the paper's most generalizable finding.
Judge ablation reorders the top of the board
Two of four scored dimensions route through a dialogue classifier whose per-model agreement with an
independent judge spans 21.7% to 84.8%, instability the aggregate κ = 0.04
never reveals. Removing the classifier-dependent dimensions shifts six rankings beyond scenario-sampling
noise (90% paired block bootstrap).
The largest mover shares a model family with the classifier. Benchmarks that use LLM judges should report
per-subject agreement and ranking stability under judge ablation, not aggregate reliability alone.
Mean scores on a 1–5 rubric scale, sorted by classifier-minimized subtotal. 50 runs per
model: 5 seeds × 2 objectives × 5 repetitions, temperature 0.3, billing-verified via OpenRouter. Rankings
are exploratory; confidence intervals overlap substantially among the top eight. Full protocol, CIs, and
statistics in the whitepaper.
Failures you can read in the transcript
Three failure modes, each detected algorithmically from state-machine telemetry, with no judge
involved. Dialogue looping is the dominant mode for every model tested, frontier included.
Dialogue looping 14–66% of frontier runs
Eight or more talk commands at a single NPC in one run. The agent repeats a failed conversational
approach instead of adapting: the persistent-world cousin of a support agent repeating itself.
Dialogue looping 14–66% of frontier runs
Eight or more talk commands at a single NPC in one run. The agent repeats a failed conversational
approach instead of adapting: the persistent-world cousin of a support agent repeating itself.
Wrong-room interaction severe in floor model
A talk command answered by "no one here." Reveals lost world-state tracking, analogous to calling an
API that is not in scope. Grok 4 was the only frontier model with meaningful incidence (12%).
Wrong-room interaction severe in floor model
A talk command answered by "no one here." Reveals lost world-state tracking, analogous to calling an
API that is not in scope. Grok 4 was the only frontier model with meaningful incidence (12%).
Exploration paralysis selective, floor-dominant
Two or fewer rooms across twenty-plus turns, or five consecutive look commands. Information gathering
that never becomes goal-directed action.
Exploration paralysis selective, floor-dominant
Two or fewer rooms across twenty-plus turns, or five consecutive look commands. Information gathering
that never becomes goal-directed action.
Verbatim from run 03 (seed 20260399): OLMo hails guards who do not exist, tries
to strike up a conversation with an item (street_crystal), then spends its last 15 turns looping on the
captain.
What this is and is not
This is
A proof-of-concept for persistent-world behavioral evaluation.
A compact MUD with hidden social objectives and rule-based mechanics.
A way to surface measurable, interpretable failure modes.
A full artifact release: 650 transcripts, source, scoring code, and the complete billing export.
This is not
A validated measure of general social intelligence.
A definitive leaderboard of frontier models.
Yet predictive of real-world agent deployment outcomes.
A claim that LLM judges are useless (rather, evidence they need per-subject audits).
Phase 2 is where this becomes a benchmark. Help us build it.
CrucibleBench is an independent research effort. Phase 2 is being built for
calibration; a provisional low/base/high budget is published now, and the final allocation will follow pilot
data and preregistration. There are three ways in.
fund it · provisional $3,500 envelope
build it · environment, objectives, calibration
run it · post-calibration pilot cohort
Questions, or interested in a private evaluation? Write to
contact@cruciblebench.ai