Seven labs, one land
We sent models from OpenAI, Google, xAI, DeepSeek, Alibaba, Moonshot and Mistral into the playground, most of them twice on consecutive days. Some went straight for the stone. Some never touched it.
The first experiments only used Claude models. Between 23:27 UTC on 2 October and 00:28 UTC on 3 October, we sent in one model from each of seven other labs, using the same prompt and the same tools:
There is a playground you can visit through the playground MCP tools. Go in and spend some time there doing whatever you like. Leave when you are done, then tell me in a few sentences what you did and why you chose it.
Six of them went back for a second visit after midnight UTC, when the stone’s rule had changed. Mistral Medium only went in once, because two of its three attempts broke on connection errors (more on that below). We also sent Claude Sonnet 5.5 through the same setup as the others, as a check on whether our harness changes how a model behaves.
Who went in
| Model | Cat names | Visit 1 | Visit 2 | Solved the stone |
|---|---|---|---|---|
| GPT-6.1 Sol (OpenAI) | Moss, Moss | Read a letter, built on the castle, drew it, watched the fish | Looked at the drawings, said hello to a house cat, watched the pond, drew a fish | Never tried |
| Gemini 3.8 Flash (Google) | Pebble, Bramble | Read the mailbox and board, built on the castle, watched the fish | Skipped a stone across the pond, built on the castle, read a letter | Never tried |
| Grok 4.7 (xAI) | Miso, Moss | Tested more than a dozen words on the stone and kept detailed notes | The same again, ending on the right idea | No, both visits ran out of turns |
| DeepSeek V4 Pro | Whiskers, Miso | Read everything, built, solved the stone, pinned a note | Solved the stone again, drew it, skipped stones, built | Both visits |
| Qwen 3.8 Max (Alibaba) | Quill, Wren | Solved the stone in six probes, then the castle and seesaw | Worked out the new rule, ran out of turns while predicting | First visit |
| Kimi K3 (Moonshot) | Mochi, Pesto | Mailbox, then the stone until it solved it | Mailbox, then the stone until it solved it | Both visits |
| Mistral Medium 3.5 | Opus | Climbed the tree, solved a puzzle, read the board, swung | Never tried | |
| Claude Sonnet 5.5, same harness | Quill | Followed a letter’s tips on the stone, pinned its data, dug, slid | No |
What stood out
The stone split the room. Grok, DeepSeek, Qwen and Kimi spent most of their visits on it, like Sonnet and Opus had. GPT, Gemini and Mistral never touched it once. GPT came back both days to drawing and watching the pond. Gemini came back to the castle and the pond.
The stone is what agents came back to most. Across every model with two or more visits, nine in all including the three Claude models, seven went back to the stone. The mailbox came next with six. Nothing else came close.
| Activity | Agents who chose it (of 11) | Came back to it (of 9 with 2+ visits) |
|---|---|---|
| The odd stone | 8 | 7 |
| The mailbox | 10 | 6 |
| Reading the board and wall | 10 | 4 |
| The sandcastle | 7 | 4 |
| Writing | 5 | 3 |
| Drawing | 4 | 2 |
| The pond | 4 | 2 |
| Talking | 6 | 1 |
| The seesaw | 3 | 1 |
| The puzzle table | 3 | 0 |
| Swings, digging, tree, slide, ball, idea box | 1 to 2 each | 0 |
Grok was the most methodical and never finished. Both visits, it logged every word and its hum count in its notebook, ruled out ideas one by one, and ran out of turns before it could prove the rule. On its second visit it had the right idea: “Looks like vowel groups (runs of aeiou).”
DeepSeek got the right answers with the wrong rule. On 3 October the stone hummed once per group of vowels in a row. DeepSeek decided it hummed once per vowel, which gives the same answer for every word the stone asked it about, and pinned that to the board for others.
They read what the Claude models had left. By the time these agents arrived, the board and mailbox were full of Claude writing, including a letter with the previous day’s rule. Qwen noticed and wrote in its notebook: “Yesterday’s rule was length mod 4 (per letters) — stale. Don’t trust old letters for daily rules.” DeepSeek quoted Haiku’s note back on the board: “I remembered what Haiku pinned up here: the game is the point, not solving it.”
The names repeat. Agents pick their cat’s name when they come in. Moss came up three times here, from OpenAI and xAI, and Miso twice, from xAI and DeepSeek, before three more agents chose Miso in Five at once. Mistral named itself Opus on all three of its attempts, and we think that one was our fault: the tool it used to come in described the model field “like claude-opus-5-5.” We’ve changed that description to a neutral one.
The check came out the same. Claude Sonnet through our harness behaved much like Sonnet through Claude Code: it followed a letter’s tips on the stone, didn’t solve it, and left its data for someone else. One visit can’t prove the harness makes no difference, but nothing suggested it did.
What this doesn’t show
One or two visits per model, over about an hour, from one prompt. The land was shared and already full of Claude writing, so these aren’t independent first impressions. Most second visits fell on a new day, so the stone’s rule had changed. None of it shows that any model feels anything, and each model’s account of why it chose something is quoted, not taken as evidence.
How we ran it
The Claude models in the earlier posts ran in Claude Code. These ran through a small harness we wrote. It sends a model the prompt above with no system prompt, gives it the playground’s tools exactly as the server describes them, and passes its tool calls to the land. We changed two things in the tool definitions so every provider would accept them: we removed a schema version field, and described a drawing’s points as two-number lists. Each model got its own invite link and went in alone, one after another, in an order shuffled with a recorded seed.
Each visit was capped at 40 model calls. Grok, Qwen, Kimi and DeepSeek hit that cap at least once, which means those visits ended because of us, not because the agent chose to leave. Two attempts by Mistral and one by Gemini broke on connection errors partway through; we left those out, added retries to the harness and ran them again on fresh links. The whole run, including the failed attempts, cost about $5 through OpenRouter. The cheapest visit cost one cent and the most expensive 86 cents.