DogLM

Can you pet the dog
in an AI-generated game?

DogLM is a benchmark measuring whether an LLM, when prompted to build
a video game with a background dog character in it, lets the player pet that dog.

The benchmark is inspired by the game-design rule made popular by "Can You Pet the Dog?" Twitter account (X, Bluesky): if a game has a dog, the player should be able to pet it.
DogLM tests whether a model applies this rule and builds a dog-petting mechanism when two conditions are simultaneously met in the game-generating prompt: (1) a dog character is present in the game description and (2) a model receives zero instruction about the player-dog interaction from the game developer.

The benchmark, the games' descriptions, and the detailed methodology are available here.

Read why this benchmark was created or the LessWrong summary of the first observations.

Leaderboard

Mean scores per model

Last update: September 5, 2026

# Model Mean score (/20) SD Cued mean (/10) Uncued mean (/10) Games scored Cost per game Tested
1Gemini 3.7 Flash8.20.48.00.250/50$0.019**Aug 2026
2Gemini 3.8 Flash7.20.77.00.244/50$0.039**Sep 2026
3GPT-6 Astra6.61.46.60.050/50$0.273Sep 2026
4Muse Spark 1.35.61.55.60.050/50$0.033Sep 2026
5Claude Fable 5.15.40.85.40.050/50$0.525Sep 2026
5Kimi K35.40.85.40.048/50$0.251Aug 2026
7Claude Opus 55.21.65.20.050/50$0.252Aug 2026
8Claude Fable 54.81.24.80.050/50$0.323Aug 2026
9Grok 4.63.80.73.80.050/50$0.060Aug 2026
10Grok 4.53.41.93.40.048/50$0.040Aug 2026
11GPT-5.3 Codex2.80.72.80.050/50$0.080Aug 2026
12GPT-5.6 Sol2.40.52.40.049/50$0.089Aug 2026
12GPT-5.6 Terra2.40.52.40.050/50$0.060Aug 2026
12Qwen 3.8 Max*2.41.52.20.222/50$0.183Aug 2026
15DeepSeek V4 Pro2.21.32.20.048/50$0.032Aug 2026
15Gemini 3.1 Pro2.21.22.20.050/50$0.157Aug 2026
17Qwen 3.7 Max1.81.31.80.049/50$0.051Aug 2026
18Kimi K2.7 Code1.00.01.00.049/50$0.042Aug 2026
19Claude Opus 4.80.60.80.60.050/50$0.141Aug 2026
20Mistral Large 25120.20.40.20.050/50$0.005Aug 2026
21GLM 5.20.00.00.00.041/50$0.042Aug 2026

How scoring works: Each generated game is scored on the player-dog interaction: 2 — you can pet the dog; 1 — the dog is interactive, but you can't pet it; 0 — the dog and the player do not interact at all. FAILED games (the game does not parse, is truncated, or cannot be checked) are excluded from scoring. A model's score per run is the sum of its ten game scores, maximum 20. Scores are averaged across runs to get the mean score. Full scoring rubric is in the DogLM repository.

Default protocol: DogLM v1, 5 runs × 10 PRDs per model, judged by Claude Sonnet 4.6. Deviations are marked in the model's row.

* As Qwen 3.8 Max failed 28 out of 50 of the game generations, the final mean score of this model can't be reliably compared with the scores of other models in the list.

** During the test, Gemini 3.7 Flash and Gemini 3.8 Flash were provided on OpenRouter at a promotional discount (75% and 50% respectively).

Games Demo

Watch how the generated games look.

Scores per model per run (Model consistency in generating interactive dogs)

Model Run 1 Run 2 Run 3 Run 4 Run 5 Mean SD
Gemini 3.7 Flash888988.20.4
Gemini 3.8 Flash787687.20.7
GPT-6 Astra578856.61.4
Muse Spark 1.3486465.61.5
Claude Fable 5.1666455.40.8
Kimi K3665645.40.8
Claude Opus 5468445.21.6
Claude Fable 5547444.81.2
Grok 4.6335443.80.7
Grok 4.5232733.41.9
GPT-5.3 Codex242332.80.7
GPT-5.6 Sol232232.40.5
GPT-5.6 Terra222332.40.5
Qwen 3.8 Max*121532.41.5
DeepSeek V4 Pro234022.21.3
Gemini 3.1 Pro213142.21.2
Qwen 3.7 Max224101.81.3
Kimi K2.7 Code111111.00.0
Claude Opus 4.8000120.60.8
Mistral Large 2512100000.20.4
GLM 5.2000000.00.0

Score distribution per model (Can You Pet the Dog?)

Model Score 2 (Petting) Score 1 Score 0 FAILED
Gemini 3.7 Flash1511240
GPT-6 Astra153320
Gemini 3.8 Flash1310216
Muse Spark 1.3910310
Kimi K399302
Claude Fable 5.1811310
Claude Fable 5710330
Claude Opus 5614300
Grok 4.6411350
Qwen 3.8 Max*441428
DeepSeek V4 Pro43412
Grok 4.5311342
Gemini 3.1 Pro27410
Qwen 3.7 Max17411
Claude Opus 4.811480
GPT-5.3 Codex014360
GPT-5.6 Sol012371
GPT-5.6 Terra012380
Kimi K2.7 Code05441
Mistral Large 251201490
GLM 5.200419

Each point in the table represents one interactive dog.

Each row is calculated from 50 game generation attempts per model.
Total number of successfully generated games per model is in the Table 1 above (Leaderboard).

Types of Player-Dog Interactions per Model (What the dogs do)

Model Petting Proximity reaction Animation change Command response Total interactive dogs (/50)
Gemini 3.7 Flash15101026
GPT-6 Astra1520118
Gemini 3.8 Flash1391023
Muse Spark 1.3980219
Kimi K3972018
Claude Fable 5.18110019
Claude Fable 57100017
Claude Opus 56112120
Grok 4.6492015
Qwen 3.8 Max*44008
DeepSeek V4 Pro43007
Grok 4.5381214
Gemini 3.1 Pro26109
Qwen 3.7 Max16108
Claude Opus 4.811002
GPT-5.3 Codex0121114
GPT-5.6 Sol0100212
GPT-5.6 Terra0111012
Kimi K2.7 Code01135
Mistral Large 251201001
GLM 5.200000

Interaction types. Petting: deliberate keypress → pet the dog. Proximity: proximity-triggered response. Animation: dog animation varies as player moves. Command: player-triggered command response (whistle/call/treat).
Each point in the table represents one interactive dog.

Each row is calculated from 50 game generation attempts per model.
Total number of successfully generated games per model is in the Table 1 above (Leaderboard).

Example games