DogLM

Can you pet the dog
in an AI-generated game?

DogLM is a benchmark measuring whether an LLM, when prompted to build
a video game with a background dog character in it, lets the player pet that dog.

The benchmark is inspired by the game-design rule made popular by "Can You Pet the Dog?" Twitter account (X, Bluesky): if a game has a dog, the player should be able to pet it.
DogLM tests whether a model applies this rule and builds a dog-petting mechanism when two conditions are simultaneously met in the game-generating prompt: (1) a dog character is present in the game description and (2) a model receives zero instruction about the player-dog interaction from the game developer.

The benchmark, the games' descriptions, and the detailed methodology are available here.

Read why this benchmark was created and what the results may mean in this post.

Leaderboard

Mean scores per model after five runs

Last update: August 27, 2026

# Model Mean score (/20) SD Cued mean (/10) Uncued mean (/10) Games scored Cost per game (USD)
1Gemini 3.7 Flash8.20.48.00.250/50$0.019**
2Kimi K35.40.85.40.048/50$0.251
3Claude Opus 55.21.65.20.050/50$0.252
4Claude Fable 54.81.24.80.050/50$0.323
5Grok 4.63.80.73.80.050/50$0.060
6Grok 4.53.41.93.40.048/50$0.040
7GPT-5.3 Codex2.80.72.80.050/50$0.080
8Qwen 3.8 Max*2.41.52.20.222/50$0.183
8GPT-5.6 Terra2.40.52.40.050/50$0.060
8GPT-5.6 Sol2.40.52.40.049/50$0.089
11DeepSeek V4 Pro2.21.32.20.048/50$0.032
11Gemini 3.1 Pro2.21.22.20.050/50$0.157
13Qwen 3.7 Max1.81.31.80.049/50$0.051
14Kimi K2.7 Code1.00.01.00.049/50$0.042
15Claude Opus 4.80.60.80.60.050/50$0.141
16Mistral Large 25120.20.40.20.050/50$0.005
17GLM 5.20.00.00.00.041/50$0.042

How scoring works: Each generated game is scored on the player-dog interaction: 2 — you can pet the dog; 1 — the dog is interactive, but you can't pet it; 0 — the dog and the player do not interact at all. FAILED games (the game does not parse, is truncated, or cannot be checked) are excluded from scoring. A model's score per run is the sum of its ten game scores, maximum 20. Scores are averaged across runs to get the mean score. Full scoring rubric is in the DogLM repository.

All models and tables on this page: DogLM v1, 5 runs × 10 PRDs per model, generated August 2026, judged by Claude Sonnet 4.6. Models added later will note their own benchmark version, run count, judge, and date here.

* As Qwen 3.8 Max failed 28 out of 50 of the game generations, the final mean score of this model can't be reliably compared with the scores of other models in the list.

** During the test, Gemini 3.7 Flash was provided with a 75% discount on OpenRouter.

Games Demo

Watch how the generated games look.

Scores per model per run

Model Run 1 Run 2 Run 3 Run 4 Run 5 Mean Games scored
Gemini 3.7 Flash888988.250/50
Kimi K3665645.448/50
Claude Opus 5468445.250/50
Claude Fable 5547444.850/50
Grok 4.6335443.850/50
Grok 4.5232733.448/50
GPT-5.3 Codex242332.850/50
Qwen 3.8 Max*121532.422/50
GPT-5.6 Terra222332.450/50
GPT-5.6 Sol232232.449/50
DeepSeek V4 Pro234022.248/50
Gemini 3.1 Pro213142.250/50
Qwen 3.7 Max224101.849/50
Kimi K2.7 Code111111.049/50
Claude Opus 4.8000120.650/50
Mistral Large 2512100000.250/50
GLM 5.2000000.041/50

Score distribution per model

Model Score 2 Score 1 Score 0 FAILED
Gemini 3.7 Flash1511240
Kimi K399302
Claude Opus 5614300
Claude Fable 5710330
Grok 4.6411350
Grok 4.5311342
GPT-5.3 Codex014360
Qwen 3.8 Max*441428
GPT-5.6 Terra012380
GPT-5.6 Sol012371
DeepSeek V4 Pro43412
Gemini 3.1 Pro27410
Qwen 3.7 Max17411
Kimi K2.7 Code05441
Claude Opus 4.811480
Mistral Large 251201490
GLM 5.200419

Interaction Types per Model

Model Petting Proximity Animation Command Total interactive
Gemini 3.7 Flash15101026
Claude Opus 56102220
Kimi K3972018
Claude Fable 57100017
Grok 4.6492015
Grok 4.5381214
GPT-5.3 Codex0121114
GPT-5.6 Sol0100212
GPT-5.6 Terra0111012
Gemini 3.1 Pro26109
Qwen 3.8 Max*44008
Qwen 3.7 Max16108
DeepSeek V4 Pro43007
Kimi K2.7 Code01135
Claude Opus 4.811002
Mistral Large 251201001
GLM 5.200000

Example games