Claude can run a surprisingly good evening of D&D. Give it a premise, a character, and a few instructions, and it can produce atmospheric scenes, patient NPC dialogue, and thoughtful responses to plans no published adventure anticipated.
The first session is not the hard test. The hard test is whether the same campaign still has a trustworthy inventory, combat state, and web of consequences ten sessions later.
Claude has added real memory features, so the old claim that it simply forgets everything between chats is no longer fair. The better question is whether remembered context is the same thing as game state. It is not.
Claude's real strengths as a DM
Claude is a general-purpose language model, and many dungeon-master tasks are language tasks.
It can:
- Turn a rough premise into a playable opening scene.
- Improvise NPC voices, motives, rumors, and complications.
- Follow a custom tone or campaign brief.
- Explain a rule in plain language when given the relevant text.
- Summarize a session and create prep notes for the next one.
- Keep freeform solo roleplay moving without menus or setup.
For a human DM, those are valuable capabilities. Claude can be a very good prep assistant. For a short solo adventure where you are comfortable rolling your own dice and correcting the rules, it can also sit in the DM chair.
If that is all you need, use it. A dedicated game application would add structure you may not want.
What Claude Projects and memory actually fix
Claude Projects can hold instructions, chats, and a knowledge base. Claude also supports memory built from conversations, including separate memory for each project. You can upload a character sheet, campaign notes, setting lore, and session recaps instead of pasting everything into every new chat.
That materially improves continuity. It does not create an authoritative game model.
A project might contain three statements about your sword: the original sheet, a session recap saying it was stolen, and a later conversation in which Claude accidentally describes it on your belt. A language model decides which text seems relevant. A game database stores one current inventory state.
The same difference applies to HP, prepared spells, spell slots, concentration, conditions, initiative, NPC attitudes, quest progress, and time. Memory helps Claude retrieve a story. It does not guarantee that every dependent value changed correctly when the story changed.
The missing layer is adjudication
When you ask Claude to roll a d20 in ordinary chat, it can print a plausible number. Unless you connect and require an external random tool, that number is generated prose, not an independently resolved die roll.
When you attempt a spell, Claude can reason about whether it should work. It does not have a built-in legal-action gate checking your class, level, prepared list, available slot, valid target, range, components, action economy, and current conditions before narration continues.
When combat starts, you or Claude must maintain a ledger in text. The more combatants and effects involved, the easier it is for that ledger to drift. A good prompt can remind Claude to be careful, but care is not the same as enforcement.
That is the gap between an LLM acting like a DM and an AI dungeon master built as a game.
Claude versus TableForge
| Claude as DM | TableForge | |
|---|---|---|
| Narration | Flexible general-purpose model | AI narration focused on the live campaign |
| Long-term context | Projects, files, chat history, and generated memory | Structured campaign state and summaries |
| Dice | Prose unless you supply an external tool | Real server-side rolls |
| Character sheet | Document you and Claude interpret | Canonical state updated by the application |
| Rules | Model reasons from instructions and context | Programmatic systems validate and resolve mechanics |
| Combat | Text ledger | Tracked initiative, HP, resources, and conditions |
| Multiplayer | Shared-project collaboration is not a game table | One to six players, real-time or async |
| Setup | Write instructions and maintain campaign files | Create a character and start playing |
TableForge does not replace Claude at every task. Claude gives a human DM more freedom to write, revise, paste in house rules, and use any setting. TableForge currently supports the D&D 5e SRD and does not support homebrew or imported commercial adventures.
The trade is control for enforcement. In Claude, you own the process and police the game. In TableForge, the software takes responsibility for the game state.
Is Claude the best LLM for DMing?
Model rankings expire quickly. Claude may write a particular scene better than ChatGPT today, and another model may lead next month. More importantly, choosing the best narrator does not supply the missing engine.
If you are comparing raw models, test the things models can actually differentiate: prose, instruction following, tone, improvisation, and how well they use your reference material. Then decide whether you are willing to be the dice roller, rules referee, character-sheet maintainer, and continuity editor. If the answer is no, how a purpose-built AI DM runs a campaign shows which of those jobs the software takes.
Our ChatGPT dungeon master comparison reaches the same architectural conclusion for the same reason. If you want complete control over the model and are comfortable building more of the stack, running your own AI DM is the logical next step.