A lot of people have tried using ChatGPT as a D&D dungeon master. It works, and recent Projects and memory features make it more practical than it used to be. This post is about where ChatGPT-as-DM is genuinely good, what the newer continuity tools solve, and what still requires a purpose-built game.
Product details below are current as of August 4, 2026.
The short version
| ChatGPT as DM | Purpose-built AI DM (TableForge) | |
|---|---|---|
| Rules adjudication | LLM improvises | Programmatic rules engine |
| Dice rolls | Prose, or a general code tool when explicitly used | Real rolls recorded as game events |
| Character sheets | Files or text you and ChatGPT reconcile | Canonical sheet updates automatically |
| Cross-session memory | Projects, files, chat history, and generated memory | Persistent structured campaign state |
| Multiplayer | Shared workspace or a human host relaying actions | Up to 6 players, real-time or async |
| Combat tracking | Text ledger maintained through conversation | Tracks initiative, conditions, HP |
| Spell slot tracking | Contextual notes or uploaded sheet | Engine tracks every slot |
| Flexibility | Any setting, house rule, or writing workflow | Supported D&D 5e SRD campaign play |
The rest of this post is the long version.
Where ChatGPT-as-DM is actually fine
We don't want to oversell the gap. ChatGPT and similar general-purpose LLMs are genuinely good at some parts of running a game.
One-shots and short scenarios. A two-hour session with a clean start and end is well within what ChatGPT can handle. The state is still small, rules drift has not accumulated, and the narration is often vivid.
Creative warmups. If you're a human DM prepping a session and want to brainstorm NPCs, locations, or plot hooks, ChatGPT is a great writing partner. This is the LLM doing exactly what LLMs are good at.
Prompt experimentation. Curious about what an AI DM could feel like? Try ChatGPT for an hour. It's a low-cost way to figure out whether the category interests you before paying for a dedicated tool.
Solo freeform play. If your idea of a game is more "collaborative fiction" than "D&D campaign," ChatGPT is fine. You'll do the dice rolls yourself, you'll handle your own character sheet, and the AI will give you atmosphere.
These are real use cases. We don't want anyone reading this and thinking ChatGPT is useless for tabletop. It isn't.
What Projects and memory improve
ChatGPT Projects group chats, files, and instructions in one workspace. Project memory can draw from conversations inside that project, while shared Projects let collaborators work from common context. ChatGPT's broader memory system can also carry useful details between conversations.
That is a real improvement. A campaign brief, current character sheet, setting guide, and session recaps no longer need to be pasted into every new chat. Calling ChatGPT completely stateless is outdated.
Memory is still selected and synthesized context, not canonical game state. A project can contain an original sheet, a recap saying a sword was stolen, and a later scene that accidentally puts it back on the character's belt. The model has to reconcile those statements. A game database stores one current inventory record.
Projects make ChatGPT a more capable campaign notebook. They do not make the notebook the rules authority.
Where ChatGPT-as-DM still breaks down
No built-in game dice. When ChatGPT prints a roll in ordinary prose, the result is part of the generated response. On plans and surfaces with a code tool, it can explicitly use that tool to produce randomness. There is still no required game workflow that commits a roll, records its expression and result, prevents a retry from changing it, and applies the outcome once before narration continues.
No rules engine. This is the big one. Ask ChatGPT to adjudicate a grapple check and it might get it right. Ask it to track concentration across a five-round fight while managing action economy and three players' reactions and it will eventually invent rulings that don't exist. Concrete things that go wrong: a paladin smites after seeing the dice (the rules require the decision first); a spell's damage type is misremembered; concentration silently survives a failed save; bonus actions and reactions get conflated.
No canonical character sheet. You can upload a sheet or maintain one in a project. ChatGPT and the players still have to keep that document consistent with the conversation. There is no transaction that automatically spends the spell slot, applies the condition, moves the item, and rejects an illegal update.
Collaboration is not campaign multiplayer. Shared Projects give several people common files and context. They do not create one party roster, turn order, permission model, campaign event stream, or interface where each player submits an action as their character. A human host still coordinates the table.
Combat drift. This is what kills most ChatGPT campaigns. Combat is the densest application of rules in D&D, and it's where LLM-as-DM systems fail most visibly. The AI loses track of who's gone, what conditions are active, how much HP someone has, whether a spell is still concentrated on. You can correct it manually, but at that point you're DMing the AI, not playing.
Does a better prompt fix it?
This is the most common follow-up question, so it's worth answering directly: no, not the parts that matter.
A well-written system prompt genuinely improves narration, tone consistency, and how well the model stays in character. It does not create a structured campaign database, a legal-action validator, or a character sheet that updates transactionally. Those are missing infrastructure, not missing instructions, and no amount of prompt engineering supplies them.
What prompting does help with is exactly the use cases above: one-shots, prep, and freeform play.
How a purpose-built AI DM is different
A tool built specifically for running tabletop campaigns handles these problems architecturally, not by hoping the LLM gets it right.
In TableForge's case: a dedicated rules engine (real code, not an LLM) adjudicates dice, combat, conditions, and spell slots. Character sheets update automatically based on engine output. Campaign memory is persistent across sessions in structured form: NPCs, locations, decisions, quest state, consequences. Multiplayer is first-class, supporting up to six players in real-time or async. The LLM still does what it's good at (narration, NPC voice, atmosphere, reacting to your choices) and stops being asked to do what it isn't good at.
The point is not that ChatGPT is bad. It offers broader writing control, any genre, uploaded house rules, and an excellent workspace for a human DM. The point is that running a long campaign is a different problem than managing useful context, and it benefits from infrastructure built for it.
How to decide
If you want to play D&D for an evening, experiment with any setting, or help a human DM prepare, try ChatGPT first. It is flexible, familiar, and requires no game-specific account.
If you want to run a campaign that holds together across sessions, with real dice and a working character sheet and the option to invite friends, you want a purpose-built tool. There are a few in the category. We make one. How a TableForge session runs shows what a purpose-built tool does with a turn that ChatGPT has to improvise.
A pragmatic order of operations: try ChatGPT for an hour. Notice where it works and where it doesn't. If the category interests you, try a free tier of a dedicated tool. Pick the one that runs your game the way you want it run.