Get the weekly newsletter that makes you better at Google Sheets, Productivity, and Finance.
TL;dr: At its DevDay on September 29, OpenAI launched Space, a shared workspace inside ChatGPT where coworkers, ChatGPT, and its AI agents work on the same files, along with Pages, a collaborative document editor. Collaborative slides and spreadsheets are announced as coming, and today a Space can already collect spreadsheets and files from connected services like Google Drive and Slack. That means the workbook is coming into the chat window, whether we are ready for it or not, and it is worth deciding now how you will vet anything that comes out of it.
Let me back up for a minute, because the context matters. Up until now, the way most of us used ChatGPT for model work was to ask for help in a chat, copy the output, and paste it into Excel or Google Sheets where our actual process lives. OpenAI is removing that step. Space holds folders, uploaded files, and Pages, which are documents that people and agents edit together in real time, and the AI sits on the canvas with everyone instead of in a side panel. It is currently live for ChatGPT Pro, Business, and Enterprise on web and desktop, and the announced direction is obvious: the office suite, rebuilt with an agent sitting in every file.
I am not writing this to review the product. Products change, the roadmap shifts, and by the time you read this the details will have moved. I am writing this because every new surface where someone builds a model is a surface you will eventually be asked to audit, version, or defend in front of your controller, and the habits for vetting those surfaces do not change as fast as the tools. The question is never whether the tool is good at drafting. The question is whether what comes out of it can stand up the same way a workbook built by a person on your team has to.
So here is the practitioner's checklist. Run it against anything an AI generates before that file becomes your model of record.
Check 1: Can you see the full formula chain? A generated workbook that shows you the answers without showing you the formulas is a black box with nice formatting. Open every tab, click through the cells, and trace where each number comes from. If the file was generated and you cannot tell which cell feeds which, you do not have a model, you have a PDF with extra steps. The bar is simple: a senior analyst who never saw the chat should be able to open the file and follow the logic without you narrating it.
Check 2: Can you reproduce the result from scratch? Delete the file and build it again from the same prompt, or hand the prompt to a colleague and ask them to build it. If the answer changes between runs, you have a model that depends on the mood of the machine, and that is not something you want feeding your forecast. Reproducibility is the minimum for anything that matters, and it is the first thing that breaks when generation replaces construction.
Check 3: Is there version history you can actually roll back to? Every finance team has a story about the file that changed at 4pm and nobody knows what changed. In Excel and Sheets we have version history, track changes, and change logs because we learned this the hard way. Before a generated workbook goes anywhere near your process, confirm that the platform keeps a real, restorable history of what was edited, by whom, and when. If the history is "the chat thread," that is not version control.
Check 4: What happened to your data when you uploaded it? This is the one people skip, and it is the one that gets you in trouble. You fed the model your actuals, your pipeline, maybe your salary planning detail. Know where that data lives now, whether it is excluded from training, who at the vendor can see it, and what your own company policy says about putting company data in a third-party workspace. OpenAI has said uploaded material is excluded from training sets and that private chats are not exposed through the shared context, which is good, but your own policy is what governs you, and you should be able to answer this question in one sentence before you upload anything.
Check 5: Would this survive a review by your controller? Imagine handing the file to the toughest reviewer in your finance org with no explanation of how it was made. If the first question they ask is "who built this and how do I know it is right," and you do not have a crisp answer, the file is not ready. This check catches everything the other four miss, because it forces you to treat the generated file like any other deliverable instead of giving it a pass because the AI made it.
Now, the concrete test plan. Do not take my word for any of this, because your own experience with one of your models will teach you more than any checklist. Pick one small model you know well, something like a monthly cash rollforward or a simple headcount forecast, and feed it to ChatGPT with the raw inputs. Ask it to build the workbook, then run the five checks against what comes back. Write down what you saw for each one, especially where the formula chain breaks down or the output shifts between runs. That writeup is worth more than this entire post, because it is calibrated to the tools you actually use and the reviewers you actually answer to.
The verdict, at least from where I sit: treat the chat workspace as a scratch pad for the first draft of a model. It is genuinely good at that, and it will get better at it fast. But a scratch pad is not a system of record, and it does not become one just because the formatting looks professional. Keep your models of record where your controls are, and let the generated files earn their way in through the same checks you would apply to anything a person built.
- Francois
Get the weekly newsletter that makes you better at Google Sheets, Productivity, and Finance.