Quick Answer: You can't retrain ChatGPT itself; you train it on your own data by giving it your documents through projects, connectors or retrieval. A project in the ChatGPT app holds up to 40 files on business plans, so large or changing document sets need retrieval (RAG) over an index. Fine-tuning shapes style and format, not facts.
If you're working out how to train ChatGPT on your own data, you probably want a chatbot that answers from your policies, product docs or support history. The model's weights stay as OpenAI shipped them. What you control is what the model reads before it answers, and that choice decides freshness, cost and who sees which document. For firms that do this work under contract, see OpenAI consulting services.
Can you actually train ChatGPT on your own data?
No, not in the sense of changing the model. Your files are read as context at question time, and OpenAI says it does not use content from Business, Enterprise or Edu workspaces to train its models by default.
Three different mechanisms often get lumped together as "training":
- Context. You attach documents to a project or a custom GPT, and ChatGPT reads the relevant parts when you ask.
- Retrieval. A connector or pipeline searches an index of your documents and passes the best passages to the model. This is RAG (retrieval-augmented generation).
- Fine-tuning. The model's weights are adjusted with example prompts and answers. OpenAI's supervised fine-tuning guide positions it for classification, output format and style, and now says the platform is winding down and closed to new users.
Only the first two put your facts in front of the model.
What are the ways to give ChatGPT your company's documents?
The quickest chatbot knowledge base is a ChatGPT project: upload your documents there, then move to retrieval once file count, change rate or permissions outgrow it.
Limits as documented by OpenAI on 28 September 2026; links in the text.
| Route | Setup effort | Document limits | Keeps up with changes? | Access control | Best for |
|---|---|---|---|---|---|
| ChatGPT project | No code | 5 files (Free), 25 (Go, Plus), 40 (Pro, Business, Enterprise, Edu) | Uploaded files: only on re-upload; pasted Drive or Slack links: read live | Owner grants chat or edit access | One team's working set |
| Custom GPT | No code; new GPTs only in Business, Enterprise, Edu | 20 files, up to 512 MB each | Only on re-upload | GPT sharing settings | A fixed reference set, until retirement |
| Connected apps with admin-managed sync | Workspace admin setup | What the admin scopes in Google Drive, SharePoint or Teams | Yes, after the source refreshes | Mirrors source permissions | Search over existing drives |
| Fine-tuning | Labeled examples plus evals | Training examples, not documents | No, retrain to change | None for documents | Tone, format, classification |
| Self-hosted assistant with retrieval | Engineering project | Your own index | On your sync schedule | Document-level rules you define | Large, sensitive or fast-changing sources |
The lowest-effort route: a project
OpenAI's Projects help page describes the setup: create a project, upload PDFs, spreadsheets or docs (10 at a time), and add instructions. You can also paste Google Drive or Slack links as sources. Write the instructions like a spec: answer only from the files, name the file used, and say so when the files don't cover the question.
When is a custom GPT enough and when do you need RAG?
A custom GPT or project is enough when the document set is small, changes rarely and every user may see every file. You need RAG once any of those breaks, and early if answers must cite the exact source.
Plan for one change. OpenAI's GPT help page says new GPT creation is closed on personal accounts and that custom GPTs are being retired in favor of plugins, with 11 December 2026 planned for affected Enterprise workspaces. Treat a custom GPT as a prototype.
Size, change rate and permissions
Twenty files is a handbook, not a support library. A help center with hundreds of articles, or a wiki that changes daily, goes stale the week you upload it. Everyone who can use the GPT or project gets answers from every uploaded file, so when HR, legal and engineering each need a different slice, the permission check has to happen at retrieval time.
If you're comparing packaged options for customer-facing bots, this roundup of AI knowledge base builders for chat and support covers the field by type.
How do you prepare documents so the answers stay accurate?
A chatbot answers from what it retrieves, so conflicting or badly formatted files produce confident wrong answers.
- Keep one source of truth per topic. Archive the old refund policy instead of uploading both versions.
- Prefer text-forward files. OpenAI's GPT guidance says complex layouts are harder to use; convert scans and slide decks to clean text.
- Structure with headings. Retrieval splits documents into passages, and a heading gives each passage context.
- Add owner and review-date metadata to every document.
- Write a test set before launch: 30 to 50 real questions with the correct answer and source file, scored after every change.
How do you keep answers current when documents change?
Give every source an owner and a re-index schedule. Uploaded project and GPT files are static copies; they change only when someone replaces them.
Connected sources lag too. OpenAI's admin-managed sync documentation says new content and permission changes can take time to appear after the source refreshes, and that individually authorized sync is no longer available.
- Sync each source at its own pace: hourly for tickets, weekly for policies.
- Run a stale-answer check after each sync and flag answers whose source is past its review date.
- Confirm takedowns: a withdrawn document must leave the index, not just the drive.
Settle where the data may go first. Whether ChatGPT is safe for confidential data depends on the plan, workspace settings and what employees paste into personal accounts.
What does a production-grade document assistant need beyond ChatGPT?
It needs retrieval that respects each user's permissions, a record of every answer, and a deployment your security team can approve. Use this checklist on your own project.
| Requirement | Question to answer |
|---|---|
| Permission-aware retrieval | Does each user see only passages from documents they can open? |
| Single sign-on | Does it use your identity provider over SAML 2.0 or OIDC? |
| Audit log | Is every question, retrieval and answer logged with a user ID? |
| Citations | Does every answer link the passage it used? |
| Evaluation | Does a test set run on every change? |
| Deployment | Does inference run in a vendor cloud, your cloud account, or on-premise? |
| Model choice | Can you switch models without rebuilding the app? |
| Delivery | Internal chat only, or also an embedded support widget and an API? |
The last row is where teams stall: they want a ChatGPT-style assistant for staff and the same answers inside their product for customers. Some firms build exactly that, as covered in this look at custom assistants with embedded customer support.
What mistakes should you avoid when training ChatGPT on your data?
- Uploading everything. Old drafts compete with the current version, and retrieval can't tell which is right.
- Ignoring permissions. In a shared project, every member can view and download every file.
- Fine-tuning for facts. Facts change faster than retrain cycles; use retrieval for facts.
- Launching without a test set. Without scored questions, "it seems fine" is your only quality bar.
- Pasting confidential data into personal accounts. On Free, Go, Plus and Pro, OpenAI may use content to improve its models if "Improve the model for everyone" is on.
How Origins AI builds private assistants on company data
Origins AI (originshq.com) is an AI-augmented engineering company that sets up self-hosted AI chatbots on a private knowledge base inside the customer's environment. Two of its products cover document assistants.
The Origins AI Velocity AI Suite is the knowledge layer. According to its product page, it ingests from 1,900+ data sources and 91+ document formats, indexes into vector databases such as Pinecone, Chroma or Weaviate, and runs with OpenAI, Anthropic, open-source or your own models.
Origins AI Chat AI is the assistant on top. Its product page lists SSO over SAML 2.0 or OIDC, access scoped by department and document, citation grounding, a conversation log with user ID and timestamp, and delivery as chat, an embeddable widget or a REST API. Origins AI's YesMadam case study reports a chatbot deployed to reduce support calls.
The company states its products run on-premise, in your own AWS, Azure or GCP account, or air-gapped. It says no data leaves your network in on-premise and air-gapped modes with self-hosted models; route requests to a hosted model and that provider's policies apply. There's no published rate card, and an implementation team handles ingestion, SSO and rollout.
Talk to an engineer
Outgrown a custom GPT? Bring the production checklist above to a call with our engineers and see what a private assistant on your documents would take.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


