Quick answer
For most businesses, yes: a closed-source model, meaning an AI model the vendor provides through an API or app without public weights, is the fastest way to start using AI. You need no GPU servers and no ML team, you pay per use, and you get access to the most capable models available. Limits show up later: when data cannot leave your company, when the model does not know your documents (that is where RAG helps), or when at scale the token bill exceeds the cost of your own infrastructure.
What is a closed-source model?
A closed-source (proprietary) model is one whose weights, the learned parameters, only the vendor has. The model runs on the vendor's servers, and you send a request through an API (the interface your application uses to talk to the model) and receive a response.
The opposite is an open-weight model: you can download the weights, run them on your own hardware and fine-tune them. We show how we fine-tune open models on our Models & fine-tuning page.
Think of it like office space. You would not build an office tower before testing whether your business works. You rent first and think about building once you know your needs and scale.
| Criterion | Closed model via API | Open model on your infrastructure | Hybrid (API + RAG or API + local model) |
|---|---|---|---|
| Time to start | Days | Weeks to months | Weeks |
| Upfront cost | Close to zero | GPU hardware or rental, deployment | Moderate: search, integration |
| Cost at large scale | Grows linearly with tokens | Mostly fixed (hardware, power, upkeep) | Depends on how traffic is split |
| Quality on general tasks | Usually the best available | Good, depends on model and hardware | Best where it matters |
| Control over data | Depends on the vendor's terms and region | Full | Sensitive data can stay local |
| Control over model version | Vendor retires old models | Full | Partial |
| Skills required | Integration, prompting, evaluation | Plus MLOps and GPU operations | Integration and search |
Who provides closed-source models in 2026?
Three model families dominate the market: GPT (OpenAI), Claude (Anthropic) and Gemini (Google). Each vendor offers several tiers: a flagship model for the hardest tasks, a balanced model for everyday work, and a small, fast model for simple high-volume tasks. Several of these models are also available through public clouds: OpenAI models on Microsoft Azure, Gemini on Google Cloud, and Claude on Amazon Bedrock, Google Cloud and Microsoft Foundry, among others (Anthropic). That makes it easier to start if you already have a contract with one of those clouds.
We deliberately do not list specific versions or prices, because they change every few months. Current rates are on the pricing pages of OpenAI, Anthropic and Google Gemini.
The practical takeaway: design your solution so the model can be swapped. Vendors regularly retire older versions and publish shutdown schedules (OpenAI, Anthropic). Your own test set and a thin layer between your application and the API turn a model change into a day of work rather than a project.
What are you actually paying for?
Tokens. A token is a piece of text, often part of a word, that the model splits input and output into. You pay separately for tokens sent to the model (question, instructions, documents) and tokens it generates (the answer), and output is usually several times more expensive than input (about five times in OpenAI's and Anthropic's current price lists).
Language matters. We measured three of our own articles that exist in Polish and English with the same content. Using the o200k_base tokenizer from OpenAI's tiktoken library, the Polish versions produced about 1.5 times as many tokens as the English ones (about 2.1 tokens per word versus about 1.3). Other models use other tokenizers, so results will vary, but the direction is the same: a budget based on English samples will be too low for content in many other languages.
Three mechanisms that genuinely lower the bill:
- Batch processing. If you do not need the answer right away (overnight reports, classifying an archive), OpenAI, Anthropic and Google charge 50% of the standard price for such requests (OpenAI, Anthropic, Google).
- Prompt caching. When many requests start with the same instructions or documents, that part can be billed at a lower rate. At Anthropic, a cache read costs at most 10% of the regular input token price (Anthropic, pricing).
- Matching the model to the task. Ticket classification or pulling fields from a form is usually fine with a small model. Save the flagship for tasks where the difference shows.
Factor in surcharges too: OpenAI adds 10% for regional processing (data residency) on models released on or after March 5, 2026 (OpenAI).
What happens to your company's data?
On API plans, vendors do not train models on customer data by default, but that is the start of the analysis, not the end. Here is how the three major vendors handle it:
- OpenAI: data sent through the API is not used to train or improve models unless the customer opts in. Abuse monitoring logs are kept for up to 30 days, and Zero Data Retention is available to eligible customers. You can also choose a processing region, including Europe (EEA and Switzerland) (OpenAI).
- Anthropic: by default it does not use inputs or outputs from commercial products (API, Claude for Work) to train models (Anthropic). On the direct API you can restrict inference to the US or leave it global; on Amazon Bedrock and Google Cloud the region is set by the endpoint you choose (Anthropic).
- Google: on the paid Gemini API, prompts and responses are not used to improve products. On the free tier they can be, but for users in the EEA, Switzerland and the UK the paid terms apply even within the free quota (Google).
Under GDPR the vendor is usually a processor, so you need a data processing agreement (DPA), clarity on where data is processed and a basis for any transfer outside the EEA. Also decide what never gets sent: medical data, national ID numbers, trade secrets. Free consumer apps have different terms than the API, which is why employees should use company accounts.
If you offer customers a chat built on such a model in the EU, remember the AI Act: since August 2, 2026, people interacting with an AI system must be told they are not talking to a human, unless that is obvious from the context (European Commission).
Which quick wins can you deploy right away?
Ones where the model works on text someone already reads or writes, and a person approves the result. The most common:
- Drafts: product descriptions, replies to routine emails, first versions of proposals.
- Summaries: meeting notes, reports, long email threads.
- Classification and routing: assigning category and priority to tickets before a person picks them up.
- Data extraction: fields from invoices, contracts and forms into a spreadsheet or system.
- Analysis: first-pass insights from spreadsheets, document comparison, brainstorming.
We work this way ourselves: we run pimento from a repository in which an AI agent acting as COO assistant writes up meeting notes, turns decisions into tasks and prepares a weekly summary.
The most important lesson from such deployments: prompt quality often matters more than model choice. Context, an example of a good answer and the expected format do more than switching to a pricier model.
When does basic use stop being enough?
When the model starts getting your company's facts wrong: procedures, products, contracts. The model knows the internet, not your documents. The fix is RAG (retrieval-augmented generation): the system first finds the right passages in your company's documents, then passes them to the model with the question. The model answers based on them and can cite the source. We go into detail in how RAG makes AI smarter.
A typical hybrid architecture:
- An application (chat, helpdesk, internal tool) receives a question.
- A search system finds matching passages in company documents.
- The question and passages go to the model via API.
- The model generates an answer with references to sources.
- The application shows the answer and, for sensitive matters, waits for a person to approve it.
You can swap the model without rebuilding everything, because company knowledge lives in the search layer, not the model. The same pattern can run fully locally: our agentic knowledge base runs on our own GPUs with local models, and every claim in an answer is footnoted to a document passage.
When should you choose an open model or your own infrastructure?
When at least one of these is true:
- Data cannot leave your company for legal or contractual reasons: medical records, customers' financial data, trade secrets.
- Volume is large and predictable, so the fixed cost of hardware is lower than a growing token bill.
- You need control over the model version, so results stay reproducible for years regardless of the vendor's retirement schedule.
- The task is narrow and specialized, and a smaller fine-tuned model does it just as well for less.
This is not theory for us. We run language models, embeddings, a reranker, and speech recognition and synthesis on our own GPU infrastructure, so we process client documents, questions and recordings without sending them to external APIs. For Fundacja MT5 we supplied a GPU cluster that sits at the client's site, and we train models on it using sleep study data. We show what AI on your own infrastructure looks like on our AI infrastructure page.
How do you decide: buy access or build?
Start with a closed-source model if you need results quickly, work with standard business data, have no ML team and want to first check whether AI solves your problem at all. Consider your own solution when data is sensitive, scale is large or you need full control. We lay out the broader framework in our guide should you build or buy your AI solution.
| Your situation | Recommended starting point |
|---|---|
| First AI project, standard data, small IT team | Closed model via API, company accounts, simple prompts |
| Model gets company facts wrong | API + RAG on company documents |
| Personal data at scale, EU region required | API with EU processing and a DPA, or a local model |
| Data that must not leave the company | Open model on your own infrastructure |
| Large, steady volume of simple tasks | Price an open model, or a small closed model in batch mode |
| Narrow specialist task | Fine-tuned open model |
How do you start in practice?
- Pick one use case with a clear cost in today's process.
- Build a test set: a few dozen real questions or documents with expected answers. Use it to compare models and prompts.
- Check the data terms: business or API plan, DPA, region, a list of data you never send.
- Run a pilot with the people who do the work, and measure time and quality before and after.
- Plan for swapping models from day one: an abstraction layer, a test set, cost monitoring.
Closed-source models are not a compromise but a sensible first step. Get value and learn what works first, then invest in custom-built solutions.
Sources
- OpenAI: API pricing
- Anthropic: Pricing
- Google: Gemini Developer API pricing
- OpenAI: Deprecations
- Anthropic: Model deprecations
- OpenAI: tiktoken
- OpenAI: Batch API
- Anthropic: Batch processing
- Google: Gemini Batch API
- Anthropic: Prompt caching
- OpenAI: Data controls in the OpenAI platform
- Anthropic: Is my data used for model training?
- Anthropic: Data residency
- Google: Gemini API Additional Terms of Service
- European Commission: Transparency obligations under Article 50 of the AI Act