Quick answer

For most companies the answer is: buy the model, build the rest. An off-the-shelf model via API is the fastest start, and your edge comes from your own data, integrations, and processes (RAG). Choose an open model on your own infrastructure when data cannot leave the company or the task is highly specialized. Training a model from scratch is a path for very few. The market is moving the same way: in 2025, 76% of enterprise AI use cases were purchased rather than built internally, up from 53% a year earlier (Menlo Ventures, 2025).

What is the build vs buy dilemma really about?

It is a trade-off between control and speed. Your team wants to add AI to a product or process, maybe a customer assistant, a document analyzer, or a recommendation engine, and the question comes up: do we build it ourselves or use an existing solution?

Building gives you control over data, cost at scale, and system behavior. Buying or integrating an existing model gets you to a working solution faster and usually costs less upfront. In practice almost nobody picks one option in its pure form. The better question is which layers to buy and which to build.

Market data suggests caution about building everything yourself. In MIT NANDA's research, external partnerships with customized tools reached deployment about 67% of the time, compared with about 33% for internally built tools (MIT NANDA, 2025). The authors note that the figures are self-reported, but the gap was consistent across interviews.

How do closed and open models differ?

A closed model is only available through an API, while an open model lets you download the weights and run it yourself.

Closed models such as GPT, Claude, or Gemini are a ready-made service. You send requests through an API, pay per use (usually per token), and the provider handles infrastructure and updates. Think of a fully furnished apartment: everything works on day one, but you cannot change the structure. We cover this path in depth in our article on closed AI models as the fastest route to AI in your company.

Open models (open weights) can be downloaded and run on your own servers or in a cloud of your choice. The model itself is free, but running it is not: you need GPUs, someone to maintain it, and monitoring.

Two terms are worth separating. Open weights means access to the trained weights. Open source AI, as defined by the Open Source Initiative, requires more: the freedom to use, study, modify, and share the system, plus access to the code and information about the training data (OSI, Open Source AI Definition 1.0). Most popular "open" models are in practice open weights. For a business, the license matters most, because it defines what you may do with the model.

Which models are open and which are API-only (October 2026)?

Model familyDownloadable weightsLicenseHow to use it
GPT (OpenAI)NoOpenAI terms of serviceAPI only
gpt-oss (OpenAI)YesApache 2.0 (model card)Your own infrastructure or cloud
Claude (Anthropic)NoAnthropic terms of serviceAPI only
Gemini (Google)NoGoogle terms of serviceAPI only
Gemma 4 (Google)YesApache 2.0 (Google, 2026)Your own infrastructure or cloud
Qwen 3.8 (Alibaba)YesMixed: smaller models such as 27B under Apache 2.0 (model card), the flagship and Flash variants under Qwen's own licenses (model card). The flagship license requires a separate license only from Model as a Service or AI Work Assistant businesses with revenue above USD 50 million over 12 consecutive months, and above 100 million monthly active users or USD 20 million in monthly revenue it requires the model name to be displayed in the interface (license)Your own infrastructure, cloud, or Qwen Cloud API
Mistral Large 3, Small 4, Ministral 3YesApache 2.0 (Mistral AI, Small 4 card)Your own infrastructure or Mistral API
Mistral Medium 3.5YesModified MIT license: no rights for companies with monthly revenue above USD 20 million (model card)Your own infrastructure (smaller companies) or Mistral API
Llama 4 (Meta)YesLlama 4 Community License (license)Your own infrastructure, subject to license terms
DeepSeek V4 and V4.1YesMIT (V4 card, V4.1 card)Your own infrastructure or API

Two practical notes. The same vendor often offers both kinds: OpenAI and Google serve their flagship models through APIs and release smaller ones as open weights. Licenses can also differ within one family: in Qwen 3.8 and at Mistral, some models are Apache 2.0 and others use the vendor's own license. The Llama license is not an open source license in the OSI sense: companies whose products had more than 700 million monthly active users must request a separate license from Meta (Llama 4 Community License). On top of that, the Llama 4 Acceptable Use Policy in Meta's official GitHub repository withholds the license rights to its multimodal models from companies based in the EU ("are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union"), except as end users of a product built on them (Llama 4 Acceptable Use Policy). Always read the license of the specific version before you commit, because terms change between releases.

What are the four ways to implement AI?

From simplest to most complex: API, RAG with an API, fine-tuning an open model, and your own model from scratch.

1. Calling an API directly

You sign up with a provider, get an API key, and start sending requests. A first prototype can run within hours. This works for quick prototypes, standard tasks (summarization, classification, drafting), and teams without AI expertise. Example: a tool that routes customer inquiries to the right department based on the message content.

2. RAG with an API

RAG (Retrieval-Augmented Generation) combines a model with search over your own data: the system first finds the relevant document passages, and only then does the model write an answer based on them. The model stays external, but the knowledge is yours. It is a good path for internal documentation assistants and specialized support bots. We explain the mechanics step by step in how RAG makes AI smarter.

3. Fine-tuning an open model

Fine-tuning means further training an existing model (such as Qwen, Mistral, or Gemma) on your own examples so that it adopts the style, format, and vocabulary of your domain. It requires more expertise and compute, but gives you a model you can run in-house, without sending data to outside providers. We show how we fine-tune open models on our Models & fine-tuning page.

4. Your own model from scratch

This is the rarest path. For scale: Mistral Large 3 was trained from scratch on 3,000 NVIDIA H200 GPUs (Mistral AI, 2025). On top of that come massive datasets and a research team. Unless you are creating fundamentally new AI capabilities, this is not your route. Companies that believe they need their own model usually need good RAG or fine-tuning.

Comparing the paths

CriterionAPIRAG with APIFine-tuning an open modelYour own model
Time to first working solutionShortestShortMediumVery long
Where the data livesWith the API providerDocuments stay with you, passages go to the APICan stay entirely with youWith you
Skills requiredA developerDevelopers and a data engineerML team and GPU infrastructureResearch team
Cost profilePay per usePay per use plus a knowledge baseGPU or cloud investment, low unit costVery high investment
Best forStandard tasksCompany knowledge that changesSpecialist language and sensitive dataNew AI capabilities

Which path should you choose? A decision checklist

The choice comes down to three questions: about the data, the task, and the team.

How sensitive is the data?

  • Public information: API.
  • Internal business data: RAG, with a GDPR-compliant data processing agreement with the API provider (GDPR, Article 28) and clear rules about what goes to the model.
  • Highly sensitive, regulated, or confidential data: an open model on your own infrastructure. We show what AI on your own infrastructure looks like on our AI infrastructure page.

How specialized is the task?

  • A common task: API.
  • Answers based on your own documents: RAG.
  • Specialist language, a fixed format, high volume: fine-tuning.

What skills and time do you have?

  • No AI team, need to validate quickly: API.
  • Developers and well-organized documents: RAG.
  • An ML team or a partner with GPU infrastructure: fine-tuning and self-hosted models.

Company size matters too, but less than the data. Startups usually begin with an API, companies with lots of documentation quickly get to RAG, and large organizations combine all paths: APIs for standard tasks, RAG for knowledge, fine-tuned models for specialized processes.

What does this look like in our projects?

We combine both sides ourselves: we use existing open models, but we build the integration, retrieval, and security layers.

  • A knowledge base on our own GPUs. Our agentic knowledge base runs on-premise with local models, because documents and questions must not leave the company. We do not train a model from scratch. The edge comes from the document pipeline, hybrid search tuned for Polish, and answers with citations to sources.
  • Our own components instead of building from scratch. We start every project from our own foundations: the agent-core framework for agents and our knowledge base engine. It is a "build once, reuse many times" option between buying and building.
  • Not every problem needs AI. Simple apps often take over part of the workload, such as merging reports or matching bank transfers to accounting records. Sometimes the best build vs buy answer is: build an ordinary application.

How do you start without locking yourself in?

Start with the simplest solution that solves the problem, but design it so you can swap the model. In practice:

  • Separate the model from the application. The layer that calls the model should let you switch providers or move to a local model without a rewrite.
  • Collect data from day one. User questions, answers, and corrections are material for quality evaluation and, later, for fine-tuning.
  • Track cost per task. As request volume grows, compare API costs with the cost of your own infrastructure.
  • Write down your data rules. What may be sent to an API, what stays in-house, and who approves exceptions.

Most successful implementations evolve: an API first, then RAG when you need company knowledge, then fine-tuning or a local model once you know exactly what you need. The best choice is not the most advanced option, but the one that fits your data, your task, and your team.

Sources