Most companies now have someone asking what the business should do with AI. The projects that stall usually begin with the technology: a team picks a model, builds a demo that impresses in a meeting, and then cannot say whether it is good enough to put in front of customers or staff. The projects that reach production begin with one specific task, a way to measure it, and data that already exists. This guide covers how to choose that first task and how to run it as a small, measured pilot.
Start From a Task, Not a Model
Write the task as a sentence about work people do today, for example "the finance team types supplier invoices into the ERP" or "customer service answers the same shipping questions all day". Then note how many times it happens per month, how long each one takes, and what a mistake costs. A task that happens thousands of times, follows a pattern, and has a person who can check the result is a good candidate. A task that happens twice a year or where a wrong answer is expensive and hard to spot is not.
Good and Poor First Use Cases
| Use case | Good first project? | Why |
|---|---|---|
| Extracting fields from invoices, receipts, or forms | Yes | High volume, easy to check against the source document |
| Answering staff questions from internal documents | Yes | Users are employees who can report wrong answers |
| Sorting and routing support tickets or emails | Yes | Results are easy to measure against past labels |
| Drafting replies or reports for a person to review | Usually | Saves time while a person stays responsible for the output |
| Fully automatic answers to customers about money or contracts | Not first | Mistakes are costly and reach customers directly |
| Predictions with little historical data | Not first | There is nothing reliable to learn from or test against |
Check the Data Before the Model
- Where does the data live, and can the project read it through an API, a database view, or regular exports?
- Is it clean enough? Scanned documents that nobody can read, or a knowledge base full of outdated policies, will produce wrong answers no matter which model you use.
- Does it contain personal data? Under the Personal Data Protection Law (UU PDP) you need a legal basis to process it and to know where it goes.
- Who owns it and can approve its use? Projects often wait weeks for this answer, so ask on day one.
- Can you collect 50 to 200 real examples with the correct answer to use as a test set?
API, Self-Hosted Model, or Both
Hosted APIs from Anthropic, OpenAI, and Google give the strongest models with no infrastructure, and they are the fastest way to find out whether the task can be done at all. Self-hosted open-weight models run with Ollama or vLLM keep data inside your network and have a fixed cost, which matters when the data is sensitive or the volume is very high. Many teams prototype with an API, measure the quality they need, and then test whether a smaller self-hosted model reaches it. If the data rules say nothing may leave your servers, start self-hosted from the beginning.
Fine-tuning a model is rarely needed for a first project. Clear instructions, a few examples in the prompt, and retrieval from your own documents (RAG) cover most business tasks, and they are cheaper to change when the rules change.
Run the Pilot
- 1Agree on one task, the people who will use the result, and the number that defines success, such as accuracy on the test set or minutes saved per document.
- 2Build the test set from real cases with the correct answers, checked by the people who do the work today.
- 3Build the simplest version that can run the test set, and measure it before adding features.
- 4Fix the largest groups of errors first, usually by improving instructions, input quality, or retrieval, and measure again after every change.
- 5Put it in front of a small group of real users with a person checking every result, and log every correction they make.
- 6Decide based on the numbers: expand, change the scope, or stop. Stopping a pilot that does not reach the target is a valid result.
What It Costs
- Model usage, billed per token on an API, or GPU servers for a self-hosted model.
- Engineering to connect the AI to your data sources, applications, and login system.
- Time from subject experts to build the test set and review results. This is often the cost teams forget.
- Running costs after launch: monitoring answer quality, updating documents and prompts, and handling model version changes.
Estimate token costs early by running the test set and multiplying the cost per case by the monthly volume. Most business tasks cost far less in tokens than the staff time they save, but a design that sends whole documents with every request can surprise you at scale.
Risks to Manage From the Start
- Wrong answers stated with confidence. Keep a person in the loop for anything that affects money, contracts, or safety, and show the source documents next to the answer.
- Data leaving the company. Check the provider terms on data retention and training, and use business or enterprise plans that exclude your data from training.
- Prompt injection, where text inside a document or a user message tries to change the instructions. Do not give the model access to actions it should not take on its own.
- Staff who do not trust or do not use the tool. Involve them in building the test set so they see where it works and where it does not.
Key takeaways
- Start with one frequent, checkable task and a number that defines success.
- Check data access, quality, and personal data rules before choosing a model.
- Prototype with an API, then test self-hosted models if data rules or volume require it.
- Run a pilot against a test set of real cases and decide based on the results.
- Keep people reviewing outputs that affect money, contracts, or customers.


