LLM development
LLM development: product integration, not a wrapper around a chat box
Large language models are components. We integrate them into applications: choosing a model, shaping the prompt and tools, handling failure, and deciding what is logged. The deliverable is a feature your staff or customers can use, with a way to change it safely.
- Model selection
- Application integration
- Evaluation harness
- Fallback behaviour

Build versus fine-tune
Most business problems are solved by retrieval, tools, and a careful prompt long before fine-tuning is justified. Fine-tuning is considered when you have enough reviewed examples and a behaviour the base model will not hold. We will not recommend it as a first step to sound advanced.
Operational concerns
Latency, cost per task, rate limits, and what happens when the provider is unavailable all belong in the design. So does the decision about which text leaves your environment. Some workloads stay on a commercial API. Some need a private deployment. That choice is made with your data rules, not with a trend.
A model inside a system
We integrate
- One model with a named task
- Tools it may call
- Evaluation on examples you care about
We do not start with
- Training a foundation model
- Every provider at once
- A hidden prompt as the only control
What changes an LLM build
Hosted or private
Where the prompt and the documents go is a design choice, not a setting at the end.
Tools
A model that can only talk is smaller than one that may create records.
Failure
Say what the product does when the model is slow, wrong, or unavailable.
What an LLM integration needs
- 01
The task
One job for the model, written as an example.
- 02
The host
Where the prompt and the documents are allowed to go.
- 03
The tools
Whether it may only talk, or may create a record.
- 04
The failure
What the product does when the model is wrong or down.
Marks a model choice is deliberate
- 01
The choice is written
Which model, and why it fits this task, is in the design, not only in a key.
- 02
Real examples
Evaluation uses samples from the work, including ones that should be refused.
- 03
The desk can wait
The response time fits the person using it. If it does not, the design changes.
- 04
A fallback exists
When the model is down or unsure, the user still has a path.
What we build
Model selection
A comparison on your examples, including cost and failure modes, not a leaderboard screenshot.
Application integration
APIs, tools the model may call, and authorization around those tools.
Evaluation harness
A repeatable set of cases so a prompt change can be accepted or rejected.
Fallback behaviour
What the product does when the model is slow, wrong, or down.
How an engagement runs
- 01
Collect real tasks
Twenty to fifty examples beat a theoretical workshop.
- 02
Integrate the smallest tool
One action the model may take, with permission checks outside the model.
- 03
Put numbers on it
Quality on the sample, cost per successful task, and the fallback rate.
Related reading
Questions we hear
Do you train foundation models?
No. We apply and integrate models. Training a foundation model is a different business, and almost no company needs it.
Can the model update our database?
Only through tools we write, with the same authorization a user would need, and usually with a confirmation when the action is hard to undo.
How do you stop invented answers?
By limiting the task, retrieving sources, quoting them, and refusing when the material is missing. We also measure the failure instead of assuming the prompt solved it.



