LLM development

LLM development: product integration, not a wrapper around a chat box

Large language models are components. We integrate them into applications: choosing a model, shaping the prompt and tools, handling failure, and deciding what is logged. The deliverable is a feature your staff or customers can use, with a way to change it safely.

  • Model selection
  • Application integration
  • Evaluation harness
  • Fallback behaviour

Build versus fine-tune

Most business problems are solved by retrieval, tools, and a careful prompt long before fine-tuning is justified. Fine-tuning is considered when you have enough reviewed examples and a behaviour the base model will not hold. We will not recommend it as a first step to sound advanced.

Operational concerns

Latency, cost per task, rate limits, and what happens when the provider is unavailable all belong in the design. So does the decision about which text leaves your environment. Some workloads stay on a commercial API. Some need a private deployment. That choice is made with your data rules, not with a trend.

A model inside a system

We integrate

  • One model with a named task
  • Tools it may call
  • Evaluation on examples you care about

We do not start with

  • Training a foundation model
  • Every provider at once
  • A hidden prompt as the only control

What changes an LLM build

01

Hosted or private

Where the prompt and the documents go is a design choice, not a setting at the end.

02

Tools

A model that can only talk is smaller than one that may create records.

03

Failure

Say what the product does when the model is slow, wrong, or unavailable.

What an LLM integration needs

  1. 01

    The task

    One job for the model, written as an example.

  2. 02

    The host

    Where the prompt and the documents are allowed to go.

  3. 03

    The tools

    Whether it may only talk, or may create a record.

  4. 04

    The failure

    What the product does when the model is wrong or down.

Marks a model choice is deliberate

  1. 01

    The choice is written

    Which model, and why it fits this task, is in the design, not only in a key.

  2. 02

    Real examples

    Evaluation uses samples from the work, including ones that should be refused.

  3. 03

    The desk can wait

    The response time fits the person using it. If it does not, the design changes.

  4. 04

    A fallback exists

    When the model is down or unsure, the user still has a path.

What we build

Model selection

A comparison on your examples, including cost and failure modes, not a leaderboard screenshot.

Application integration

APIs, tools the model may call, and authorization around those tools.

Evaluation harness

A repeatable set of cases so a prompt change can be accepted or rejected.

Fallback behaviour

What the product does when the model is slow, wrong, or down.

How an engagement runs

  1. 01

    Collect real tasks

    Twenty to fifty examples beat a theoretical workshop.

  2. 02

    Integrate the smallest tool

    One action the model may take, with permission checks outside the model.

  3. 03

    Put numbers on it

    Quality on the sample, cost per successful task, and the fallback rate.

Questions we hear

Do you train foundation models?

No. We apply and integrate models. Training a foundation model is a different business, and almost no company needs it.

Can the model update our database?

Only through tools we write, with the same authorization a user would need, and usually with a confirmation when the action is hard to undo.

How do you stop invented answers?

By limiting the task, retrieving sources, quoting them, and refusing when the material is missing. We also measure the failure instead of assuming the prompt solved it.